Showing posts with label confirmation. Show all posts
Showing posts with label confirmation. Show all posts

Tuesday, June 12, 2018

Yet another counterexample to Nicod's Principle

Nicod’s Principle says that the claim that all Fs are Gs is confirmed by each instance.

Here’s yet another counterexample. Consider the claim:

  1. All unicorns are male.

We take this claim to be true, albeit vacuously so, since there are no unicorns.

But suppose an instance of (1), namely a male unicorn, were found. We would immediately conclude that (1) is probably false. For if there is a male unicorn, likely there is a female one as well.

The problem here is that when we learn of Sam that it is a male unicorn, we also learn that there are unicorns. And as soon as we learned that there are unicorns, that undercut the reason we had for believing (1), namely that we thought (1) was vacuously true.

Saturday, November 18, 2017

Bayesianism and anomaly

One part of the problem of anomaly is this. If a well-established scientific theory seems to predict something contrary to what we observe, we tend to stick to the theory, with barely a change in credence, while being dubious of the auxiliary hypotheses. What, if anything, justifies this procedure?

Here’s my setup. We have a well-established scientific theory T and (conjoined) auxiliary hypotheses A, and T together with A uncontroversially entails the denial of some piece of observational evidence E which we uncontroversially have (“the anomaly”). The auxiliary hypotheses will typically include claims about the experimental setup, the calibration of equipment, the lack of further causal influences, mathematical claims about the derivation of not-E from T and the above, and maybe some final catch-all thesis like the material conditional that if T and all the other auxiliary hypotheses obtain, then E does not obtain.

For simplicity I will suppose that A and T are independent, though of course that simplifying assumption is rarely true.

I suspect that often this happens: T is much better confirmed than A. For T tends to be a unified theoretical body that has been confirmed as a whole by a multitude of different kinds of observations, while A is a conjunction of a large number of claims that have been individually confirmed. Suppose, say, that P(T)=0.999 while P(A)=0.9, where all my probabilities are implicitly conditional on some background K. Given the observation E, and the fact that T and A entail its negation, we now know that the conjunction of T and A is false. But we don’t know where the falsehood lies. Here’s a quick and intuitive thought. There is a region of probability space where the conjunction of T and A is false. That area is divided into three sub-regions:

  1. T is true and A is false

  2. T is false and A is true

  3. both are false.

The initial probabilities of the three regions are, respectively, 0.0999, 0.0009999 and 0.0001. We know we are in one of these three regions, and that’s all we now know. Most likely we are in the first one, and the probability that we are in that one given that we are in one of the three is around 0.99. So our credence in T has gone down from three nines (0.999) to two nines (0.99), but it’s still high, so we get to hold on to T.

Still, this answer isn’t optimistic. A move from 0.999 to 0.99 is actually an enormous decrease in confidence.

But there is a much more optimistic thought. Note that the above wasn’t a real Bayesian calculation, just a rough informal intuition. The tip-off is that I said nothing about the conditional probabilities of E on the relevant hypotheses, i.e., the “likelihoods”.

Now setup ensures:

  1. P(E|A ∧ T)=0.

What can we say about the other relevant likelihoods? Well, if some auxiliary hypothesis is false, then E is up for grabs. So, conservatively:

  1. P(E|∼A ∧ T)=0.5
  2. P(E|∼A ∧ ∼T)=0.5

But here is something that I think is really, really interesting. I think that in typical cases where T is a well-established scientific theory and A ∧ T entails the negation of E, the probability P(E|A ∧ ∼T) is still low.

The reason is that all the evidence that we have gathered for T even better confirms the hypothesis that T holds to a high degree of approximation in most cases. Thus, even if T is false, the typical predictions of T, assuming they have conservative error bounds, are likely to still be true. Newtonian physics is false, but even conditionally on its being false we take individual predictions of Newtonian physics to have a high probability. Thus, conservatively:

  1. P(E|A ∧ ∼T)=0.1

Very well, let’s put all our assumptions together, including the ones about A and T being independent and the values of P(A) and P(T). Here’s what we get:

  1. P(E|T)=P(E|A ∧ T)P(A|T)+P(E|∼A ∧ T)P(∼A|T)=0.05
  2. P(E|∼T)=P(E|A ∧ ∼T)P(A|∼T)+P(E|∼A ∧ ∼T)P(∼A|∼T) = 0.14.

Plugging this into Bayes’ theorem, we get P(T|E)=0.997. So our credence has crept down, but only a little: from 0.999 to 0.997. This is much more optimistic (and conservative) than the big move from 0.999 to 0.99 that the intuitive calculation predicted.

So, if I am right, at least one of the reasons why anomalies don’t do much damage to scientific theories is that when the scientific theory T is well-confirmed, the anomaly is not only surprising on the theory, but it is surprising on the denial of the theory—because the background includes the data that makes T “well-confirmed” and would make E surprising even if we knew that T was false.

Note that this argument works less well if the anomalous case is significantly different from the cases that went into the confirmation of T. In such a case, there might be much less reason to think E won’t occur if T is false. And that means that anomalies are more powerful as evidence against a theory the more distant they are from the situations we explored before when we were confirming T. This, I think, matches our intuitions: We would put almost no weight in someone finding an anomaly in the course of an undergraduate physics lab—not just because an undergraduate student is likely doing it (it could be the professor testing the equipment, though), but because this is ground well-gone over, where we expect the theory’s predictions to hold even if the theory is false. But if new observations of the center of our galaxy don’t fit our theory, that is much more compelling—in a regime so different from many of our previous observations, we might well expect that things would be different if our theory were false.

And this helps with the second half of the problem of anomaly: How do we keep from holding on to T too long in the light of contrary evidence, how do we allow anomalies to have a rightful place in undermining theories? The answer is: To undermine a theory effectively, we need anomalies that occur in situations significantly different from those that have already been explored.

Note that this post weakens, but does not destroy, the central arguments of this paper.

Thursday, June 16, 2016

Measure of confirmation

Suppose I played a lottery that involved my picking ten digits. The digits I picked are: 7509994361. At this point, my probability that I won, assuming I had to get every digit right to win, is 10−10. Suppose now that I am listening to the radio and the winning number's digits are announced: 7, 5, 0, 9, 9, 9, 4, 3, 6 and 1. My sequence of probabilities that I am the winner after hearing the successive digits will approximately be: 10−9, 10−8, 10−7, 10−6, 10−5, 10−4, 0.001, 0.01, 0.1 and approximately 1.

If we think that to get a significant amount of evidential confirmation for a hypothesis we need to give a significant absolute increment to the probability, only the final two or three digits gave significant confirmation to the hypothesis that I won, as the last digit increased my probability by 0.9, the one before by 0.09, and the one before by 0.009. But it seems quite wrong to me to say that as I am listening to the digits being announced one by one, I get no significant evidence until the last two or three digits. In fact, I think that each digit I hear gives me an equal amount of evidence for the hypothesis that I won.

Here's an argument for why it's wrong to say that a significant absolute increment of probability is needed to get significant confirmation. Let A be the collective evidence of learning digits 1 through 7 and let B be the evidence of learning digits 8 through 10. Then on the absolute increment view, A provides me with no significant confirmation that I won, but when I learn B after learning A, then B provides me with significant confirmation. However, had I simply learned B, without learning A first, that wouldn't have afforded me significant confirmation--my probability of winning would still have been 10−7. So the combination of two pieces of evidence, both insignificant on their own, gives me conclusive evidence.

There is nothing absurd about this on its own. Sometimes two insignificant pieces of evidence combine to conclusive evidence. Learning (C) that the winning number is written on a piece of paper is insignificant on its own; learning (D) that 7509994361 is written on the piece of paper is insignificant on its own; but combined they are conclusive. Yes: but that is a special case, where the two pieces of evidence fit together in a special way. We can say something about how they fit together: they fail to be statistically independent conditionally on the hypothesis I am considering. Conditionally on 7509994361 being the winning number, the event that 7509994361 is what is written on the paper and the event that what is written on the paper is the winning number are far from independent. But in my previous example, A and B are independent conditionally on the hypothesis I am considering, as well as being independent conditionally on the negation of that hypothesis. They aren't pieces of evidence that interlock like C and D are.

Friday, November 13, 2015

Bayesian divergence

Suppose I am considering two different hypotheses, and I am sure exactly one of them is true. On H, the coin I toss is chancy, with different tosses being independent, and has a chance 1/2 of landing heads and a chance 1/2 of landing tails. On N, the way the coin falls is completely brute and unexplained--it's "fundamental chaos", in the sense of my ACPA talk. So, now, you observe n instances of the coin being tossed, about half of which are heads and half of which are tails. Intuitively, that should support H. But if N is an option, if the prior probability of N is non-zero, we actually get Bayesian divergence as n increases: we get further and further from confirmation of H.

Here's why. Let E be my total evidence--the full sequence of n observed tosses. By Bayes' Theorem we should have:

P(H|E) = P(E|H)P(H)/[P(E|H)P(H) + P(E|N)P(N)].
But there is a problem: P(E|N) is undefined. What shall we do about this? Well, it is completely undefined. Thus, we should take it to be an interval of probabilities, the full interval [0,1] from 0 to 1. The posterior probability P(H|E), then, will also be an interval between:
P(E|H)P(H)/[P(E|H)P(H) + (0)·P(N)] = 1
and
P(E|H)P(H)/[P(E|H)P(H) + (1)·P(N)] ≤ P(E|H)/P(N) = 2n / P(N).
(Remember that E is a sequence of n fair and independent tosses if H is true.) Thus, as the number of observations increases, the posterior probability for the "sensible" hypothesis H gets to be an interval [a,1], where a is very small. But something whose probability is almost the whole interval [0,1] is not rationally confirmed. So the more data we have, the further we are from confirmation.

This means that no-explanation hypotheses like N are pernicious to Bayesians: if they are not ruled out as having zero or infinitesimal probability from the outset, they undercut science in a way that is worse and worse the more data we get.

Fortunately, we have the Principle of Sufficient Reason which can rule out hypotheses like N.

Tuesday, January 20, 2015

A natural hierarchy of levels of credence

Let a0=1/2. Let an=1−(1−an−1)2. Thus, approximately, the first couple of values are: a0=0.5, a1=0.75, a2=0.938, a3=0.996, a4=0.99998 and a5=0.9999999998.

Think of these as thresholds for credences. There is an interesting property that the items in this sequence have: If a perfect Bayesian rational agent assigns credence an to a proposition p, then she has credence at least an−1 that her credence in p will never fall below an−1. Thus, the agent who assigns credence a1=0.75 to p thinks it's at least as likely as not (a0=0.5) that her credence will always stay at the level of being at least as likely as not. Moreover, it is easy to show that no lower values for an have this property.

So each of these thresholds for belief has the property that it gives the previous degree of confidence in the belief not dipping below the previous degree of confidence. This is a pretty natural property, and it seems like it would provide a good way of dividing up our confidence in our beliefs. Credence from a1 on indicates, maybe, that "likely" the proposition is true. From a2 on, maybe we can be "fairly confident". From a3 on, maybe we can be "quite confident". The level a4 gives us "pretty sure", while a5 maybe makes us "sure", in the ordinary non-philosophical sense of "sure".

How we label the thresholds is a matter of words. But taking these thresholds as natural division points may be a good way to organize beliefs.

Note that each of these thresholds requires approximately twice as much evidence as the preceding one, if we measure the amount of evidence according to the logarithm of the Bayes factor.

Monday, February 24, 2014

"If there are so many, then probably there are more"

Suppose the police have found one person involved in the JFK assassination. Then simplicity grounds may give us significant reason to think that that one person is the sole killer. But suppose that they have found 15 people involved. Then while the hypothesis H15 that there were exactly 15 conspirators is simpler than the hypothesis Hn that there were exactly n for n>15, nonetheless barring special evidence that they got them all, we should suspect that there are more conspirators at large. With that large number, it's just not that likely that all were caught.

Why is this? I think it's because even though prior probabilities decrease with complexity, the increment of complexity from H15 to, say, H16 or H17 is much smaller than the increment of complexity from H1 to H2. Maybe P(H2)≈0.2P(H1). But surely we do not have P(H16)≈0.2P(H15). Rather, we have a modest decrease, maybe P(H16)≈0.9P(H15) and P(H17)≈0.9P(H16). If so, then P(H16)+P(H17)≈1.7P(H15). Unless we receive specific evidence that favors H15 over H16 and H17, something like this will be true of the posterior probabilities, and so the disjunction of H16 and H17 will be significantly more likely that H15.

Thus we have a heuristic. If our information is that there are at least n items of some kind, but we have no evidence that there are no more, then when n is small, say 1 or 2 or maybe 3, it may be reasonable to think there are no more items of that kind. But if n is bigger—my intuition is that the switch-around is around 6—then under these conditions it is reasonable to think there are more. If there are so many, then probably there are more. And this just follows from the fact that the increase in complexity from 1 to 2 is great, and from 2 to 3 is significant, but from 6 to 7 or maybe even 4 to 5 it's not very large.

This is all just intuitive, since I do not have any precise way to assign prior probabilities. But staying at this intuitive level, we get some nice intuitive applications:

  • If after thorough investigation we have found only one kind of good that could justify God's permitting evil, then we have significant evidence that it's the only such good. And if some evil is no justified by that kind of good, then that gives significant evidence that it's not justified. But suppose we've found six, say. And it's easy to find at least six: (1) exercise of virtues that deal with evils; (2) significant freedom; (3) preservation of laws of nature; (4) opportunities to go beyond justice via forgiveness[note 1]; (5) adding variety to life; (6) punishment; (7) the great goods of the Incarnation and sacrifice of the cross. So we have good reason to think there are more permission-of-evil justifying goods that we have not yet found. (Alston makes this point.)
  • Suppose our best definition of knowledge has three clauses. Then we might reasonably suspect that we've got the definition. But it is likely, given Gettier stuff, that one needs at least four clauses. But for any proposed definition with four clauses, we should be much more cautious to think we've got them all.
  • Suppose we think we have four fundamental kinds of truths, as Chalmers does (physics, qualia, indexicals and that's all). Then we shouldn't be confident that we've got them all. But once we realize that the list leaves out severel kinds (e.g., morality, mathematics, intentions and intentionality, pace Chalmers), our confidence that we have them all should be low.
  • If our best physics says that there are two fundamental laws, we have some reason to think we've got it all. But if it says that there six, we should be dubious.

Friday, February 10, 2012

Rawls and rationally intractable disagreement

Let me preface by saying I am not a political philosopher, and this may be off-base. Start by granting this claim for the sake of the argument:

  1. The disagreement between comprehensive views is very long-standing and there is no progress to agreement, except when non-rational, coercive methods are applied to generate agreement or for other merely sociological reasons there happens to be cultural homogeneity.
(I don't know how to characterize merely sociological reasons.) Now consider two possible explanations of (1):
  1. People's idiosyncratic or culturally-based preferences, as well as their presently-held comprehensive views, often significantly bias them in their disagreements between comprehensive views.
  2. It is not possible to resolve the disagreement between comprehensive views by reason alone.
We now have at least three options: only (2) explains (1); only (3) explains (1); both (2) and (3) explain (1). How would we decide between these? Well, first observe that (2) is a truism. Moreover, (2) clearly is at least a part of the explanation of (1). So of the three options, the two that remain are:
  • both (2) and (3) explain (1)
  • (2) by itself explains (1)
I understand that it is important to Rawls' project that (3) be a part of the explanation of (1), because it is important to Rawls' project that (3) be true, and apparently the main evidence he adduces for (3) is that it explains (1).

But now the question whether (1) is explains by (2) and (3), or simply by (2), is to a significant degree an empirical question.

And there is an obvious experiment to test between these options. Take a bunch of intelligent and rational people without idiosyncratic and culturally-based preferences who do not adhere to any comprehensive views, and see if they come to agree to on a comprehensive view or against all of them--if they do, then (3) is not a part of an explanation of (1), and if they don't, then (3) is a part of an explanation of (1). And we cannot at present rule out the possibility that such an experiment would rule in favor of the hypothesis that (2) by itself explains (1).

But now note that this experiment is precisely the original situation of deliberation under the veil of ignorance. And note that we can say directly this. If it is an empirically open possibility that agreement on a comprehensive view or against all of them would arise in the original situation, then it seems to be an open possibility that the delegates would legislate in accordance with a comprehensive view or in ways that significantly impugn the freedom to follow comprehensive views. And that's unacceptable to Rawls.

Sound-bite version: Please don't infer that a debate would be unsettled in an idealized situation from the fact that it's unsettled in the real world.

But I probably don't know what I'm talking about.

Monday, January 3, 2011

Conclusive evidence and confirmation

Fitelson proposes the following principle: "If E constitutes conclusive evidence for H1, but E constitutes less than conclusive evidence for H2 (where it is assumed that E, H1 and H2 are all contingent), then E favors H1 over H2."

The principle is false. You witness Jones being shot and after approaching you observe that he is dead. Let E be this evidence. Let H1 be the hypothesis that Jones is dead. Let H2 be the hypothesis that Jones has been killed. Then E is conclusive evidence for H1 and less than conclusive evidence for H2. (That you observed that Jones is dead entails H1. But on one reading of "Jones being shot", E is compatible with the hypothesis that Jones was already dead when he was shot—we can say "He was shot after he died." And on any reading, E is compatible with Jones having died of some other cause, coincidentally.) But it is wrong to say that E favors the hypothesis that Jones is dead over the hypothesis that Jones was killed.

Sunday, January 2, 2011

A derivation of the likelihood-ratio measure of confirmation

Let CK(E,H) be the degree of confirmation that E lends H given background K. Here is a derivation of the likelihood-ratio measure. The derivation is, I think, more compeling than Milne's.
We need several assumptions. For simplicity, write PK(A)=P(A|K) and PK(A|B)=P(A|BK). Our first assumption is uncontroversial and everybody accepts it. (For simplicity, I shall also suppose throughout that we're working with events such that none of the relevant Boolean combinations have probability zero or one.)
  1. CK(E,H) is a continuous function of the probabilities PK(−) where the blank can be filled in by any boolean combination of E and H.
Everybody in the measure of confirmation business accepts (1). We now need two more complex theses to get some interesting results. One is this:
  1. If I is independent of all Boolean combinations of E, H and K, then CK(IE,H)=CK(E,H).
An event I like that is obviously irrelevant, and so E and IE confirm H equally. This should be uncontroversial, though it is sufficient to refute the Eells-Jeffrey measure. The next thesis is more controversial:
  1. If E and F are events that are conditionally independent given HK as well as given (~H)K, then CKE(F,H)=CK(F,H).
This thesis says that independent evidence has the same evidential force no matter in which order it comes in. For instance, if you flip a coin twice to gather evidence for whether the coin is biased in favor of heads, and you get heads twice, each heads result provides exactly the same confirmation.
Now we get some substantive results. The first one is easy.
Theorem 1. If (1), then CK(E,H) is a function solely of PK(H), PK(E|H) and PK(E|~H), i.e., there is a function f such that CK(E,H)=f(PK(E|H),PK(E|~H),PK(H)).
This is because all the relevant Boolean combinations can be written in terms of these three probabilities.
I am omitting the proofs of the next couple of Theorems, except to note that obviously Theorem 4 follows from Theorems 2 and 3. I haven't written any of proofs out, but I am confident of theoremhood (maybe with some minor additional assumption). Of course, I could be wrong.
Theorem 2. If (1) and (2), then CK(E,H) is a function solely of PK(H) and the likelihood ratio PK(E|H)/PK(E|~H).
Observe that Theorem 2 refutes the Eells-Jeffrey likelihood-difference measure, given that (1) and (2) are so plausible.
Theorem 3. If (1) and (3), then CK(E,H) is a function solely of the likelihoods PK(E|H) and PK(E|~H).
Theorem 4. If (1), (2) and (3), then CK(E,H) is a function solely of the likelihood ratio PK(E|H)/PK(E|~H).
Therefore, given (1)-(3), CK(E,H)=f(PK(E|H)/PK(E|~H)) for some function f. It is obvious that f must then be an increasing function (this needs some additional assumptions). If all that is of interest is the comparison of degrees of confirmation, this is all we need. But perhaps we want to combine degrees of confirmation. This could be done additively or multiplicatively, i.e.,
  1. If E and F are conditionally independent given HK and (~H)K, then CK(EF,H)=CK(E,H)+CKE(F,H)
or
  1. If E and F are conditionally independent given HK and (~H)K, then CK(EF,H)=CK(E,H)CKE(F,H).
Now we get two final results.
Theorem 5. Assume (1), (2), (3) and (4). Then there is a constant c such that CK(E,H)=clog(PK(E|H)/PK(E|~H)).
Theorem 6. Assume (1), (2), (3) and (5). Then there is a constant c such that CK(E,H)=(PK(E|H)/PK(E|~H))c.
And for simplicity in both cases we should normalize by setting c=1.