Showing posts with label reflection principle. Show all posts
Showing posts with label reflection principle. Show all posts

Thursday, May 23, 2024

A supertasked Sleeping Beauty

One of the unattractive ingredients of the Sleeping Beauty problem is that Beauty gets memory wipes. One might think that normal probabilistic reasoning presupposes no loss of evidence, and weird things happen when evidence is lost. In particular, thirding in Sleeping Beauty is supposed to be a counterexample to Van Fraassen’s reflection principle, that if you know for sure you will have a rational credence of p, you should already have one. But that principle only applies to rational credences, and it has been claimed that forgetting makes one not be rational.

Anyway, it occurred to me that a causal infinitist can manufacture something like a version of Sleeping Beauty with no loss of evidence.

Suppose that:

  • On heads, Beauty is woken up at 8 + 1/n hours for n = 2, 4, 6, ... (i.e., at 8.5 hours or 8:30, at 8.25 hours or 8:15, at 8.66… hours or 8:10, and so on).

  • On tails, Beauty is woken up at 8 + 1/n hours for n = 1, 2, 3, ... (i.e. at 9:00, 8:30, 8:20, 8:15, 8:10, …).

Each time Beauty is woken up, she remembers infinitely many wakeups. There is no forgetting. Intuitively she has twice as many wakeups on tails, which would suggest that the probability of heads is 1/3. If so, we have a counterexample to the reflection principle with no loss of memory.

Alas, though, the “twice as many” intuition is fishy, given that both infinities have the same cardinality. So we’ve traded the forgetting problem for an infinity problem.

Still, there may be a way of avoiding the infinity problem. Suppose a second independent fair coin is tossed. We then proceed as follows:

  • On heads+heads, Beauty is woken up at 8 + 1/n hours for n = 2, 4, 6, ...

  • On heads+tails, Beauty is woken up at 8 + 1/n hours for n = 1, 3, 5, ...

  • On tails+whatever, Beauty is woken up at 8 + 1/n hours for n = 1, 2, 3, ....

Then when Beauty wakes up, she can engage in standard Bayesian reasoning. She can stipulatively rigidly define t1 to be the current time. Then the probability of her waking up at t1 if the first coin is heads is 1/2, and the probability of her waking up at t1 if the first coin is tails is 1. And so by Bayes, it seems her credence in heads should be 1/3.

There is now neither forgetting nor fishy infinity stuff.

That said, one can specify that the reflection principle only applies if one can be sure ahead of time that one will at a specific time have a specific rational credence. I think one can do some further modifying of the above cases to handle that (e.g., one can maybe use time-dilation to set up a case where in one reference frame the wakeups for heads+heads are at different times from the wakeups for heads+tails, but in another frame they are the same).

All that said, the above stories all involve a supertask, so they require causal infinitism, which I reject.

Thursday, May 4, 2023

Reflection and null probability

Suppose a number Z is chosen uniformly randomly in (0, 1] (i.e., 0 is not allowed but 1 is) and an independent fair coin is flipped. Then the number X is defined as follows. If the coin is heads, then X = Z; otherwise, X = 2Z.

  • At t0, you have no information about what Z, X and the coin toss result are, but you know the above setup.

  • At t2, you learn the exact value of X.

Here’s the puzzling thing. At t2, when you are informed that X = x (for some specific value of x) your total evidence since t0 is:

  • Ex: Either the coin landed heads and Z = x, or the coin landed tails and Z = x/2.

Now, if x > 1, then when you learn Ex, you know for sure that the coin was tails.

On the other hand, if x ≤ 1, then Ex gives you no information about whether the coin landed heads or tails. For Z is chosen uniformly and independently of the coin toss, and so as long as both x and x/2 are within the range of possibilities for Z, learning Ex seems to tell you nothing about the coin toss. For instance, if you learn:

  • E1/4: Either the coin landed heads and Z = 1/4, or the coin landed tails and Z = 1/8,

that seems to give you no information about whether the coin landed heads or tails.

Now add one more stage:

  • At t1, you are informed whether x ≤ 1 or x > 1.

Suppose that at t1 what you learn is that x ≤ 1. That is clearly evidence for the heads hypothesis (since x > 1 would conclusively prove the tails hypothesis). In fact, standard Bayesian reasoning implies you will assign probability 2/3 to heads and 1/3 to tails at this point.

But now we have a puzzle. For at t1, you assign credence 2/3 to heads, but the above reasoning shows you that at t2, you will assign credence 1/2 to heads. For at t2 your total eidence since t0 will be summed up by Ex for some specific x ≤ 1 (Ex already includes the information given to you at t1). And we saw that if x ≤ 1, then Ex conveys no evidence about whether the coin was heads or tails, so your credence in heads at t2 must be the same as at t1.

So at t1 you assign 2/3 to heads, but you know that when you receive further more specific evidence, you will move to assign 1/2 to heads. This is counterintuitive, violates van Fraassen’s reflection principle, and lays you open to a Dutch Book.

What went wrong? I don’t really know! This has been really puzzling me. I have four solutions, but none makes me very happy.

The first is to insist that Ex has zero probability and hence we simply cannot probabilistically update on it. (At most we can take P(H|Ex) to be an almost-everywhere defined function of x, but that does not provide a meaningful result for any particular value of x.)

The second is to say that true uniformity of distribution is impossible. One can have the kind of uniformity that measure theorists talk about (basically, translation invariance), but that’s not enough to yield non-trivial comparisons of the probabilities of individual values of Z (we assumed that x and x/2 were equally likely options for Z if X ≤ 1).

The third is some sort of finitist thesis that rules out probabilistic scenarios with infinitely many possible outcomes, like the choice of Z.

The fourth is to bite the bullet, deny the reflection principle, and accept the Dutch Book.

Wednesday, April 26, 2023

Cable Guy and van Fraassen's Reflection Principle

Van Fraassen’s Reflection Principle (RP) says that if you are sure you will have a specific credence at a specific future time, you should have that credence now. To avoid easy counterexamples, the RP needs some qualifications such that there is no loss of memory, no irrationality, no suspicion of either, full knowledge of one’s own credences at any time, etc.

Suppose:

  1. Time can be continuous and causal finitism is false.

  2. There are non-zero infinitesimal probabilities.

Then we have an interesting argument against van Fraassen’s Reflection Principle. Start by letting RP+ be the strengthened version of RP which says that, with the same qualifications as needed for RP, if you are sure you will have at least credence r at a specific future time, then you should have at least credence r now. I claim:

  1. If RP is true, so is RP+.

This is pretty intuitive. I think one can actually give a decent argument for (3) beyond its intuitiveness, and I’ll do that in the appendix to the post.

Now, let’s use Cable Guy to give a counterexample to RP+ assuming (1) and (2). Recall that in the Cable Guy (CG) paradox, you know that CG will show at one exact time uniformly randomly distributed between 8:00 and 16:00, with 8:00 excluded and 16:00 included. You want to know if CG is coming in the afternoon, which is stipulated to be between 12:00 (exclusive) and 16:00 (inclusive). You know there will come a time, say one shortly after 8:00, when CG hasn’t yet shown up. At that time, you will have evidence that CG is coming in the afternoon—the fact that they haven’t shown up between 8:00 and, say, 8:00+δ for some δ > 0 increases the probability that CG is coming in the afternoon. So even before 8:00, you know that there will come a time when your credence in the afternoon hypothesis will be higher than it is now, assuming you’re going to be rational and observing continuously (this uses (1)). But clearly before 8:00 your credence should be 1/2.

This is not yet a counterexample to RP+ for two reasons. First, there isn’t a specific time such that you know ahead of time for sure your credence will be higher than 1/2, and, second, there isn’t a specific credence bigger than 1/2 that you know for sure you will have. We now need to do some tricksy stuff to overcome these two barriers to a counterexample to RP+.

The specific time barrier is actually pretty easy. Suppose that a continuous (i.e., not based on frames, but truly continuously recording—this may require other laws of physics than we have) video tape is being made of your front door. You aren’t yourself observing your front door. You are out of the country, and will return around 17:00. At that point, you will have no new information on whether CG showed up in the afternoon or before the afternoon. An associate will then play the tape back to you. The associate will begin playing the tape back strictly between 17:59:59 and 18:00:00, with the start of the playback so chosen that that exactly at 18:00:00, CG won’t have shown up in the playback. However, you don’t get to see the clock after your return, so you can’t get any information from noticing the exact time at which playback starts. Thus, exactly at 18:00:00 you won’t know that it is exactly 18:00:00. However, exactly at 18:00:00, your credence that CG came in the afternoon will be bigger than 1/2, because you will know that the tape has already been playing for a certain period of time and CG hasn’t shown up yet on the tape. Thus, you know ahead of time that exactly at 18:00:00 your credence in the afternoon hypothesis will be higher than 1/2.

But you don’t know how much higher it will be. Overcoming that requires a second trick. Suppose that your associate is guaranteed to start the tape playback a non-infinitesimal amount of time before 18:00:00. Then at 18:00:00 your credence in the afternoon hypothesis will be more than 1/2 + α for any infinitesimal α. By RP+, before the tape playback, your credence in the afternoon hypothesis should be at least 1/2 + α for every infinitesimal α. But this is absurd: it should be exactly 1/2.

So, we now have a full counterexample to RP+, assuming infinitesimal probabilities and the coherence of the CG setup (i.e., something like (1)). At exactly 18:00:00, with no irrationality, memory loss or the like involved (ignorance of what time it is not irrational nor a type of memory loss), you will have a credence at least 1/2 + α for some positive infinitesimal α, but right now your credence should be exactly 1/2.

Appendix: Here’s an argument that if RP is true, so is RP+. For simplicity, I will work with real-valued probabilities. Suppose all the qualifications of RP hold, and you now are sure that at t1 your credence in p will be at least r. Let X be a real number uniformly randomly chosen between 0 and 1 independently of p and any evidence you will acquire by t1. Let Ct(q) be your credence in q at t. Let u be the following proposition: X < r/Ct(p) and p is true. Then at t1, your credence in u will be (r/Ct(p))Ct(p) = r (where we use the fact that r ≤ Ct(p)). Hence, by RP your credence now in u should be r. But since u is a conjunction of two propositions, one of them being p, your credence now in p should be at least r.

(One may rightly worry about difficulties in dropping the restriction that we are working with real-valued probabilities.)

Friday, January 16, 2015

A tale of two forensic scientists

Curley and Arrow are forensic scientist expert witnesses involved in very similar court cases. Each receives a sum of money from a lawyer for one side in the case. Arrow will spend the money doing repeated tests until the money (which of course also pays for his time) runs out, and then he will present to the court the full test data that she found. Curley, on the other hand, will do tests until such time as either the money runs out or he has reached an evidential level sufficient for the court to come to the decision that the lawyer who hired him wants. Of course, Curley isn't planning to commit perjury: he will truthfully report all the tests he actually did, and hopes that the court won't ask him why he stopped when he did. Curley reasons that his method of proceeding has two advantages over Arrow's:

  1. if he stops the experiments early, his profits are higher and he has more time for waterskiing; and
  2. he is more likely to confirm the hypothesis that his lawyer wants confirmed, and hence he is more likely to get repeat business from this lawyer.

Now here is a surprising fact which is a consequence of the martingale property of Bayesian investigations (or of a generalization of van Fraassen's Reflection Principle). When Curley and Arrow reflect on what their final credences will be with respect to what they are each testing, if they are perfect Bayesian agents, their expectation for their future credence equals their current credence. This thought may lead Curley to think himself quite innocent in his procedures. After all, on average, he expects to end up with the same credence as he would if he followed Arrow's more onerous procedures.

So why do we think Curley crooked? It's because we do not just care about the expected values of credences. We care about whether credences reach particular thresholds. In the case at hand, we care about whether the credence reaches the threshold that correlates with a particular court decision. And Curley's method does increase the probability that that credence level will be reached.

What happens is that Curley, while favoring a particular conclusion, sacrifices the possibility of reaching evidence that confirms that conclusion to a degree significantly higher than his desired threshold, for the sake of increasing the probability of reaching the threshold. For when he stops his experiments once the level of confirmation has reached the desired threshold, he is giving up on the possibility—useless to him or to the side that hired him—that the level of confirmation will go up even higher.

I think it helps that in real life we don't know what the thresholds are. Real-life experts don't know just how much evidence is needed, and so there is some an incentive to try to get a higher level of confirmation, rather than to stop once one has reached a threshold. But of course in the above I stipulated there was a clear and known threshold.

Van Fraassen's reflection principle

Van Fraassen's reflection principle (RP) says that if a rational agent is certain she will assign credence r to p, then she should now assign r to p.

As I was writing on being-sure yesterday, I was struck by the fact (and I wasn't the first person to be struck by it) that for Bayesian agents, the RP is a special case of the fact that the sequence of continually updated credences forms a martingale whose filtration is defined by the information one is updating on the basis of.

Indeed, martingale considerations give us the following generalization of RP:

  • (ERP) For any future time t, assuming I am certain that I will remain rational, my current credence in p should equal the expected value of my credence at t.
(Van Fraassen himself formulates ERP in addition to RP.) In RP, it is assumed that I am certain that my credence at t will be r, and of course then the expected value of that future credence is r. But ERP generalizes this to cases where I don't know exactly what my future credence will be.

But we can get an even further generalization of RP. I understand that ERP and RP apply when there is a specific future time at which one knows what one's credence will be. But suppose instead we have some method of determining a variable future time T. The one restriction on that determination is that it can only depend on the data available to us up to and including that time. For instance, we might not know exactly when we will perform some set of experiments in the next couple of years, and we might let T be a time at which those experiments have been performed. The generalization of ERP then is:

  • (GERP) For any variable future time T in a future human life bounded by the normal bounds on human life and such that whether T has been reached is guaranteed to be dependent only on data gathered up to time T, my current credence in p should equal the expected value of my credence at T.
This follows from Doob's optional sampling theorem (given that human life has a normal upper bound of about 200 years) and the martingale property of Bayesian epistemic lives.

Now GERP seems like a quite innocent generalization of ERP when we are merely thinking about the fact that we don't know when we will do an experiment. But now imagine a slightly crooked scientist out to prove a pet theory. She gets a research grant that suffices for a sequence of a dozen experiments. She is not so crooked that she will fake experimental data or believe contrary to the evidence, but she resolves that she will stop experimenting as soon as she has enough experiments to confirm her theory—or at the end of the dozen experiments, if worst comes to worst. This is intuitively a failure of scientific integrity—she seems to be biasing one's research plan to favor the pet theory. One might think that the slightly crooked scientist would be irrational to set her current credence according to her expected value of her credence at her chosen stopping point. But according to GERP, that's exactly what she should do. Indeed, according to GERP, the expected value of her credence at the end of a series of experiments does not depend on how she chooses when to stop the experiments. Nonetheless, she is being crooked, as I hope to explain in a future post.