Showing posts with label probability. Show all posts
Showing posts with label probability. Show all posts

Tuesday, September 15, 2026

Molinism and correlations

Assume Molinism.

Suppose T is a set of character traits that makes one be equipoised between freely taking and refusing a $5000 bribe for a city contract. Let C1, C2, ... be distinct and highly specific possible circumstances of Curley being offered such a bribe, differing in morally unimportant ways that are not very relevant to Curley’s character. Perhaps he is offered the bribe in an Italian restaurant, or in a Thai restaurant, or while walking in the park. Perhaps the bribe is being offered by a man or by a woman. Maybe it’s in the morning or the evening. Etc.

Consider the counterfactual of free will:

  • Qi: Were Curley to have T and be offered a $5000 bribe for a city contract in Ci, he would accept the bribe.

Intuitively, P(Qi) ≈ 1/2 for any i. Let’s accept that intuition. But now here is an interesting question: What kinds of statistical correlation is there between Q1, Q2, ...?

The intuition behind van Inwagen’s re-run thought experiment says that if Curley were offered a sequence of exactly similar bribes, with memory erased in between, then Curley’s acceptance/rejection decisions would behave like an independent sequence of random variables. That intuition suggests that Q1, Q2, ... are also statistically independent.

On the other hand, one might have the conflicting intuition that Qi and Qj are more correlated when the circumstances Ci and Cj are more similar.

On the third hand, one might think that although each of Qi has probability around 1/2, the probabilities of conjunctions of the Qi are undefined, and hence it makes no sense to talk about their correlations.

These seem to me to be the three most plausible views on the correlation question. Call these the Independence, Similarity, and Undefined views.

On Molinism plus Independence, God has a vast amount of providential power. He is nearly certain to be able to get Curley (or any other free agent) to freely do anything he wants, simply by choosing minor and unimportant features of circumstances. Given independence, it is extremely unlikely that all the Qi have the same truth value. Thus, God can just choose the circumstances to get the truth value he wants—to get Curley to accept or to get Curley to reject the bribe. In particular, on Independence, free will does not do much to help the theist.

On the Similarity view, there are some more serious constraints of God’s providential power. However, I think there is another problem with the Similarity view. On Similarity, there is presumably some complicated function from the degree of similarity between circumstances to the degree of correlation (say, measured by covariance) between the truth values of the corresponding counterfactuals. Where does this function come from? Normally correlations between events are explained by laws of nature. But counterfactuals of free will are prior to laws of nature. I suppose the correlation function would have to be something necessary. But it does seem mysterious.

I think that if I were a Molinist, I would find the Undefined view somewhat appealing. The version of it that seems most plausible would be that P(QiQj) is not defined as a number, but can be represented as the interval [max(0,P(Qi)+P(Qj)−1),min(P(Qi),P(Qj))]. But I still find it a bit odd that P(Qi) and P(Qj) are defined but P(QiQj) is not. Nonetheless, this seems possible.

Monday, August 31, 2026

Subjective time and self-locating probability

Suppose I am sluggish mentally in the morning and function fast mentally in the afternoons. I find myself in a dark room with no information about what time it is. Consider the two hypotheses MORN, that it is between 7am and 8am and AFT, that it is between 1pm and 2pm. Which is more likely, or are they equally likely?

It seems obvious that if I function mentally the same way in the morning and afternoon, MORN and AFT should be equally likely.

But now suppose I live in a world where every morning all processes in the solar system slow down by a factor of two, returning to the normal speed in the afternoon. Then, surely, for all practical purposes, the period from 7am to 8am is half an hour long! And so it seems right to say P(MORN) = (1/2)P(AFT). But the case of my own mental sluggishness seems to be just a more localized version of the solar system slowing down. Thus, in my original story we should say that likewise P(MORN) < P(AFT).

Here is one way to imagine mental functioning slowing down: my thinking is divided into discrete moments (like a computer’s internal clock, where basic operations take a clock cycle), and in the morning there are fewer of these discrete moments of thinking. If so, then it seems reasonable to say that the probability that the current time is between x and y is proportional to the number of discrete moments of thought between x and y, and hence P(MORN) < P(AFT), there being fewer moments of thought in the morning.

But I have a hard time getting good intuitions about the continuous case.

Preference and choice disposition

Here is an initially plausible thesis about deliberation that I've long felt is right:

  1. Normally, when choosing between options, you are more likely to choose an option X than an option Y if and only if you prefer X to Y.

Now consider a normal case where you must choose between four options:

  1. Cookies and Earl Gray
  2. Cookies and chamomile
  3. Crackers and Earl Gray
  4. Crackers and chamomile.

Suppose that you mildly prefer cookies to crackers and Earl Gray to chamomile, so that you have a 60% chance of choosing cookies over crackers and 60% chance of choosing Earl Gray over chamomile. Further, suppose you have no preferences about the combination as such. Then the chances of your four options are:

  • P(A) = 36%

  • P(B) = 24%

  • P(C) = 24%

  • P(D) = 16%.

Clearly, you can choose A. But notice an odd thing. You can organize your deliberation to start as follows: First choose between A and the disjunction B or C or D. (Maybe then if you chose B or C or D, you decide between B and C or D. And finally if you chose C or D, you decide between C and D.) However, here is another plausible principle of rationality:

  1. Normally, if you prefer A to every disjunct of a disjunction, you prefer A to the disjunction.

Organizing your deliberation as above is not abnormal. So, (2) should still apply. You prefer A to B, and you prefer A to C, and you prefer A to D. So you should prefer A to the disjunction B or C or D, and hence you prefer A over the disjunction B or C or D. Thus, by (1) you are more likely to choose A than the disjunction B or C or D. In particular, the chance of your choosing A isn’t 36%, but must be more than 50%.

Hence, (1) predicts that the probabilities of our choices greatly depend on how one coarse-grains the options in one’s deliberation. This shouldn’t be. I thus reject (1). Which makes me sad, as it was a very neat thesis.

Tuesday, August 25, 2026

More on infinite chances at salvation

Yesterday, I offered a Toy Model of how God could give infinite chances to someone and they could still reject him. The crucial assumption of the model was that it might take multiple good choices to dig oneself out of the moral hole that one’s wickedness got one into.

But one can also have a model without that assumption. Suppose that every day I make a choice for or against God. If I choose for God, I am saved and God makes me virtuous. If I choose against God, I become habituated into a bit more vice. How much more? Well, we can measure a person’s virtue by how likely they are to choose for God. So let’s suppose that there is some number β strictly between 0 and 1 such that each time I choose against God, my goodness—as measured by the probability of my next choice being for God—is multiplied by β. We can think of β as a kind of wicked habituation constant.

What is the probability that I will always reject God? Suppose my initial level of goodness is α. Then the probability that I will reject God on day 1 is 1 − α. If I do that, then on day 2 my level of goodness will be αβ, and the probability that I will reject God will be 1 − αβ. And if I do that, then on day 3 my level of goodness will be αβ2, and the probability that I will reject God will be 1 − αβ2. So the probability that I will always reject God is:

  • P(R∞) = (1−α)(1−βα)(1−β2α)....

Easy mathematical fact: as long as α < 1 (i.e., I am not maximally good on day 1) and β < 1 (the bad choice does decrease my goodness), P(R∞) is bigger than zero. In other words, on this model, I have infinite chances, and each choice has the possibility of resulting in my complete redemption, and yet I have a non-zero probability of choosing badly for eternity.

Let’s make up some conservative numbers. Suppose the infinite chances regimen starts with death, and that β = 0.99: one gets 1% worse each time one chooses against God—and that’s 1% of my previous level of goodness, rather than an absolute 1%. Presumably people die with various levels of goodness between 0 and 1. Suppose one in a thousand people dies with a goodness level α ≤ 0.1. Now if one has goodness level α = 0.1 when one starts the infinite chances regimen at death, and β = 0.99, it turns out (or so Mathematica calculates QPochhammer[.1,.99]) that the probability of damnation is about 0.000035. So with 8 billion people, there would be 8 million dying with goodness level at most 0.1, and so we might well expect at least 8, 000, 000 × 0.000035 = 280 of them to be damned even if they got the infinite chances regimen. So on this toy model, we would expect some people now alive to be damned.

Now it seems plausible that it is compatible with God’s goodness that he would allow one in a thousand people to die with a goodness level of 0.1, and that he would allow a character deterioration of at least 1% per posthumous choice against God—this would give people a real ability to affect their character, while yet being very generous with redemptiveness. In other words, even on an infinite chances model, it is compatible with God’s goodness that hundreds of people be forever separated from God. But if hundreds of people damned is compatible with God’s goodness on an infinite chances model, why wouldn’t hundreds of people damned also be compatible with God’s goodness apart from an infinite chances model, as long as it was also a model where God is very generous in redemptiveness and gives people a real ability to affect their character?

Wednesday, June 17, 2026

Self-locating evidence and bearers of epistemic good

In the case of non-epistemic goods, it’s an obvious feature of life that someone there is a choice to be made by an individual between their own first-order good and the first-order good of the community—each requires the sacrifice of the other. In the case of epistemic goods, this is less obvious.

In the pragmatic case, the typical reason for such competition between goods is due to limited resources. This, of course, also happens in the epistemic sphere. Suppose Alice is much more intellectually talented than Bob, but only Bob has the money to go to university. If Bob spends the money on himself, he will gain private epistemic goods, but will contribute little epistemically to society as a whole. But if he gives the money to Alice, she may become a brilliant scholar or scientist, significantly contributing to society’s knowledge.

More interesting than these, however, are cases of competition between private and communal epistemic goods that are not due to epistemic resources. I find it interesting that some cases of self-locating evidence appear to be such.

Suppose there are ten billion people in the world, currently isolated from one another. A device produced by a mad scientist has a 99.9% chance at noon today of triggering a death ray that randomly kills 99.9999% of the world population. Noon has just passed. You are still alive. Should you think the device worked? Sleeping Beauty style arguments say “No”. This time I want to think about this in terms of individual epistemic goods. In N runs of the device, 0.001N runs will have you survive because the device doesn’t trigger and 0.999 ⋅ 0.000001N runs will have you survive despite the device triggering. Thus, the vast majority of the runs where you survive are runs where the device didn’t trigger. Hence, it’s best for you individually to adopt the epistemic policy of thinking the device didn’t trigger.

But on the other hand, suppose we all adopt the epistemic policy of thinking the device didn’t trigger. Then 99.9% of the time, we are unanimously collectively wrong. And if we all adopt the epistemic policy of thinking the device did trigger, then 99.9% of the time, we are unanimously collectively right. It seems thus that if we look at the epistemic goods of society, then a policy of thinking the device did trigger is best.

If this is right, it points to a potential diagnosis of why the problems about self-locating evidence (doomsday, multiverses, Sleeping Beauty, etc.) are so difficult. For there may be different bearers of epistemic goods at play—say, society vs. the individual—and it could be that different answers are appropriate depending on whose goods we are pursuing. Maybe.

Friday, June 5, 2026

Knowledge and induction

Assume that in fact all ravens are black. Suppose you are sequentiallly observing ravens, and noting each one to be black. After observing n ravens, your evidence that the next raven is black will typically be significantly better than the evidence that all ravens are black. Now at some point, say after observing nA ravens, your evidence that all ravens are black will rise to the level of knowledge. Thus, plausibly, at an earlier point in the sequence, call it nN, your evidence that the next raven is black will have risen to the level of knowledge.

Suppose now you have observed nA − 1 ravens, and you have been handed a raven in an opaque box, which you are certain you are about to open. Since nN < nA, at this point you have reached nN. Hence:

  1. You do not know that all ravens are black.

  2. You do know that the next raven is black.

  3. You know that when you observe the next raven, you will have sufficient evidence for knowledge that all ravens are black.

But note that while you know you will have sufficient evidence for knowledge that all ravens are black, you don’t know that you will know that all ravens are black. There is nothing deeply surprising about this distinction. We might well say about someone who has been subjected to misleading or Gettiered evidence that they have sufficient evidence to know something but nonetheless they don’t know, though the case at hand feels different.

One interesting thing about this case, as I read it, is that it contradicts the thesis that K = E, i.e., that knowledge is evidence. For if knowledge is evidence, and you know that the next raven is black, then you already have the evidence you will gain by observing the next raven, and hence you are already in the position to know that all ravens are black.

Another interesting thing is that it shows that you can know something and nonetheless it be rational for you to investigate it. For you know that the next raven is black, but it’s worth investigating further, since it is only upon observation that your knowledge of the next raven’s blackness turns into the kind of evidence that gives you knowledge that all ravens are black.

All this might make one think that I have misconstrued the epistemic facts, and it is false that there can be a point nN prior to nA at which you know that the next raven is black. Here is one way to back up my intuition that there can be such a point nN < nA. Suppose that we know for sure we live in a world where the color distributions of birds are always uncorrelated between the males and the females of the species, so that information about the color of members of one sex are irrelevant to the color of members of the other sex. Also assume that you know for sure that ravens have equal numbers of each sex, that you are observing ravens in an alternative female-male-female-male-… sequence, and that your priors for the color distributions of the two sexes of ravens are the same. Then if pM and pF are the probabilities that all male ravens are black and all female ravens are black, and pA is the probability that all ravens are black, then at any given point in the observation sequence pA = pMpF. Let nM and nF be the points in the sequence where you know that all male and all female ravens are black, respectively. Then, nM < nA and nF < nA, since pA = pMpF is significantly smaller than either pM and pF at all points in the sequence except when we’ve observed all the ravens of one sex, and since pM and pF rise fairly gradually as we go through the sequence. Thus, at the point nA − 1, we will have already reached knowledge that all the male ravens are black and the knowledge that all the female ravens are black. In particular, then, we know that the next raven is black, since if nA is even, the nAth raven is male and we know all male ravens are black, and if nA is odd, then the nAth raven is female, and we know that all female ravens are black.

Friday, April 24, 2026

More on wagers for the perfectly rational

Consider a choice between two wagers on a fair coin:

  • W1: on heads, you get $1 if you are perfectly rational and $3 if you are not

  • W2: on tails, you get $2 if you are perfectly rational and $1 if you are not.

Suppose you are perfectly rational, and that it’s a part of perfect rationality that you know for sure you’re perfectly rational. It’s obvious you should go for W2. But let’s calculate. We immediately run into the zero-probability problem that I’ve lately been thinking about. For if you’re perfectly rational, the probability that you go for W1 is zero, so E(U|W1) seems to be undefined. Of course, E(U|W2) is unproblematically half of $2, or $1, but you can’t say whether that beats “undefined” or not.

Suppose you think: Maybe E(U|W1) is undefined in classical probability, but maybe I can use some other way of defining it, say using Popper functions.

Well, let’s think about what E(U|W1) “should be”. So imagine that you actually go for W1. Now, only an imperfectly rational agent would go for W1. So, if you were to go for W1, you would get $3 on heads, so your expected payoff would be $1.50, which beats anybody’s expected payoff for W2. So, formally, E(U|W1) is undefined, but if you close your eyes to that and think intuitively, you get E(U|W1) equally $1.50, which yields the wrong result that as a perfectly rational agent you should go for W1.

What if we say that a perfectly rational agent need not know for sure that they are perfectly rational? Suppose, say, you are perfectly rational agent who is 0.99 sure you are perfectly rational. Then E(U|W1) and E(U|W2) are both well-defined. But what are they? Well, it’s intuitively clear that if you are 0.99 sure that you are perfectly rational, you should go for W2. But supposing that’s right, then W1 entails you are not perfectly rational, and since P(W1) = 0.01, the expectation E(U|W1) is well-defined, and must be equal to $1.50. Oops!

This line of reasoning assumed evidential decision theory. What if you go for causal decision theory? Well, there are two causal hypotheses: R (you are perfectly rational) and Rc (you are not) with P(R) = 0.99 and P(Rc) = 0.01. So now your causal expected utility on W1 equals

  • CE(U|W1) = 0.99E(U|W1∩R) + 0.01E(U|W2∩Rc).

What is this? Well, W1 ∩ R is the empty set! But conditionalizing on an empty set is not a merely technical problem in the way that conditionalizing on a specific zero-probability outcome of a continuous spinner is. Rather, it is simply nonsense. So the first summand is undefined, and hence the sum is undefined. Thus you simply cannot make a decision with causal decision theory here.

It’s obvious that if you’re nearly sure you’re perfectly rational you should go for W2. But neither evidential nor causal decision theory gives a way to that conclusion.

[By the way, the reason I set up W1 and W2 as I did, with one having the payoff on heads and the other on tails, was to ensure that we didn’t have domination. For one might reasonably say that a perfectly rational agent will try to decide on grounds of domination first, before resorting to probabilities.]

Thursday, April 23, 2026

Good's Theorem, perfect rationality, and conditioning on zero probability events

Recently, I found myself puzzled by the difficulty in applying “classical” evidential decision theory to a perfectly rational agent. The problem was that the rational agent decides whether to do A or B based on a comparison between the conditional expectations E(U|A) and E(U|B) of the utility function U. But supposing that in fact E(U|A) > E(U|B), the perfectly rational agent has no chance of doing B, so P(B) = 0, and hence E(U|B) is undefined.

But then I thought this isn’t a big deal, because we aren’t perfectly rational agents, so we always have a chance of screwing up and hence P(B) > 0 even if E(U|B) is much less than E(U|A).

I am not entirely satisfied with this. After all, you might think: “I may be pretty imperfect, but if I am choosing between a donut D and a year of torture T, I have zero chance of choosing the year of torture. But then E(U|T) is undefined, so how am I being rational in this choice? Maybe that’s a good objection, maybe not.

But here is another reason why the “We’re imperfect” solution isn’t completely ideal. We want to say that Good’s Theorem tells us something important about rationality—namely, that more information makes rational agents make better decisions. Good’s Theorem is usually interpreted as saying that under some independence conditions, the expected value of a perfectly rational choice given more information is no less than that of a perfectly rational choice given less information. Notice that this is obviously false in the case of an imperfectly rational agent. Thus, we have to make sense of “What a perfectly rational agent would choose” to make sense of the standard interpretation of Good’s Theorem. Moreover, in the setting of Good’s Theorem, the perfectly rational agent has to be choosing based on expected utilities—and that’s precisely what generates the zero-probability-conditioning problem.

Now, the Theorem is still true as an abstract bit of mathematics. But the application is difficult if we can’t make sense of a perfectly rational agent who is certain to maximize expected utility.

Likely we can extend Good’s Theorem to talk about the limiting case of imperfect agents getting more and more perfect. But it would be nice if we didn’t have to.

Monday, April 13, 2026

A double lottery and non-normalized probabilities

Suppose a positive integer N is generated by a fair lottery.

Then, a random integer K is chosen between 1 and N (inclusive).

What information does this give you about N?

Obviously you now know that N ≥ K. Anything else?

Consider some specific pair of numbers n ≥ k, and suppose we’ve found out that K = k. What’s the probability that N = n? Of course P(N=n|K=k) = 0/0. But what if we do this as a limiting procedure. Suppose first that N is randomly chosen between 1 and M where M ≥ n, and let PM be the probabilities for this case. Then

  • PM(N=n|K=k) = (1/M)(1/n)/[(1/M)Σj=kMj−1] = (1/n)/Σj=kMj−1.

Take the limit as M goes to infinity. Since Σn=k∞j−1 = ∞, the limit is zero, so we don’t have a meaningful distribution for N.

On the other hand, what if we independently choose two random integers K1 and K2 between 1 and N? Suppose n ≥ ki for i = 1, 2. Let k* = max (k1,k2). Then:

  • PM(N=n|K1=k1,K2=k2) = (1/M)(1/n2)/[(1/M)Σj=k*Mj−2] = (1/n2)/Σj=k*Mj−2.

Take the limit as M → ∞ and call that P(N=n|K1=k1,K2=k2). The limit behaves like ck*/n2, for a constant c > 0, and generates a well-defined probability for N = n.

With zero samples, we don’t have a well-defined probability for N. With one sample, we still don’t. But with two samples (or more), now we do. This is a rummy thing: how is it that sampling turns probabilistic nonsense into sense?

This is making me more friendly to using non-normalized probabilities. After all, the fair lottery for N is easily modeled by the constant probability p0(n) = 1. With one sample N = k, we have p1(n) = 1/n for n ≥ k and p1(n) = 0 for n < k. With two samples k1, k2, we have p2(n) = 1/n2 for n ≥ max (k1,k2) and p2(n) otherwise. All this makes perfect sense. And there is a lovely mathematical feature of non-normalized probabilities: conditionalization is conjunction. The conditional probability of an event A on event B is just the probability of A ∩ B.

Non-normalized probabilities aren’t going to solve all problems with infinite fair lotteries. For instance, I toss a fair coin and generate a number N with the following rule. On heads, I choose N with my fair lottery on the positive integers. On tails, I choose N such that the probability of N = n is 2−n (e.g., I toss an independent fair coin and let N be the number of the first toss that gives heads). What’s my non-normalized probability p(x,n), where x is heads or tails and n is a positive integer? We surely want ∑np(H,n) = ∑np(T,n): the total probability of the heads options equals the total probability of the tails options. But clearly p(T,n) has to exponentially decrease so ∑np(T,n) is finite and non-zero. On the other hand, p(H,n) is constant, so ∑np(H,n) is zero or infinity. So they can’t be equal.

But I wonder if one could say something like this: Non-normalized probabilities make sense in certain cases, and in those cases it’s reasonable to use them?

Thursday, April 9, 2026

Epistemic utilities and death

In the previous post, I proved that we get a proper scoring rule if we compute epistemic utilities as follows. We start with our current credence assignment, consider what credence assignment we will have in the future after we update on some further evidence, and then score that. I then suggested that one could get a lifetime epistemic utility by adding up the epistemic utilities over all the moments of life, and as long as death wasn’t random—as long as the lifespan was fixed—this would generate a proper scoring rule. I then said that if death is random (as it is) it might be the case that you don’t get a proper scoring rule.

My conjecture was wrong. You still get a proper scoring rule despite random death. It’s easy to see this. The basic idea is this. Suppose that you might die the next moment. This partitions the probability space into two subsets, D and L, for death and life. Your current credence is p. Next moment, on D, your credence either doesn’t exist (because you don’t exist or you exist in some supernatural state where you don’t have credences) or doesn’t count (because I am after lifetime credences). Thus, the appropriate way to do a forward-looking scoring of your credence p is to score it s(pL) on L and 0 on D, where pL is the result of conditionalizing your credence on the evidence L (after all, if you are alive, you will conditionalize on being alive), and s is some proper scoring rule. In other words, your forward looking score is sL(p) = 1L ⋅ s(pL).

Is this score proper? Yes! For by propriety of s we have:

  • EpL(s(pL)) ≥ EpL(s(qL)).

But this is the same as:

  • (p(L))−1Ep(1L⋅s(pL)) ≥ (p(L))−1Ep(1L⋅s(qL)).

Multiplying both sides by p(L) (I am assuming a non-zero probability of survival), we get:

  • Ep(sL(p)) ≥ Ep(sL(q)).

We can combine this with a more complex set of future investigations as in the previous post, and things will still work.

It is crucial to the above argument that when you’re alive, you can tell you’re alive. I suppose that’s not always true. When you’re asleep, you are alive, but can’t tell you’re alive. So to generalize beyond the above toy example, replace death with unconsciousness or something like that.

Forward-looking scoring rules

An accuracy scoring rule assigns a score to a probability function representing an agent’s credences, ostensibly measuring how close that probability function is to the truth. The score s(p) of a probability function p is a random variable, because the value of the score depends on what is actually true, i.e., on where we are in the probability space.

A proper scoring rule (on probabilistic credences) satisfies the propriety inequality

  1. Eps(p) ≥ Eps(q)

which says that the expected score of your current lights—your current credences p—by your current lights is optimal: you won’t improve your expected score (by your current lights) by switching to a different credence q.

You can think of a proper scoring rule as representing the epistemic utility of having a credence p.

But now let’s think about things dynamically. In the future, you will receive additional evidence. As a good Bayesian agent, you will update on this evidence by conditionalization. Perhaps instead of thinking about maximizing your current score, you should think about maximizing your future score. Maybe your true epistemic utility is the score you will end up with after all the future evidence is in.

A simple model of this is as follows. There is some finite partition I = (I1,...,In) of your probability space Ω with each cell Ii of the partition representing a possibility for what you might learn given future evidence. Your current credence function is p, and p(Ii) > 0 for all i. There is then a random credence function pI where pI(ω) is the credence function you will have once the evidene is in if you are at ω ∈ Ω. In other words, pI(ω)(A) = p(A∣Ii) where Ii is the member of the partition that contains ω. (Technically, the function that maps ω to pI(ω)(A) is equal to the conditional probability p(A∣G) where G is the algebra generated by I.)

Now, given a proper scoring rule s, define a new scoring rule sI as follows:

  1. sI(p)(ω) = s(pI(ω))(ω).

Your sI-score for p at ω then represents the score you will have at ω once you learn which cell of the partition I you are in.

Theorem: The scoring rule sI is proper if s is proper.

Note that sI won’t be strictly proper (i.e., (1) won’t always have strict inequality when p and q are distinct) if I has two or more cells, because pI and qI are going to be the same if p and q assign different probabilities to the cells, but have the same conditional probabilities on each cell. But it might still be the case that sI is strictly proper with respect to some relevant subfield of Ω—that needs some further investigation.

Suppose now you are a Bayesian agent who is guaranteed to consciously live for n moments. In each moment, new information comes in. Thus, we have a sequence J0, ..., Jn of finer and finer partitions, with J0 being the trivial partition, and with pJk representing the credence you will have at time k. Your overall epistemic lifetime score is then:

  1. sΣ(p) = ∑ksJk(p).

It follows from the Theorem that sΣ is a proper scoring rule if s is. And if J0 is the trivial partition, then sJ0 = s, and so if s is strictly proper, then the lifetime score sΣ is strictly proper, since the sum of a strictly proper rule and a proper rule is strictly proper. So, lifetime scores are strictly proper if they are constructed from an instantaneous score—in the above toy model.

Alas, the toy model is not fully adequate, because it is random when we will die, and so our lifespan doesn’t have a fixed sequence of moments. Once we take into account the randomness of when we will die, the overall epistemic lifetime score might stop being proper: this needs further investigation.

Proof of Theorem: By the Greaves and Wallace Theorem, an optimal method of updating credences with respect to expected proper score is by Bayesian conditionalization. Apply the Greaves and Wallace Theorem to the scoring rule s and the starting credence p with the following two strategies:

A. Bayesian conditionalization on the true cell of I.

B. Switch your credence from p to q, then apply Bayesian conditionalization on the true cell of I.

Saying that (A) is at least as good as (B) is equivalent to the the proper scoring rule inequality (1) for sI.

Wednesday, October 22, 2025

Probably most people you know are more social than you

You might observe:

  1. Most of the people I know are more social than me.

And then you might beat up on yourself, concluding:

  1. I am less social than most people.

But the inference from (1) to (2) is obviously fallacious.

For whether you know a person is a function of how social you are and how social they are. Thus, the sample of people in (1) suffers from an evident sampling bias: it is skewed towards people who are more likely to be social.

How strong is this bias? Well, here is a model. There are N people. Each person has a sociality score between 0 and 1. Each person knows themselves. For each pair of distinct people, we independently decide if they know each other, with a probability equal to the average of their sociality scores. Then we calculate the fraction of people who have the property that most people they know have a higher sociality score.

Computer simulation gives us about 59% for N = 1000 or N = 1500 with sociality scores uniformly distributed from 0 to 1. I haven’t bothered to come up with a closed form solution.

So the bias isn’t that strong, but indeed most people are such that most people they know are more social than they are.

I just saw this more thorough related study.

Monday, October 20, 2025

Another infinite dice game

Suppose infinitely many people independently roll a fair die. Before they get to see the result, they will need to guess whether the die shows a six or a non-six. If they guess right, they get a cookie; if they guess wrong, an electric shock.

But here’s another part of the story. An angel has considered all possible sequences of fair die outcomes for the infinitely many people, and defined the equivalence relation ∼ on the sequences, where α ∼ β if and only if the sequences α and β differ in at most finitely many places. Furthermore, the angel has chosen a set T that contains exactly one sequence from each ∼-equivalence class. Before anybody guesses, the angel is going to look at everyone’s dice and announce the unique member α of T that is ∼-equivalent to the actual die rolls.

Consider two strategies:

  1. Ignore what the angel says and say “not six” regardless.

  2. Guess in accordance with the unique member α: if α says you have six, you guess “six”, and otherwise you guess “not six”.

When the two strategies disagree for a person, there is a good argument that the person should go with strategy (1). For without the information from the angel, the person should go with strategy (1). But the information received from the angel is irrelevant to each individual x, because which ∼-equivalence class the actual sequence of rolls falls into depends only on rolls other than x’s. And following strategy (1) in repeats of the game results in one getting a cookie five out of six times on average.

However, if everyone follows strategy (2), then it is guaranteed that in each game only finitely many people get a shock and everyone else gets a cookie.

This seems to be an interesting case where self-interest gets everyone to go for strategy (1), but everyone going for strategy (2) is better for the common good. There are, of course, many such games, such as Tragedy of the Commons or the Prisoner’s Dilemma, but what is weird about the present game is that there is no interaction between the players—each one’s payoff is independent of what any of the other players do.

(This is a variant of a game in my infinity book, but the difference is that the game in my infinity book only worked assuming a certain rare event happened, while this game works more generally.)

My official line on games like this is that their paradoxicality is evidence for causal finitism, which thesis rules them out.

Friday, September 19, 2025

Random causation across temporal gaps

Suppose that causation across temporal gaps is possible: that an object x can have a direct effect in a future time, with no intermediate causes. Given that a cause clearly can have a random effect—say, you press a button and you get a green light or a red light at random—then it should also be possible for a cause to have an effect at a random future time.

Now imagine a button that, when pressed, causes a flash of light at a random time in the future, from tomorrow onward, with the probability that the flash happens in n days being 1/2n.

This is not very different from a button that, when pressed, triggers a sequence of fair coin tosses, one per day, with a beep that goes off as soon as heads comes up. The probability that there will be a beep in n days is 1/2n.

But there is still an important difference between the flash and the beep, even though they are probabilistically isomorphic. The flash is guaranteed but the beep is not (it is possible to get tails everyday). On open-future views, it is true that the flash will happen but not true that the beep will.

One could imagine the flash method being used by God in connection with indefinite-time future promises like “One day I’m going to make a flash of light.” God can just create the button that causes the flash to happen on a random future day and then trigger the button.

Thursday, September 11, 2025

Why do we like being confident?

We like being more confident. We enjoy having credences closer to 0 or 1. Even if the proposition we are confident in is one that is such that it is a bad thing that it is true, the confidence itself, abstracted from the badness of the state of affairs reported by the proposition, is something we enjoy.

Here is a potential justification of this attitude in many cases. We can think of the epistemic utility of one’s credence r in a proposition p as measured by an accuracy scoring rule given by two functions T(r) and F(r), where T(r) gives the value of having credence r in p when p is actually true and F(r) gives the value when p is actually false. Most people thinking about scoring rules think they should satisfy the technical condition of being strictly proper. But strict propriety implies that the function V(r) = rT(r) + (1−r)F(r) is strictly convex. Now suppose the scoring rule is also symmetric, so that T(r) = F(1−r). Then V(r) is a strictly convex function that is symmetric about r = 1/2. Such a function has its minimum at r = 1/2, and is strictly decreasing on [0,1/2] and strictly increasing on [1/2,1]. But the function V(r) measures your expectation of your epistemic utility. How happy you are about your credence, perhaps, corresponds to your expectation of your epistemic utility. So you are most unhappy at credence 1/2, and you get happier that closer you are to 0 or 1.

OK, it’s surely not that?!

Monday, June 2, 2025

Shuffling an infinite deck

Suppose infinitely many blindfolded people, including yourself, are uniformly randomly arranged on positions one meter apart numbered 1, 2, 3, 4, ….

Intuition: The probability that you’re on an even-numbered position is 1/2 and that you’re on a position divisible by four is 1/4.

But then, while asleep, the people are rearranged according to the following rule. The people on each even-numbered position 2n are moved to position 4n. The people on the odd numbered positions are then shifted leftward as needed to fill up the positions not divisible by 4. Thus, we have the following movements:

  • 1 → 1

  • 2 → 4

  • 3 → 2

  • 4 → 8

  • 5 → 3

  • 6 → 12

  • 7 → 5

  • 8 → 16

  • 9 → 6

  • and so on.

If the initial intuition was correct, then the probability that now you’re on a position that’s divisible by four is 1/2, since you’re now on a position divisible by four if and only if initially you were on a position divisible by two. Thus it seems that now people are no longer uniformly randomly arranged, since for a uniform arrangement you’d expect your probability of being in a position divisible by four to be 1/4.

This shows an interesting difference between shuffling a finite and an infinite deck of cards. If you shuffle a finite deck of cards that’s already uniformly distributed, it remains uniformly distributed no matter what algorithm you use to shuffle it, as long as you do so in a content-agnostic way (i.e., you don’t look at the faces of the cards). But if you shuffle an infinite deck of distinct cards that’s uniformly distributed in a content-agnostic way, you can destroy the uniform distribution, for instance by doubling the probability that a specific card is in a position divisible by four.

I am inclined to take this as evidence that the whole concept of a “uniformly shuffled” infinite deck of cards is confused.

Friday, May 23, 2025

Hyperreal infinitesimal probabilities and definability

In order to assign non-zero probabilities to such things as a lottery ticket in an infinite fair lottery or hitting a specific point on a target with a uniformly distributed dart throw, some people have proposed using non-zero infinitesimal probabilities in a hyperreal field. Hajek and Easwaran criticized this on the grounds that we cannot mathematically specify a specific hyperreal field for the infinitesimal probability. If that were right, then if there are hyperreal infinitesimal probabilities for such a situation, nonetheless we would not be able to say what they are. But it’s not quite right: there is a hyperreal field that is "definable", or fully specifiable in the language of ZFC set theory.

However, for Hajek-Easwaran argument against hyperreal infinitesimal probabilities to work, we don’t need that the hyperreal field be non-definable. All we need is that the pair (*R,α) be non-definable, where *R is a hyperreal field and α is the non-zero infinitesimal assigned to something specific (say, a single ticket or the center of the target).

But here is a fun fact, much of the proof of which comes from some remarks that Michael Nielsen sent me:

Theorem: Assume ZFC is consistent. Then ZFC is consistent with there not being any definable pair (*R,α) where *R is a hyperreal field and α is a non-zero infinitesimal in that field.

[Proof: Solovay showed there is a model of ZFC where every definable set is measurable. But every free ultrafilter on the powerset of the naturals is nonmeasurable. However, an infinite integer in a hyperreal field defines a free ultrafilter on the naturals—given an infinite integer M, say that a subset A of the naturals is a member of the ultrafilter iff |M| ∈ *A. And a non-zero infinitesimal defines an infinite integer—say, as the floor of its reciprocal.]

Given the Theorem, without going beyond ZFC, we cannot count on being able to define a specific hyperreal non-zero infinitesimal probability for situations like a ticket infinite lottery or hitting the center of a target. Thus, if a friend of hyperreal infinitesimal probabilities wants to be able to define one, they must go beyond ZFC (ZFC plus constructibility will do).

Monday, April 28, 2025

Probabilities of regresses of chickens

Suppose we have a backwards-infinite sequence of asexually reproducing chickens, ..., c−3, c−2, c−1, c0 with cn having a chance pn of producing a new chicken cn + 1 (chicken c0 may or may not have succeeded; the earlier ones have succeeded). Suppose that the pn are all strictly between 0 and 1, and that the infinite product p−1p−2p−3... equals some number p strictly between 0 and 1.

Intuitively, we should be surprised that chicken c0 exists if p is low and not surprised if p is high. If we have observed c0 and are considering theories as to what the chances pn are, other things being equal, we should prefer the theories on which the product p is high to ones on which it’s low.

But what exactly does p measure? It seems to be some kind of a chance of us getting c0. But it doesn’t measure the unconditional probability of getting an infinite sequence of chickens leading up to c0. For that is very tiny indeed, since it is extremely unlikely that the world would contain chickens at all. It seems to be a kind of conditional probability. Let qn be the proposition that chicken cn exists. Then P(q0∣qn) = p0p−1p−2...pn, and so p is the limit of the conditional probabilities P(q0∣qn). It is plausible thus to think of p as a conditional probability of q0 on q−∞, which is the infinite disjunction of all the qn.

But q−∞ is a rather odd proposition. It is grounded in qn for every finite n, assuming that a disjunction, even an infinite one, is grounded in its true disjuncts. Thus every one of the qn is explanatorily prior to q−∞. But this means that P(q0∣q−∞) is actually a conditional probability of q0 on something that isn’t explanatorily prior to q0—indeed, that is explanatorily posterior to q0. This challenges the interpretation of p as a chance of getting chicken c0.

I am not quite sure what conclusion to draw from the above argument. Maybe it offers some support for causal finitism, by suggesting that things are weird when you have a backwards infinite causal sequence?

Friday, April 11, 2025

Unreliable Grim Reapers

As usual, Fred is alive at 10 am, and there is an infinite sequence of Grim Reapers, where the nth has an alarm set for 60/n minutes after 10 am, and if the alarm goes off, it checks if Fred is dead, and swings its scythe at Fred if and only if Fred is alive. But here’s the twist. These Grim Reapers are unreliable killers. The probability that the nth Reaper’s swing would succeed in killing Fred is 1/np, where p is some positive real number, the same for each Reaper, and independently of all other relevant events.

Here’s the fun thing. It seems possible for Fred to survive the whole ordeal. All it takes is for every Grim Reaper to fail at killing Fred. Nothing absurd happens then. Moreover, it seems this isn’t the only way for absurdity to be avoided in this case. We could also suppose that the nth Reaper kills Fred, while Reapers n + 1, n + 2, … all fail.

Suppose we adopt what seems the best alternative to Causal Finitism, namely the Inconsistent Pair response to the original Grim Reaper paradox, which says that the reason the original paradox is impossible is simply because it embodies an Inconsistent set of propositions—some Reaper has to kill Fred and none can. If that’s what’s wrong with the original Grim Reaper paradox, then it seems we have to accept my Unreliable Reaper story as possible.

But things are a little bit more complicated. The only way to avoid paradox in the Unreliable Reaper story is if there is some n ≥ 0 such that all the Reapers starting with Reaper n + 1 fail. But now suppose that 0 < p ≤ 1. Then the event that all the Reapers starting with Reaper n + 1 fail is less than or equal to (1−1/(n+1)p)(1−1/(n+2)p)(1−1/(n+3)p)... = 0 (this is because Σk 1/kp = ∞ if p ≤ 1). Thus the probability that we have avoided paradox is 0. Hence, if we have to avoid paradox, a specific zero probability event—namely, the event of paradox-avoidance—has to happen (the probability of a countable disjunction of zero probability events is zero). But if it has to happen, it can’t be probability zero, but must be probability one!

Perhaps here we bring back the Inconsistent Pair response. We say that my Unreliable Reaper story is impossible if p ≤ 1, because if p ≤ 1, then a zero probability event has probability one, which is inconsistent. No such problem occurs if p > 1. Thus, on this version of the Inconsistent Pair response, my Unreliable Reaper story is impossible if the success probability of the nth Reaper is 1/np for p ≤ 1 but possible if p > 1. And that’s pretty counterintuitive.

Wednesday, March 19, 2025

Provability and truth

The most common argument that mathematical truth is not provability uses Tarski’s indefinability of truth theorem or Goedel’s first incompleteness theorem. But while this is a powerful argument, it won’t convince an intuitionist who rejects the law of excluded middle. Plus it’s interesting to see if a different argument can be constructed.

Here is one. It’s much less conclusive than the Tarski-Goedel approach. But it does seem to have at least a little bit of force. Sometimes we have experimental evidence (at least of the computer-based kind) for a mathematical claim. For instance, perhaps, you have defined some probabilistic setup, and you wonder what the expected value of some quantity Q is. You now set up an apparatus that implements the probabilistic setup, and you calculate the average value of your observations of Q. After a billion runs, the average value is 3.141597. It’s very reasonable to conclude that the last digit is a random deviation, and that the mathematically expected value of Q is actually π.

But is it reasonable to conclude that it’s likely provable that the expected value of Q is π? I don’t see why it would be. Or, at least, we should be much less confident that it’s provable than that the expected value is π. Hence, provability is not truth.