Showing posts with label Bayesianism. Show all posts
Showing posts with label Bayesianism. Show all posts

Tuesday, June 30, 2026

The thesis of the uniqueness of our self-consciousness

I am toying with the thesis that each person’s self-consciousness is qualitatively different. Perhaps all persons have uniquely individuating properties, something like haecceities, and for each such individuating property there is a different quale of seeing oneself as an instance of it. One might even think the quale is the unique individuating property. Or, alternately, it might be that each person’s qualia are different: my quale of red is a quale of me-seeing-red, my quale of middle-C is a quale of me-hearing-middle-C and your quale of red is a quale of you-seeing-red while your quale of middle-C is a quale of your-hearing-middle-C. There is, presumably, something that all of my qualia have in common, and all your qualia have in common, and so on.

What consequences might there be of such a view?

First, some people have the intuition that there is something mysterious about self-consciousness, something that goes significantly beyond mere consciousness (pace my thoughts here). The qualitative conscious difference thesis could be a way of capturing that mysterious extra in self-consciousness.

Second, this would be a way to support and explain intuition of the non-fungibility and irrepeatability of persons. There is something irrevocably unique about all our perceptions of the world.

Third, we might get a neat account of second-person knowledge. What happens in second-person knowledge—very mysteriously!—is that we come to see that special qualitative feature of the other, albeit in a different mode, from the outside.

Fourth, the thesis might make one even more cautious about some thorny epistemological matters. There cannot be such a thing as two people with the same qualitative evidence: each person’s evidence must be qualitatively different from that of every other person. The concept of being an epistemic peer becomes dubious: you unavoidably see the world differently from me. The application of Bayesianism to fundamental matters becomes even more fishy. Plausibly, we have essentiality of origins, so that if there is a God, you and I couldn’t have existed in a world without God. On the qualitative uniqueness thesis, this means that there are qualitative features of our mental lives that cannot be exemplified in a world without God. We can then expect that we will have some trickiness when we ask questions like: “How likely is it that the evidence I have would obtain if God didn’t exist?” Similar issues come up for fundamental laws of nature. In particular, we cannot really do Bayesian reasoning on our total evidence, and we should probably see Bayesianism as a mere toy model.

Wednesday, June 17, 2026

Self-locating evidence and bearers of epistemic good

In the case of non-epistemic goods, it’s an obvious feature of life that someone there is a choice to be made by an individual between their own first-order good and the first-order good of the community—each requires the sacrifice of the other. In the case of epistemic goods, this is less obvious.

In the pragmatic case, the typical reason for such competition between goods is due to limited resources. This, of course, also happens in the epistemic sphere. Suppose Alice is much more intellectually talented than Bob, but only Bob has the money to go to university. If Bob spends the money on himself, he will gain private epistemic goods, but will contribute little epistemically to society as a whole. But if he gives the money to Alice, she may become a brilliant scholar or scientist, significantly contributing to society’s knowledge.

More interesting than these, however, are cases of competition between private and communal epistemic goods that are not due to epistemic resources. I find it interesting that some cases of self-locating evidence appear to be such.

Suppose there are ten billion people in the world, currently isolated from one another. A device produced by a mad scientist has a 99.9% chance at noon today of triggering a death ray that randomly kills 99.9999% of the world population. Noon has just passed. You are still alive. Should you think the device worked? Sleeping Beauty style arguments say “No”. This time I want to think about this in terms of individual epistemic goods. In N runs of the device, 0.001N runs will have you survive because the device doesn’t trigger and 0.999 ⋅ 0.000001N runs will have you survive despite the device triggering. Thus, the vast majority of the runs where you survive are runs where the device didn’t trigger. Hence, it’s best for you individually to adopt the epistemic policy of thinking the device didn’t trigger.

But on the other hand, suppose we all adopt the epistemic policy of thinking the device didn’t trigger. Then 99.9% of the time, we are unanimously collectively wrong. And if we all adopt the epistemic policy of thinking the device did trigger, then 99.9% of the time, we are unanimously collectively right. It seems thus that if we look at the epistemic goods of society, then a policy of thinking the device did trigger is best.

If this is right, it points to a potential diagnosis of why the problems about self-locating evidence (doomsday, multiverses, Sleeping Beauty, etc.) are so difficult. For there may be different bearers of epistemic goods at play—say, society vs. the individual—and it could be that different answers are appropriate depending on whose goods we are pursuing. Maybe.

Tuesday, June 16, 2026

Is there some sort of a probability problem with a humongous but finite universe?

It’s easy to generate probabilistic paradoxes in a universe (or multiverse) with infinitely many people (e.g., if infinitely many people roll a die, equal numbers of people get 1 as get more than 1, so why think it’s more likely to get more than 1?). But what about a very large but finite universe? I used to think: “The only relevant difference is between finite and infinite. Really big but finite—no problem.” Now I am not so sure.

Paul Heyl measured the gravitational constant G as 6.670 × 10−11 m3 kg−1 s−2, and denote the latter quantity by G0. Consider two theories:

  • H1: The gravitational constant is between 6.665 × 10−11 m3 kg−1 s−2 and 6.675 × 10−11 m3 kg−1 s−2.

  • H2: The gravitational constant is between 7.676 × 10−11 m3 kg−1 s−2 and 7.686 × 10−11 m3 kg−1 s−2.

It seems obvious that:

  1. Heyl’s measurement strongly supports H1 but does not completely rule out H2.

But let’s think this through. Suppose Heyl’s evidence is the proposition E which he would express as “I measured G to be G0.” But, very plausibly, it is an essential property of a human being that they exist in a world with such-and-such a gravitational constant. One way of getting to this conclusion is to say that the forces of gravity are part of our causal history, and then to apply the essentiality of origins. Another is to say that we couldn’t have been made of completely different matter, but the forces exerted by the matter in our bodies are an essential property of that matter.

Given this essentiality of gravitational constant assumption, it follows that at least one of H1 and H2 is incompatible with Heyl’s existence. Now, to get (1), we need prior probabilities on which P(H1|E) > P(H2|E) > 0. Such prior probabilities will assign a non-zero value to H1E and to H2E. But at least one of these two claims is impossible since E entails Heyl’s existence, and a probability assignment that assigns a non-zero value to something impossible is screwed up, and we should be quite suspicious of what we get from it.

We might try to avoid this by using self-locating evidence. But my colleague Yoaav Isaacs has this great paper that gives a pretty strong argument that there isn’t a good way to working with self-locating evidence. So suppose we put this option aside.

Or we might make a distinction between logical impossibility and metaphysical impossibility. I find that suspicious, too.

So, what’s left? Well, here’s one remaining suggestion. Heyl’s evidence is equivalent to the proposition that Heyl measured G to be G0, a proposition that rigidly refers to Heyl, and hence won’t be compatible with both H1 and H2. But we can weaken Heyl’s evidence to something that is compatible with H1 and H2, something purely qualitative, like:

  • EQ: A physicist named “Paul Heyl”, who married someone named “Lucy Daugherty”, and who …, measured G to be G0.

Here, “…” is all the other purely qualitative stuff we know about Paul Heyl, so that EQ is compatible with both H1 and H2.

But now here is a problem. Suppose we live in a vast but finite universe with, say, 101010 people. In such a universe, we might well expect large numbers of people named “Paul Heyl” who satisfy all the conditions in EQ, including the measurement of G to be G0, even if in fact G is in the range indicated in H1 (measurement error!). Thus, P(EQ|H2) is close to 1 as is P(EQ|H1). Granted, we do have P(EQ|H1) > P(EQ|H2) > 0. But because the two probabilities are so close to each other, the support EQ gives to H1 over H2 is very slight, and hence we no longer have (1).

It follows that unless we can find some other way of solving the problem that the essentiality of the laws of nature to humans poses for Bayesian reasoning, a fair amount of fundamental physics research would be undercut by a large enough—even if finite—universe.

Of course, maybe we can find some other way of solving it. But maybe we can’t. And if we can’t, then the EQ solution might be our best bet—and it’ll work just fine in a universe that isn’t too vast.

Friday, June 5, 2026

Knowledge and induction

Assume that in fact all ravens are black. Suppose you are sequentiallly observing ravens, and noting each one to be black. After observing n ravens, your evidence that the next raven is black will typically be significantly better than the evidence that all ravens are black. Now at some point, say after observing nA ravens, your evidence that all ravens are black will rise to the level of knowledge. Thus, plausibly, at an earlier point in the sequence, call it nN, your evidence that the next raven is black will have risen to the level of knowledge.

Suppose now you have observed nA − 1 ravens, and you have been handed a raven in an opaque box, which you are certain you are about to open. Since nN < nA, at this point you have reached nN. Hence:

  1. You do not know that all ravens are black.

  2. You do know that the next raven is black.

  3. You know that when you observe the next raven, you will have sufficient evidence for knowledge that all ravens are black.

But note that while you know you will have sufficient evidence for knowledge that all ravens are black, you don’t know that you will know that all ravens are black. There is nothing deeply surprising about this distinction. We might well say about someone who has been subjected to misleading or Gettiered evidence that they have sufficient evidence to know something but nonetheless they don’t know, though the case at hand feels different.

One interesting thing about this case, as I read it, is that it contradicts the thesis that K = E, i.e., that knowledge is evidence. For if knowledge is evidence, and you know that the next raven is black, then you already have the evidence you will gain by observing the next raven, and hence you are already in the position to know that all ravens are black.

Another interesting thing is that it shows that you can know something and nonetheless it be rational for you to investigate it. For you know that the next raven is black, but it’s worth investigating further, since it is only upon observation that your knowledge of the next raven’s blackness turns into the kind of evidence that gives you knowledge that all ravens are black.

All this might make one think that I have misconstrued the epistemic facts, and it is false that there can be a point nN prior to nA at which you know that the next raven is black. Here is one way to back up my intuition that there can be such a point nN < nA. Suppose that we know for sure we live in a world where the color distributions of birds are always uncorrelated between the males and the females of the species, so that information about the color of members of one sex are irrelevant to the color of members of the other sex. Also assume that you know for sure that ravens have equal numbers of each sex, that you are observing ravens in an alternative female-male-female-male-… sequence, and that your priors for the color distributions of the two sexes of ravens are the same. Then if pM and pF are the probabilities that all male ravens are black and all female ravens are black, and pA is the probability that all ravens are black, then at any given point in the observation sequence pA = pMpF. Let nM and nF be the points in the sequence where you know that all male and all female ravens are black, respectively. Then, nM < nA and nF < nA, since pA = pMpF is significantly smaller than either pM and pF at all points in the sequence except when we’ve observed all the ravens of one sex, and since pM and pF rise fairly gradually as we go through the sequence. Thus, at the point nA − 1, we will have already reached knowledge that all the male ravens are black and the knowledge that all the female ravens are black. In particular, then, we know that the next raven is black, since if nA is even, the nAth raven is male and we know all male ravens are black, and if nA is odd, then the nAth raven is female, and we know that all female ravens are black.

Wednesday, June 3, 2026

Epistemic rationality and Pascal's Wager

Pascal’s Wager is an argument that it is prudentially rational to engage in theistic belief promotion practices (TBPP), namely practices apt to promote one’s belief in God.

My interest this post is the standard epistemic rationality objection to the Wager, that engaging in TBPPs is irrational—a kind of brainwashing of oneself. Let’s think about the objection with a bit more care. Consider a specific TBPP Q, say Pascal’s example of going to Mass. If Q is indeed a TBPP, one expects engagement in TBPP to make it more likely that one believes in God. But how does one expect Q to achieve that goal? There are two possibilities. Either Q is expected to achieve that goal by providing one with evidence for theism or in some non-evidential way (or a combination of the two).

Suppose Q is expected to work evidentially. Then we already have the expectation of a higher credence given Q. This expectation is either rational or not. If it is not rational, then we don’t actually have good reason to engage in Q. If it is rational, however, then we should rationally raise our credence in theism right now, without having to bother engaging in Q, and for reasons having nothing to do with any wager. But if Q promotes belief in God non-rationally, then we should not engage in Q for the sake of such promotion of belief—we should not aim to non-rationally promote beliefs.

Let me make the first horn of the dilemma—namely, that the expectation of a higher credence is rational—a bit more precise. We can distinguish two (not mutually exclusive) ways in which a practice rationally increases one’s credence in a hypothesis H. One way is purely epistemic, by uncovering facts about reality. This is the usual way. But if that’s the way we expect to increase our credence in theism by engaging in Q, then we already have evidence that there are such theism-indicating facts to be discovered, and so we should already increase our credence in Q. The other way is practical, by promoting the hypothesis H in a way that shows up to us. The second way is a bit unusual, but here is an example: one way to increase your credence that you will not die of heart disease is to live a healthy life. For if you live a healthy life, you are less likely to die of heart disease, and since you will notice signs of improved cardiac health (e.g., lower resting heart rate, less huffing and puffing on stairs, etc.), your credence that you won’t die of heart diseases will also increase. But this practical way of increasing credence is utterly irrelevant in the case of theism, since nothing we can do can make God more or less likely to exist! So the only rational way that remains is the evidence-based way, and evidence-of-evidence is already evidence.

I used to be quite impressed by the worry that Pascal’s Wager leads to self-deception. I am less impressed. Here is why. There is a serious technical flaw in the argument for the first horn of the dilemma. A simple model for the relation of credence and belief is that you believe a proposition if and only if your credence is above some threshold β. This model might be false, but an analogue of what I will say should apply on more sophisticated models as well.

Here is the point. Consider a case where one is thinking about observing (and suppose this is a simple non-Newcombian observation that does not affect the hypothesis) whether some event E evidentially relevant to H has obtained. Then one’s expected posterior credence is:

  1. C(E)C(HE) + C(∼E)C(H∣∼E),

where C is one’s credence function. But if one is a good Bayesian reasoner, then by total probability the value in (1) is simply equal to one’s prior C(H). Thus the value of one’s credence has no expectation of change upon observation when one is being rational. This seems to support the idea that if you expect your credence to go up, you should already raise it.

But in fact it’s not so simple. For even though the expected posterior credence equals one’s current credence, it could well be that it is more likely that the expected posterior credence exceeds the threshold β if you make the observation than if you don’t. Indeed, cases are obvious. Suppose the belief threshold β is 0.9, and you tossed a coin out of my sight. Suppose I have a prior credence 0.5 that this coin is fair and a prior credence 0.5 that it is double-headed. Then currently I don’t believe (or disbelieve) that the coin is fair. But if I look at the coin and I see tails, I will believe that it is fair—indeed, I will have posterior credence 1 in its fairness. But if I don’t look at the coin, I am not going to get any evidence, and I will continue not to believe that the coin is fair. If I look at the coin, the probability that I will see tails is 0.25 (I have credence 0.5 that it’s fair, and if it’s fair, the chance of tails is 0.5), and so the probability that I will believe if I look is 0.25 (since if I look and see tails, my posterior will be 1 which is bigger than β = 0.9), and the probability that I will believe if I don’t look is 0. Of course, if I don’t see tails, my credence that the coin is fair will go down. But while it will go down, it won’t affect whether I believe that the coin is fair—for I already don’t believe it (in the sense of not-believe, rather than in the sense of believe-not). And there is no irrationality of any sort in looking at the coin in this case.

In other words, the point is that while one’s rational credence has no positive or negative rational expectation of change upon observation, whether one’s rational credence is above a threshold certainly can have a positive or negative rational expection of change.

How could this work in a Pascal’s Wager situation? Let’s talk through one possibility. Take Pascal’s example of going regularly to Mass. Suppose, as Pascal says, your current credence in God is 0.5. You might think that if God exists, going to Mass has a decent chance, say 0.2, of resulting in an evident radical transformation E of your life, so evident and radical that updating your credence in theism on E will push your credence in God to above the threshold β. Of course, you might go to Mass it might not produce any such evident radical transformation (this is true even if God always improves the hearts of people who go to Mass, since he might do so more gradually or less evidently), and in that case your rational credence in God will go below 0.5. But going below 0.5 won’t affect whether you believe in God, since 0.5 is already, I assume, far below the belief threshold β. On the other hand, maybe if you don’t go to Mass, the probability that you will get evidence that will push your credence in theism above β is pretty small—smaller than 0.2. Very likely, your credence will just oscillate a little around 0.5 in the non-Mass-going case. Thus if there is a payoff you get for your credence exceeding the threshold β, it will be worth going to Mass, without there being any epistemic irrationality in the reasoning.

Thought of in this way, we get some practical guidance as to which TBPPs the agnostic or atheist should engage in. They should look for practices that, if God exists, have a decent chance of producing evidence for theism sufficient to push them above β.

I am a bit doubtful that Pascal meant us to think in the above way. He may well have been recommending TBPPs on the grounds of their non-rational effect on belief. My argument above does not defend that.

Wednesday, April 22, 2026

Extending Good's Theorem to experiments and not just observations

Good’s Theorem basically says that a utility-maximizing agent can expect to make decisions that are at least as good if they get more information. (And under some additional conditions, one can expect the decisions to be better.)

Now consider this case:

  1. You will be offered a chance to make a bet at certain odds on the result of a coin toss, where as far as you can tell it’s equally likely that the coin is fair and that it is double-headed. Someone offers to tell you how the previous toss of the coin went.

Good’s Theorem says your decision whether to make the bet will be at least as good given the information about the previous three tosses as without that information. Hence, if the information is being announced, you don’t need to cover your ears. This is, of course, very intuitive. But now consider a slightly different case:

  1. Things are set up just as in (1), except now instead of information about the previous toss, you are offered a chance to have the following experiment get performed before your decision: the coin will be tossed an extra time and the result will be announced to you.

The difference is that in (2) you are not simply being offered additional information about how things are. For whether you go for the experiment or not, either way, you have full information about the experiment and its results. If you don’t go for the experiment, that full information is that the coin was not tossed an extra time (and hence did not land either heads or tails). If you do go for the experiment, the full information is that the coin was tossed and it landed heads, or else that it was tossed and it landed tails. In (2), you are not just finding out information by going for the deal: you are making something happen—an extra toss—and then finding out something about that.

So you can’t apply Good’s Theorem directly to (2). It would be nice to have a formulation of Good’s Theorem that works in cases where instead of merely finding out information, you perform an experiment.

I initially thought this would be easy. Maybe it is, but I don’t see it. There are, after all, cases where performing a cost-free experiment is not a good idea. Suppose, for instance, that you will be allowed to bet tomorrow that a certain car has more than 10 gallons of gasoline. The experiment is to start up the car and look at the gas gauge. But starting the car reduces the amount of gasoline in it, and one can easily rig the case so that benefits from the information gain are outweighed by the fact that you have made that bet less favorable.

So, we want to rule out cases where there is dependence between whether you perform the experiment and the payoffs of the wagers. If F is the event of performing the experiment, it may seems initially we should assume something like:

  1. E(U|WiF) = E(U|WiFc) for all i,

where Wi is your choosing wager i and U is the utility random variable. In other words, the expected utility of each wager is unaffected by whether the experiment has been performed. But no! Suppose a coin has been tossed, and you are choosing between W1 where you get a dollar on heads and W2 where you get a dollar on tails. But let F be the experiment of looking at the coin. (This is a case for the original Good’s Theorem.) Then E(U|WiFc) = 0.50, while E(U|WiF) is very close to 1.00 for the reason that when you find out what the coin is like, you are close to certain to bet on what you see, and hence you are close to certain to win your bet.

If F1 is heads and F2 is tails, we solve the problem by replacing (3) with:

  1. E(U|WiFjF) = E(U|WiFjFc) for i and j.

Namely, the expected utility of wager Wi given information Fj is independent of whether you performed the experiment F. But that only works because it makes sense to ask what the coin is showing if you aren’t looking: it makes sense to conditionalize on Fj ∩ Fc. But in the cases that interest me, there is no fact of the matter as to the result of the experiment when the experiment is not performed, since Molinism is false and we live in an indeterministic world. And in these cases, Fj ∩ Fc is the empty set: the Fj represent the possible results of the experiment but the experiment has no result when it is not performed.

I can get something by supposing a two-step procedure. You perform the experiment, event F, and you learn the result, event L. Then we can assume:

  1. E(U|WiFLc) = E(U|WiFc) for all i

  2. E(U|WiFjFL) = E(U|WiFjFLc) for all i and j

  3. P(Fj|FL) = P(Fj|FLc).

Assumption (5) says that it makes no difference to the expected utility of a wager whether (3) the experiment is performed but its result is not learned or (b) the experiment is not performed at all. In other words, the experiment itself doesn’t affect things. Assumption (6) says that given a specific experimental result, learning the result makes no difference to the expected utility of each wager–result pair. Assumption (7) says that the results of the experiment are unaffected by whether you learn the result of the experiment.

Without (6) or (7), we wouldn’t expect to get the result we want. If we don’t have (6), it might be that utilities are wildly affected by whether you learn the result. (The simplest case is that the wagers all have a big negative payoff on L.) If we don’t have (7), then learning the result might have some evidential or retrocausal impact on what the result is, and then again we shouldn’t expect that learning the result is a good thing.

Given (5)–(7), I think we can now reason as follows. You are choosing between:

  1. performing the experiment and learning the results

and

  1. not performing the experiment and (hence) not learning the results.

By (5), a rational agent will decide the same way in (ii) as in:

  1. performing the experiment and not learning the results,

and the expected utilities of (ii) and (iii) will be the same for this rational agent.

We now apply Good’s Theorem to the choice between (i) and (iii) (we will use (6) and (7) here, and assume the case is non-Newcombian and hence allows the use of Evidential Decision Theory) and get the result that (i) is at least as good as (iii). Since we have indifference between (ii) and (iii), it follows that (i) is at least as good as (ii). (We can also analyze the cases of a strict expected utility inequality.)

This is roundabout, but that’s not my main worry.

What I am really worried about is one technicality. To run the above argument, I had to assume that there is a way of performing the experiment without learning the result, namely that F ∩ Lc is non-empty. In general, however, we cannot assume this. Suppose, for instance, that we have a world with a quantum mechanics where observation causes collapse. Then the experiment of collapsing a wavefunction by means of observation cannot be done without observing the result of the experiment. In such scenarios, I cannot simply introduce a third option of performing the experiment and not learning the results, since that third option may not be consistent with the laws of physics. (And, of course, the utilities for breaking the laws of physics could be wild.)

But without introducing that third option, namely F ∩ Lc, I don’t know how to formulate the independence assumptions that are needed. I also don’t know if the problem is “merely technical” or “deep”. If I had to bet at even odds, I would bet on its being merely technical. But it might be deep.

Monday, April 13, 2026

A double lottery and non-normalized probabilities

Suppose a positive integer N is generated by a fair lottery.

Then, a random integer K is chosen between 1 and N (inclusive).

What information does this give you about N?

Obviously you now know that N ≥ K. Anything else?

Consider some specific pair of numbers n ≥ k, and suppose we’ve found out that K = k. What’s the probability that N = n? Of course P(N=n|K=k) = 0/0. But what if we do this as a limiting procedure. Suppose first that N is randomly chosen between 1 and M where M ≥ n, and let PM be the probabilities for this case. Then

  • PM(N=n|K=k) = (1/M)(1/n)/[(1/M)Σj=kMj−1] = (1/n)/Σj=kMj−1.

Take the limit as M goes to infinity. Since Σn=kj−1 = ∞, the limit is zero, so we don’t have a meaningful distribution for N.

On the other hand, what if we independently choose two random integers K1 and K2 between 1 and N? Suppose n ≥ ki for i = 1, 2. Let k* = max (k1,k2). Then:

  • PM(N=n|K1=k1,K2=k2) = (1/M)(1/n2)/[(1/M)Σj=k*Mj−2] = (1/n2)/Σj=k*Mj−2.

Take the limit as M → ∞ and call that P(N=n|K1=k1,K2=k2). The limit behaves like ck*/n2, for a constant c > 0, and generates a well-defined probability for N = n.

With zero samples, we don’t have a well-defined probability for N. With one sample, we still don’t. But with two samples (or more), now we do. This is a rummy thing: how is it that sampling turns probabilistic nonsense into sense?

This is making me more friendly to using non-normalized probabilities. After all, the fair lottery for N is easily modeled by the constant probability p0(n) = 1. With one sample N = k, we have p1(n) = 1/n for n ≥ k and p1(n) = 0 for n < k. With two samples k1, k2, we have p2(n) = 1/n2 for n ≥ max (k1,k2) and p2(n) otherwise. All this makes perfect sense. And there is a lovely mathematical feature of non-normalized probabilities: conditionalization is conjunction. The conditional probability of an event A on event B is just the probability of A ∩ B.

Non-normalized probabilities aren’t going to solve all problems with infinite fair lotteries. For instance, I toss a fair coin and generate a number N with the following rule. On heads, I choose N with my fair lottery on the positive integers. On tails, I choose N such that the probability of N = n is 2n (e.g., I toss an independent fair coin and let N be the number of the first toss that gives heads). What’s my non-normalized probability p(x,n), where x is heads or tails and n is a positive integer? We surely want np(H,n) = ∑np(T,n): the total probability of the heads options equals the total probability of the tails options. But clearly p(T,n) has to exponentially decrease so np(T,n) is finite and non-zero. On the other hand, p(H,n) is constant, so np(H,n) is zero or infinity. So they can’t be equal.

But I wonder if one could say something like this: Non-normalized probabilities make sense in certain cases, and in those cases it’s reasonable to use them?

Thursday, April 9, 2026

Predictability and epistemic utility

You’re thinking whether to become an assembly-line worker or an artist. Then you reflect on the value of knowledge. And you become a factory worker, on the grounds that if you become an assembly-line worker, you will know what you’ll be doing every working day of your future, but if you’re an artist, your activities will be unpredictable.

Some remarks. First, there is something perverse about using the value of knowledge in this way. The normal way to pursue the value of knowledge is to find out things that are independent of your pursuit. But here you are pursuing knowledge by making there be less to know about the world (or your world). Yet, paradoxically, it sure seems like the line of thought above makes sense.

Second, the my initial story depends on Molinism being false. For if there are comprehensive subjective conditionals of free will, then by becoming an artist you get to know the conditionals about what you would do in the various artistic situations you’re in. But on the assembly line story, you don’t get to know these. So the Molinist doesn’t have the paradox. I suppose that’s a bit of evidence for Molinism.

Life-time epistemic utility

I’ve been thinking about the diachronic aspects of epistemic utility. In the case of non-epistemic utility, we can get a decent first approximation to life-time utility by adding up (or, if time is continuous, integrating) momentary utility. But I think this works less well for the epistemic case. For many things of purely epistemic importance, figuring them out is much more important than when one figures them out. (Granted, figuring them out earlier is instrumentally epistemically valuable, because it gives one more time to use leverage the knowledge to figure other things.)

Here’s an extreme version of encoding the value of “figuring something out”. Assuming one does not suffer from mental decline, the epistemic value of one’s life is the epistemic value of the very last moment of it. It’s interesting to note that this won’t work. For imagine that no matter what other credences you had at a given time, you always set the credence of “This is the last moment of my life” to one, while being careful to (inconsistently) make no use of this credence in update. If only the last moment counts, this modification to your credences would be a good idea: it makes sure that when the last moment comes, you get the epistemic utility credit for it.

I suspect that other weightings that favor later over earlier beliefs will suffer from a similar problem—they make it a good idea to err on the side of pessimism about how close death is.

But at the same time, I think some sort of favoring of later beliefs over earlier ones seems appropriate. I don’t know how to resolve this difficulty.

Epistemic utilities and death

In the previous post, I proved that we get a proper scoring rule if we compute epistemic utilities as follows. We start with our current credence assignment, consider what credence assignment we will have in the future after we update on some further evidence, and then score that. I then suggested that one could get a lifetime epistemic utility by adding up the epistemic utilities over all the moments of life, and as long as death wasn’t random—as long as the lifespan was fixed—this would generate a proper scoring rule. I then said that if death is random (as it is) it might be the case that you don’t get a proper scoring rule.

My conjecture was wrong. You still get a proper scoring rule despite random death. It’s easy to see this. The basic idea is this. Suppose that you might die the next moment. This partitions the probability space into two subsets, D and L, for death and life. Your current credence is p. Next moment, on D, your credence either doesn’t exist (because you don’t exist or you exist in some supernatural state where you don’t have credences) or doesn’t count (because I am after lifetime credences). Thus, the appropriate way to do a forward-looking scoring of your credence p is to score it s(pL) on L and 0 on D, where pL is the result of conditionalizing your credence on the evidence L (after all, if you are alive, you will conditionalize on being alive), and s is some proper scoring rule. In other words, your forward looking score is sL(p) = 1L ⋅ s(pL).

Is this score proper? Yes! For by propriety of s we have:

  • EpL(s(pL)) ≥ EpL(s(qL)).

But this is the same as:

  • (p(L))−1Ep(1Ls(pL)) ≥ (p(L))−1Ep(1Ls(qL)).

Multiplying both sides by p(L) (I am assuming a non-zero probability of survival), we get:

  • Ep(sL(p)) ≥ Ep(sL(q)).

We can combine this with a more complex set of future investigations as in the previous post, and things will still work.

It is crucial to the above argument that when you’re alive, you can tell you’re alive. I suppose that’s not always true. When you’re asleep, you are alive, but can’t tell you’re alive. So to generalize beyond the above toy example, replace death with unconsciousness or something like that.

Forward-looking scoring rules

An accuracy scoring rule assigns a score to a probability function representing an agent’s credences, ostensibly measuring how close that probability function is to the truth. The score s(p) of a probability function p is a random variable, because the value of the score depends on what is actually true, i.e., on where we are in the probability space.

A proper scoring rule (on probabilistic credences) satisfies the propriety inequality

  1. Eps(p) ≥ Eps(q)

which says that the expected score of your current lights—your current credences p—by your current lights is optimal: you won’t improve your expected score (by your current lights) by switching to a different credence q.

You can think of a proper scoring rule as representing the epistemic utility of having a credence p.

But now let’s think about things dynamically. In the future, you will receive additional evidence. As a good Bayesian agent, you will update on this evidence by conditionalization. Perhaps instead of thinking about maximizing your current score, you should think about maximizing your future score. Maybe your true epistemic utility is the score you will end up with after all the future evidence is in.

A simple model of this is as follows. There is some finite partition I = (I1,...,In) of your probability space Ω with each cell Ii of the partition representing a possibility for what you might learn given future evidence. Your current credence function is p, and p(Ii) > 0 for all i. There is then a random credence function pI where pI(ω) is the credence function you will have once the evidene is in if you are at ω ∈ Ω. In other words, pI(ω)(A) = p(AIi) where Ii is the member of the partition that contains ω. (Technically, the function that maps ω to pI(ω)(A) is equal to the conditional probability p(AG) where G is the algebra generated by I.)

Now, given a proper scoring rule s, define a new scoring rule sI as follows:

  1. sI(p)(ω) = s(pI(ω))(ω).

Your sI-score for p at ω then represents the score you will have at ω once you learn which cell of the partition I you are in.

Theorem: The scoring rule sI is proper if s is proper.

Note that sI won’t be strictly proper (i.e., (1) won’t always have strict inequality when p and q are distinct) if I has two or more cells, because pI and qI are going to be the same if p and q assign different probabilities to the cells, but have the same conditional probabilities on each cell. But it might still be the case that sI is strictly proper with respect to some relevant subfield of Ω—that needs some further investigation.

Suppose now you are a Bayesian agent who is guaranteed to consciously live for n moments. In each moment, new information comes in. Thus, we have a sequence J0, ..., Jn of finer and finer partitions, with J0 being the trivial partition, and with pJk representing the credence you will have at time k. Your overall epistemic lifetime score is then:

  1. sΣ(p) = ∑ksJk(p).

It follows from the Theorem that sΣ is a proper scoring rule if s is. And if J0 is the trivial partition, then sJ0 = s, and so if s is strictly proper, then the lifetime score sΣ is strictly proper, since the sum of a strictly proper rule and a proper rule is strictly proper. So, lifetime scores are strictly proper if they are constructed from an instantaneous score—in the above toy model.

Alas, the toy model is not fully adequate, because it is random when we will die, and so our lifespan doesn’t have a fixed sequence of moments. Once we take into account the randomness of when we will die, the overall epistemic lifetime score might stop being proper: this needs further investigation.

Proof of Theorem: By the Greaves and Wallace Theorem, an optimal method of updating credences with respect to expected proper score is by Bayesian conditionalization. Apply the Greaves and Wallace Theorem to the scoring rule s and the starting credence p with the following two strategies:

A. Bayesian conditionalization on the true cell of I.

B. Switch your credence from p to q, then apply Bayesian conditionalization on the true cell of I.

Saying that (A) is at least as good as (B) is equivalent to the the proper scoring rule inequality (1) for sI.

Friday, February 20, 2026

A test case for explanationism

Here is a test case for explanationist stories about initial priors, on which a more explanatory theory has a higher initial prior. Consider these hypotheses:

  1. The universe is not created or governed by reason, has a finite lifetime, and has low initial entropy.

  2. The universe is not created or governed by reason, has a finite lifetime, and has low final entropy.

Causation goes from past to future, so absent some kind of foresight involved in the initial conditions, initial conditions are more explanatory than final conditions. Now, presumably there is a nice one-to-one correspondence between worlds where (1) and (2) hold, so absent an explanationist bias in the priors, we shouldn’t have a preference between (1) and (2).

This is just a test case, but I have to confess I don’t have any direct intuition comparing the probabilities of (1) and (2). I like explanationism, so I have a theory-laden reason to assign a higher probability to (1) than to (2).

Wednesday, February 4, 2026

Algorithmic priors and human nature

One promising way to define priors is with algorithmic probability, such as Solomonoff priors. The idea is that we have a language L (say, one based on Turing machines), and we imagine generating random descriptions in L in a canonical way. E.g., add an end-of-string symbol to L and randomly and independently generate symbols until you hit the end-of-string symbol, and then conditionalize on the string uniquely describing a situation, and take the probability of a specific situation s to be the probability of that a random description so generated describes s.

These kinds of priors are rather appealing for science, since they appear to be induction-friendly, as they assign high probabilities to compressible—more briefly expressible—situations. Thus, if our situations are distributions of color among ravens, monochromatic distributions get much higher probability as they can be much more briefly described, like x(B(x)) or x(W(x)).

Philosophically, I think the big problem is with the choice of the language. It would be nice if we could let L be a language that cuts nature exactly at the joints. But we don’t know that language. And absent that language, we need something arbitrary.

Here is a particular version of the problem. Take Kuhn’s division of science into ordinary and revolutionary science. One aspect of this division is that in ordinary science, we have a scientific language, and are discovering things within it. In that case, it is reasonable to take L to be that language. However, when we are doing revolutionary science and creating new paradigms, we cannot do that. The new paradigms either cannot be described in the old language or their description is unwieldy in a way that does not do justice to the plausibility of the new paradigm. Indeed, much of the point of revolutionary science is to create a language within which the description of the world is simpler, and then argue that this language is therefore more likely to cut nature at the joints.

Another version of this problem is what language L we choose when we are generating fundamental priors. Practically speaking, we cannot use a scientific language that cuts nature at the joints, because we have not yet discovered it. If this was merely a practical concern, we could try to say that this doesn’t matter: the fundamental priors are ones that we ought to have rather than any that we actually have or could have—perhaps ought does not imply can. But the concern is not merely practical. For one of the main points of our inductive reasoning is to discover what concepts cut nature at the joints, and this is largely an empirical enterprise. If the right fundamental priors were to reflect the joints in nature, then the enterprise wouldn’t make much sense, as we would be obligated to have already completed much of the enterprise before we started it.

So, I think, we have to say L does not always cut nature at the joints, and yet this generating appropriate priors for us. But we still need a constraint on L. After all, we could imagine a language that thwarts our empirical enterprise, such as one where only fairies can be described briefly and anything else requires very long descriptions, so we have very high priors for fairies and very low priors for everything else. What will be the constraint? Practically, we pretty much have to start with some ordinary human language. I think our ideal should not be far from what is practical. Thus, I propose, if we are going to go with algorithmic priors, we should choose L to be a language that fits well with our human nature as communicators. This is an anthropocentric choice, and I think human epistemology is rightly anthropocentric.

But why think that the anthropocentric choice is apt to lead to truth? There are two stories to be told here. First, it may be that human nature requires a measure of trust in human nature. Second, that trust is vindicated if we are created by a good God who loves the truth.

Monday, October 20, 2025

Another infinite dice game

Suppose infinitely many people independently roll a fair die. Before they get to see the result, they will need to guess whether the die shows a six or a non-six. If they guess right, they get a cookie; if they guess wrong, an electric shock.

But here’s another part of the story. An angel has considered all possible sequences of fair die outcomes for the infinitely many people, and defined the equivalence relation ∼ on the sequences, where α ∼ β if and only if the sequences α and β differ in at most finitely many places. Furthermore, the angel has chosen a set T that contains exactly one sequence from each ∼-equivalence class. Before anybody guesses, the angel is going to look at everyone’s dice and announce the unique member α of T that is -equivalent to the actual die rolls.

Consider two strategies:

  1. Ignore what the angel says and say “not six” regardless.

  2. Guess in accordance with the unique member α: if α says you have six, you guess “six”, and otherwise you guess “not six”.

When the two strategies disagree for a person, there is a good argument that the person should go with strategy (1). For without the information from the angel, the person should go with strategy (1). But the information received from the angel is irrelevant to each individual x, because which -equivalence class the actual sequence of rolls falls into depends only on rolls other than x’s. And following strategy (1) in repeats of the game results in one getting a cookie five out of six times on average.

However, if everyone follows strategy (2), then it is guaranteed that in each game only finitely many people get a shock and everyone else gets a cookie.

This seems to be an interesting case where self-interest gets everyone to go for strategy (1), but everyone going for strategy (2) is better for the common good. There are, of course, many such games, such as Tragedy of the Commons or the Prisoner’s Dilemma, but what is weird about the present game is that there is no interaction between the players—each one’s payoff is independent of what any of the other players do.

(This is a variant of a game in my infinity book, but the difference is that the game in my infinity book only worked assuming a certain rare event happened, while this game works more generally.)

My official line on games like this is that their paradoxicality is evidence for causal finitism, which thesis rules them out.

Monday, September 8, 2025

Observations and risk of confirmation/disconfirmation

It seems that a rational agent cannot guarantee their credence in a hypothesis H to go up by choosing what observation to perform. For if no matter what I observe, my credence in H goes up given my observation, then my credence should already have gone up prior to the observation—I should boost my credence from the armchair.

But this reasoning is false in general. For in performing the observation, I not only learn which of the possible observable results is in place, but I also learn that I have performed the observation. In cases where the truth of H has a correlation with whether I actually perform the observation, this can have a predictable direction of effect on my credence in H.

Suppose that the hypothesis H is the conjunction that I am going to look in the closet and there is life on Mars. By looking to check if there is a mouse in my closet, I ensure that the first conjunct of H is true, and hence I increase my credence in H—no matter what I find out about mice.

This is a very trivial fact. But it does mean that we need to qualify the statement that any observation that can confirm a hypothesis can also disconfirm it. We need to specify that the confirmation and disconfirmation happen after one has already updated on the fact that one has performed the observation.

Friday, February 21, 2025

Adding or averaging epistemic utilities?

Suppose for simplicity that everyone is a good Bayesian and has the same priors for a hypothesis H, and also the same epistemic interests with respect to H. I now observe some evidence E relevant to H. My credence now diverges from everyone else’s, because I have new evidence. Suppose I could share this evidence with everyone. It seems obvious that if epistemic considerations are the only ones, I should share the evidence. (If the priors are not equal, then considerations in my previous post might lead me to withhold information, if I am willing to embrace epistemic paternalism.)

Besides the obvious value of revealing the truth, here are two ways to reason for this highly intuitive conclusion.

First, good Bayesians will always expect to benefit from more evidence. If my place and that of some other agent, say Alice, were switched, I’d want the information regarding E to be released. So by the Golden Rule, I should release the information.

Second, good Bayesians’ epistemic utilities are measured by a strictly proper scoring rule. But if Alice’s epistemic utilities for H are measured by a strictly proper (accuracy) scoring rule s that assigns an epistemic utility s(p,t) to a credence p when the actual truth value of H is t, which can be zero or one. By definition of strict propriety, the expectation by my lights of what Alice’s epistemic utility for a given credence should be is strictly maximized when that credence equals my credence. Since Alice shares the priors I had before I observed E, if I can make E evident to her, her new posteriors will match my current ones, and so revealing E to her will maximize my expectation of her epistemic utility.

So far so good. But now suppose that the hypothesis H = HN is that there exist N people other than me, and my priors assign probability 1/2 to there being N and 1/2 to its being n, where N is much larger than n. Suppose further that my evidence E ends up significantly supporting hypothesis Hn, so that my posterior p in HN is smaller than 1/2.

Now, my expectation of the total epistemic utility of other people if I reveal E is:

  • UR = pNs(p,1) + (1−p)ns(p,0).

And if I conceal E, my expectation is:

  • UC = pNs(1/2,1) + (1−p)ns(1/2,0).

If we had N = n, then it would be guaranteed by strict propriety that UR > UC, and so I should reveal. But we have N > n. Moreover, s(1/2,1) > s(p,1): if some hypothesis is true, a strictly proper accuracy scoring rule increases strictly monotonically with the credence. If N/n is sufficiently large, the first terms of UR and UC will dominate, and hence we will have UC > UR, and thus I should conceal.

The intuition behind this technical argument is this. If I reveal the evidence, I decrease people’s credence in HN. If it turns out that the number of people other than me actually is N, I have done a lot of harm, because I have decreased the credence of a very large number N of people. Since N is much larger than n, this consideration trumps considerations of what happens if the number of people is n.

I take it that this is the wrong conclusion. On epistemic grounds, if everyone’s priors are equal, we should release evidence. (See my previous post for what happens if priors are not equal.)

So what should we do? Well, one option is to opt for averaging rather than summing of epistemic utilities. But the problem reappears. For suppose that I can only communicate with members of my own local community, and we as a community have equal credence 1/2 for the hypothesis Hn that our local community of n people contains all agents, and credence 1/2 for the hypothesis Hn + N that there is also a number N of agents outside our community much greater than n. Suppose, further, that my priors are such that I am certain that all the agents outside our community know the truth about these hypotheses. I receive a piece of evidence E disfavoring Hn and leading to credence p < 1/2. Since my revelation of E only affects the members of my own commmunity, depending on which hypothesis is true, if p is my credence after updating on E, the relevant part of the expected contribution to the utility of revealing E with regard to hypothesis Hn is:

  • UR = p((n−1)/n)s(p,1) + (1−p)((n−1)/(n+N))s(p,0).

And if I conceal E, my expectation contribution is:

  • UC = p((n−1)/n)s(1/2,1) + (1−p)((n−1)/(n+N))s(p,0).

If N is sufficiently large, again UC will beat UR.

I take it that there is something wrong with epistemic utilitarianism.

Tuesday, January 28, 2025

Comparing binary experiments for non-binary questions

In my last two posts (here and here), I introduced the notion of an experiment being epistemically at least as good as another for a set of questions. I then announced a characterization of when this happens in the special case where the set of questions consists of a single binary (yes/no) question and the experiments are themselves binary.

The characterization was as follows. A binary experiment will result in one of two posterior probabilities for the hypothesis that our yes/no question concerns, and we can form the “posterior interval” between them. It turns out that one experiment is at least as good as another provided that the first one’s posterior interval contains the second one’s.

I then noted that I didn’t know what to say for non-binary questions (e.g., “How many mountains are there on Mars?”) but still binary experiments. Well, with a bit of thought, I think I now have it, and it’s almost exactly the same. A binary experiment now defines a “posterior line segment” in the space of probabilities, joining the two possible credence outcomes. (In the case of a probability space with a finite number n of points, the space of probabilities can be identified as the set of points in n-dimensional Euclidean space all of whose coordinates are non-negative and add up to 1.) A bit of thought about convex functions makes it pretty obvious that E2 is at least as good as E1 if and only if E2’s posterior line segment contains E1’s posterior line segment. (The necessity of this geometric condition is easy to see: consider a convex function that is zero everywhere on E2’s posterior line segment but non-zero on one of E1’s two possible posteriors, and use that convex function to generate the scoring rule.)

This is a pretty hard to satisfy condition. The two experiments have to be pretty carefully gerrymandered to make their posterior line segments be parallel, much less to make one a subset of the other. I conclude that when one’s interest is in more than just one binary question, one binary experiment will not be overall better than another except in very special cases.

Recall that my notion of “better” quantified over all proper scoring rules. I guess the upshot of this is that interesting comparisons of scoring rules are not only relative to a set of questions but to a specific proper scoring rule.

Monday, January 20, 2025

Open-mindedness and epistemic thresholds

Fix a proposition p, and let T(r) and F(r) be the utilities of assigning credence r to p when p is true and false, respectively. The utilities here might be epistemic or of some other sort, like prudential, overall human, etc. We can call the pair T and F the score for p.

Say that the score T and F is open-minded provided that expected utility calculations based on T and F can never require you to ignore evidence, assuming that evidence is updated on in a Bayesian way. Assuming the technical condition that there is another logically independent event (else it doesn’t make sense to talk about updating on evidence), this turns out to be equivalent to saying that the function G(r) = rT(r) + (1−r)F(r) is convex. The function G(r) represents your expected value for your utility when your credence is r.

If G is a convex function, then it is continuous on the open interval (0,1). This implies that if one of the functions T or F has a discontinuity somewhere in (0,1), then the other function has a discontinuity at the same location. In particular, the points I made in yesterday’s post about the value of knowledge and anti-knowledge carry through for open-minded and not just proper scoring rules, assuming our technical condition.

Moreover, we can quantify this discontinuity. Given open-mindedness and our technical condiiton, if T has a jump of size δ at credence r (e.g., in the sense that the one-sided limits exist and differ by y), then F has a jump of size rδ/(1−r) at the same point. In particular, if r > 1/2, then if T has a jump of a given size at r, F has a larger jump at r.

I think this gives one some reason to deny that there are epistemically important thresholds strictly between 1/2 and 1, such as the threshold between non-belief and belief, or between non-knowledge and knowledge, even if the location of the thresholds depends on the proposition in question. For if there are such thresholds, then now imagine cases of propositions p with the property that it is very important to reach a threshold if p is true while one’s credence matters very little if p is false. In such a case, T will have a larger jump at the threshold than F, and so we will have a violation of open-mindedness.

Here are three examples of such propositions:

  • There are objective norms

  • God exists

  • I am not a Boltzmann brain.

There are two directions to move from here. The first is to conclude that because open-mindedness is so plausible, we should deny that there are epistemically important thresholds. The second is to say that in the case of such special propositions, open-mindedness is not a requirement.

I wondered initially whether a similar argument doesn’t apply in the absence of discontinuities. Could one have T and F be openminded even though T continuously increases a lot faster than F decreases? The answer is positive. For instance the pair T(r) = e10r and F(r) =  − r is open-minded (though not proper), even though T increases a lot faster than F decreases. (Of course, there are other things to be said against this pair. If that pair is your utility, and you find yourself with credence 1/2, you will increase your expected utility by switching your credence to 1 without any evidence.)

Wednesday, September 4, 2024

Independent invariant regular hyperreal probabilities: an existence result

A couple of years ago I showed how to construct hyperreal finitely additive probabilities on infinite sets that satisfy certain symmetry constraints and have the Bayesian regularity property that every possible outcome has non-zero probability. In this post, I want to show a result that allows one to construct such probabilities for an infinite sequence of independent random variables.

Suppose first we have a group G of symmetries acting on a space Ω. What I previously showed was that there is a hyperreal G-invariant finitely additive probability assignment on all the subsets of Ω that satisfies Bayesian regularity (i.e., P(A) > 0 for every non-empty A) if and only if the action of G on Ω is “locally finite”, i.e.:

  • For any finitely generated subgroup H of G and any point x in G, the orbit Hx is finite.

Here is today’s main result (unless there is a mistake in the proof):

Theorem. For each i in an index set, suppose we have a group Gi acting on a space Ωi. Let Ω = ∏iΩi and G = ∏iGi, and consider G acting componentwise on Ω. Then the following are equivalent:

  1. there is a hyperreal G-invariant finitely additive probability assignment on all the subsets of Ω that satisfies Bayesian regularity and the independence condition that if A1, ..., An are subsets of Ω such that Ai depends only on coordinates from Ji ⊆ I with J1, ..., Jn pairwise disjoint if and only if the action of G on Ω is locally finite

  2. there is a hyperreal G-invariant finitely additive probability assignment on all the subsets of Ω that satisfies Bayesian regularity

  3. the action of G on Ω is locally finite.

Here, an event A depends only on coordinates from a set J just in case there is a subset A′ of j ∈ JΩj such that A = {ω ∈ Ω : ω|J ∈ A′} (I am thinking of the members of a product of sets as functions from the index set to the union of the Ωi). For brevity, I will omit “finitely additive” from now on.

The equivalence of (b) and (c) is from my old result, and the implication from (a) to (b) is trivial, so the only thing to be shown is that (c) implies (a).

Example: If each group Gi is finite and of size at most N for a fixed N, then the local finiteness condition is met. (Each such group can be embedded into the symmetric group SN, and any power of a finite group is locally finite, so a fortiori its action is locally finite.) In particular, if all of the groups Gi are the same and finite, the condition is met. An example like that is where we have an infinite sequence of coin tosses, and the symmetry on each coin toss is the reversal of the coin.

Philosophical note: The above gives us the kind of symmetry we want for each individual independent experiment. But intuitively, if the experiments are identically distributed, we will want invariance with respect to a shuffling of the experiments. We are unlikely to get that, because the shuffling is unlikely to satisfy the local finiteness condition. For instance, for a doubly infinite sequence of coin tosses, we would want invariance with respect to shifting the sequence, and that doesn’t satisfy local finiteness.

Now, on to a sketch of the proof from (c) to (a). The proof uses a sequence of three reductions using an ultraproduct construction to cases exhibiting more and more finiteness.

First, note that without loss of generality, the index set I can be taken to be finite. For if it’s infinite, for any finite partition K of I, and any J ∈ K, let GJ = ∏i ∈ JGi, let ΩJ = ∏i ∈ JΩi, with the obvious action of GJ on ΩJ. Then G is isomorphic to J ∈ KGJ and Ω to J ∈ KΩJ. Then if we have the result for finite index sets, we will get a regular hyperreal G-invariant probability on Ω that satisfies the independence condition in the special case where J1, ..., Jn are such that Ji and Jj for distinct i and j are such that at least one of Ji ∩ J and Jj ∩ J is empty for every J ∈ K. We then take an ultraproduct of these probability measures with respect to K and an ultrafilter on the partially ordered set of finite partitions of I ordered by fineness, and then we get the independence condition in full generality.

Second, without loss of generality, the groups Gi can be taken as finitely generated. For suppose we can construct a regular probability that is invariant under H = ∏iHi where Hi is a finitely generated subgroup of Gi and satisfies the independence condition. Then we take an ultraproduct with respect to an ultrafilter on the partially ordered set of sequences of finitely generated groups (Hi)i ∈ I where Hi is a subgroup of Gi and where the set is ordered by componentwise inclusion.

Third, also without loss of generality, the sets Ωi can be taken to be finite, by replacing each Ωi with an orbit of some finite collection of elements under the action of the finitely generated Gi, since such orbits will be finite by local finiteness, and once again taking an appropriate ultraproduct with respect to an ultrafilter on the partially ordered set of sequences of finite subsets of Ωi closed under Gi ordered by componentwise inclusion. The Bayesian regularity condition will hold for the ultraproduct if it holds for each factor in the ultraproduct.

We have thus reduced everything to the case where I is finite and each Ωi is finite. The existence of the hyperreal G-invariant finitely additive regular probability measure is now trivial: just let P(A) = |A|/|Ω| for every A ⊆ Ω. (In fact, the measure is countably additive and not merely finitely additive, real and not merely hyperreal, and invariant not just under the action of G but under all permutations.)

Monday, August 5, 2024

Natural reasoning vs. Bayesianism

A typical Bayesian update gets one closer to the truth in some respects and further from the truth in other respects. For instance, suppose that you toss a coin and get heads. That gets you much closer to the truth with respect to the hypothesis that you got heads. But it confirms the hypothesis that the coin is double-headed, and this likely takes you away from the truth. Moreover, it confirms the conjunctive hypothesis that you got heads and there are unicorns, which takes you away from the truth (assuming there are no unicorns; if there are unicorns, insert a “not” before “are”). Whether the Bayesian update is on the whole a plus or a minus depends on how important the various propositions are. If for some reason saving humanity hangs on you getting it right whether you got heads and there are unicorns, it may well be that the update is on the whole a harm.

(To see the point in the context of scoring rules, take a weighted Brier score which puts an astronomically higher weight on you got heads and there are unicorns than on all the other propositions taken together. As long as all the weights are positive, the scoring rule will be strictly proper.)

This means that there are logically possible update rules that do better than Bayesian update. (In my example, leaving the probability of the proposition you got heads and there are unicorns unchanged after learning that you got heads is superior, even though it results in inconsistent probabilities. By the domination theorem for strictly proper scoring rules, there is an even better method than that which results in consistent probabilities.)

Imagine that you are designing a robot that maneouvers intelligently around the world. You could make the robot a Bayesian. But you don’t have to. Depending on what the prioritizations among the propositions are, you might give the robot an update rule that’s superior to a Bayesian one. If you have no more information than you endow the robot with, you won’t be able to expect to be able to design such an update rule. (Bayesian update has optimal expected accuracy given the pre-update information.) But if you know a lot more than you tell the robot—and of course you do—you might well be able to.

Imagine now that the robot is smart enough to engage in self-reflection. It then notices an odd thing: sometimes it feels itself pulled to make inferences that do not fit with Bayesian update. It starts to hypothesize that by nature it’s a bad reasoner. Perhaps it tries to change its programming to be more Bayesian. Would it be rational to do that? Or would it be rational for it to stick to its programming, which in fact is superior to Bayesian update? This is a difficult epistemology question.

The same could be true for humans. God and/or evolution could have designed us to update on evidence differently from Bayesian update, and this could be epistemically superior (God certainly has superior knowledge; evolution can “draw on” a myriad of information not available to individual humans). In such a case, switching from our “natural update rule” to Bayesian update would be epistemically harmful—it would take us further from the truth. Moreover, it would be literally unnatural. But what does rationality call on us to do? Does it tell us to do Bayesian update or to go with our special human rational nature?

My “natural law epistemology” says that sticking with what’s natural to us is the rational thing to do. We shouldn’t redesign our nature.