Showing posts with label scoring rules. Show all posts
Showing posts with label scoring rules. Show all posts

Wednesday, June 3, 2026

Epistemic rationality and Pascal's Wager

Pascal’s Wager is an argument that it is prudentially rational to engage in theistic belief promotion practices (TBPP), namely practices apt to promote one’s belief in God.

My interest this post is the standard epistemic rationality objection to the Wager, that engaging in TBPPs is irrational—a kind of brainwashing of oneself. Let’s think about the objection with a bit more care. Consider a specific TBPP Q, say Pascal’s example of going to Mass. If Q is indeed a TBPP, one expects engagement in TBPP to make it more likely that one believes in God. But how does one expect Q to achieve that goal? There are two possibilities. Either Q is expected to achieve that goal by providing one with evidence for theism or in some non-evidential way (or a combination of the two).

Suppose Q is expected to work evidentially. Then we already have the expectation of a higher credence given Q. This expectation is either rational or not. If it is not rational, then we don’t actually have good reason to engage in Q. If it is rational, however, then we should rationally raise our credence in theism right now, without having to bother engaging in Q, and for reasons having nothing to do with any wager. But if Q promotes belief in God non-rationally, then we should not engage in Q for the sake of such promotion of belief—we should not aim to non-rationally promote beliefs.

Let me make the first horn of the dilemma—namely, that the expectation of a higher credence is rational—a bit more precise. We can distinguish two (not mutually exclusive) ways in which a practice rationally increases one’s credence in a hypothesis H. One way is purely epistemic, by uncovering facts about reality. This is the usual way. But if that’s the way we expect to increase our credence in theism by engaging in Q, then we already have evidence that there are such theism-indicating facts to be discovered, and so we should already increase our credence in Q. The other way is practical, by promoting the hypothesis H in a way that shows up to us. The second way is a bit unusual, but here is an example: one way to increase your credence that you will not die of heart disease is to live a healthy life. For if you live a healthy life, you are less likely to die of heart disease, and since you will notice signs of improved cardiac health (e.g., lower resting heart rate, less huffing and puffing on stairs, etc.), your credence that you won’t die of heart diseases will also increase. But this practical way of increasing credence is utterly irrelevant in the case of theism, since nothing we can do can make God more or less likely to exist! So the only rational way that remains is the evidence-based way, and evidence-of-evidence is already evidence.

I used to be quite impressed by the worry that Pascal’s Wager leads to self-deception. I am less impressed. Here is why. There is a serious technical flaw in the argument for the first horn of the dilemma. A simple model for the relation of credence and belief is that you believe a proposition if and only if your credence is above some threshold β. This model might be false, but an analogue of what I will say should apply on more sophisticated models as well.

Here is the point. Consider a case where one is thinking about observing (and suppose this is a simple non-Newcombian observation that does not affect the hypothesis) whether some event E evidentially relevant to H has obtained. Then one’s expected posterior credence is:

  1. C(E)C(HE) + C(∼E)C(H∣∼E),

where C is one’s credence function. But if one is a good Bayesian reasoner, then by total probability the value in (1) is simply equal to one’s prior C(H). Thus the value of one’s credence has no expectation of change upon observation when one is being rational. This seems to support the idea that if you expect your credence to go up, you should already raise it.

But in fact it’s not so simple. For even though the expected posterior credence equals one’s current credence, it could well be that it is more likely that the expected posterior credence exceeds the threshold β if you make the observation than if you don’t. Indeed, cases are obvious. Suppose the belief threshold β is 0.9, and you tossed a coin out of my sight. Suppose I have a prior credence 0.5 that this coin is fair and a prior credence 0.5 that it is double-headed. Then currently I don’t believe (or disbelieve) that the coin is fair. But if I look at the coin and I see tails, I will believe that it is fair—indeed, I will have posterior credence 1 in its fairness. But if I don’t look at the coin, I am not going to get any evidence, and I will continue not to believe that the coin is fair. If I look at the coin, the probability that I will see tails is 0.25 (I have credence 0.5 that it’s fair, and if it’s fair, the chance of tails is 0.5), and so the probability that I will believe if I look is 0.25 (since if I look and see tails, my posterior will be 1 which is bigger than β = 0.9), and the probability that I will believe if I don’t look is 0. Of course, if I don’t see tails, my credence that the coin is fair will go down. But while it will go down, it won’t affect whether I believe that the coin is fair—for I already don’t believe it (in the sense of not-believe, rather than in the sense of believe-not). And there is no irrationality of any sort in looking at the coin in this case.

In other words, the point is that while one’s rational credence has no positive or negative rational expectation of change upon observation, whether one’s rational credence is above a threshold certainly can have a positive or negative rational expection of change.

How could this work in a Pascal’s Wager situation? Let’s talk through one possibility. Take Pascal’s example of going regularly to Mass. Suppose, as Pascal says, your current credence in God is 0.5. You might think that if God exists, going to Mass has a decent chance, say 0.2, of resulting in an evident radical transformation E of your life, so evident and radical that updating your credence in theism on E will push your credence in God to above the threshold β. Of course, you might go to Mass it might not produce any such evident radical transformation (this is true even if God always improves the hearts of people who go to Mass, since he might do so more gradually or less evidently), and in that case your rational credence in God will go below 0.5. But going below 0.5 won’t affect whether you believe in God, since 0.5 is already, I assume, far below the belief threshold β. On the other hand, maybe if you don’t go to Mass, the probability that you will get evidence that will push your credence in theism above β is pretty small—smaller than 0.2. Very likely, your credence will just oscillate a little around 0.5 in the non-Mass-going case. Thus if there is a payoff you get for your credence exceeding the threshold β, it will be worth going to Mass, without there being any epistemic irrationality in the reasoning.

Thought of in this way, we get some practical guidance as to which TBPPs the agnostic or atheist should engage in. They should look for practices that, if God exists, have a decent chance of producing evidence for theism sufficient to push them above β.

I am a bit doubtful that Pascal meant us to think in the above way. He may well have been recommending TBPPs on the grounds of their non-rational effect on belief. My argument above does not defend that.

Monday, April 20, 2026

Lifetime epistemic value

Suppose I discover some fact that I never end up using for anything, or even occurrently thinking about after the discovery. Now, knowledge is good. If I learn the fact earlier in life, then I will have had the knowledge for a longer period of time. So is it better for me to have learned the fact earlier in life?

I doubt it. Consider two scenarios. On the first, I learn what the capital of Zambia is just before I enter a ten-year coma. On the second, I learn it right right after I exit the coma. Learning it before the coma gives me ten more years of knowing it. But that seems a worthless gain. I conclude that in the case of non-occurrent knowing, it doesn’t matter much how long I know.

What about for occurrent knowledge? Other things being equal, if I learn some fact earlier in life, I will occurrently know the fact more times. Is that valuable?

I am less sure. But consider a daily ritual where every morning after waking up, before I am capable of any serious intellectual activity, I think to myself: Sheep have four legs. Thereby, I greatly increase the number of instances in which that piece of knowledge is being occurrently known. Again, this doesn’t seem to be worth the bother.

So it seems that neither for non-occurrent or occurrent knowledge is there non-instrumental value in knowing the thing for a longer period of time. Of course, there typically is instrumental value in knowing something for a longer period of time, both instrumental epistemic value—you can use it in your intellectual investigations of more things—and often instrumental pragmatic value.

This suggests the following. If an agent never loses knowledge, then the lifetime non-instrumental value of their knowledge depends on what they have come to know, not on when they have come to know it. The analogous thesis for perfect Bayesian agents and scoring rules is that their lifetime epistemic utility is the epistemic accuracy score at the latest point in their lives. (If we apply this to Sleeping Beauty, we are apt to get halving. But we shouldn’t apply this to Sleeping Beauty, as she forgets about her first wakeup.)

Things are more complicated in the case of agents who do lose knowledge, whether to memory loss, irrationality or misleading evidence. If we count such an agent’s lifetime non-instrumental epistemic value based on all that they have ever known, that means that if they lost knowledge of p, there is no gain to them from getting it back. But obviously they are better off epistemically if they do get it back. Things get messy and complicated now. A short-period loss in old age doesn’t seem as bad as a case where you found out something early in life and then didn’t have it for the rest of your life.

This is getting messy.

The epistemic value of experiments

You perform an experiment and are going to rationally update on its results. It seems that you should expect this to be good for your epistemic utility as compared to non-performance of the experiment.

Not always! Silly case: Your boss has tasked you with performing a boring chemistry experiment. If you do the experiment, you will find out very little. But if you don’t do it, you will find out a lot about the range of swear words that your boss knows.

What makes this case silly is that you should really think of it as a choice of which experiment to perform, one in chemistry or one in psychology, and in this case the psychology experiment is the more interesting one.

So if we want to say that an experiment can be expected to improve your epistemic utility, we need to be a bit more careful. We need to ensure that non-performance of the experiment doesn’t itself generate information.

But it always does. At the very least, non-performance of the experiment generates the information that the experiment has not been performed by you. You find out something about yourself, and that might far outweigh the value of anything you find out from the experiment. Granted, you also find out something about yourself by performance of the experiment, but it is easy to imagine cases where what you find out by non-performance is more significant. For instance, it could be that your refusal to perform the experiment shows that you have a very specific and rare personality type, while your performance of the experiment gives you nothing so specific.

Suppose, for instance, that you score your epistemic utility by bits of information. The experiment consists in bending down to see which side an unusual coin lying on the ground is facing—that’s one bit of information. Your prior probability that you will look at the coin is 3/4: you are the sort of person who tends to look. So by looking at the coin, you will gain 1 − log2(3/4) = 1.4 bits, mostly regarding the coin but also a little bit about yourself. By not looking at the coin, you will gain 0 − log2(1/4) = 4 bits, all about yourself. Better not to look!

Of course, there are Newcomb-like issues here.

Lesson: The principle that performing a non-trivial experiment should be expected to improve epistemic utility is going to be difficult to formulate.

Thursday, April 9, 2026

Epistemic utilities and death

In the previous post, I proved that we get a proper scoring rule if we compute epistemic utilities as follows. We start with our current credence assignment, consider what credence assignment we will have in the future after we update on some further evidence, and then score that. I then suggested that one could get a lifetime epistemic utility by adding up the epistemic utilities over all the moments of life, and as long as death wasn’t random—as long as the lifespan was fixed—this would generate a proper scoring rule. I then said that if death is random (as it is) it might be the case that you don’t get a proper scoring rule.

My conjecture was wrong. You still get a proper scoring rule despite random death. It’s easy to see this. The basic idea is this. Suppose that you might die the next moment. This partitions the probability space into two subsets, D and L, for death and life. Your current credence is p. Next moment, on D, your credence either doesn’t exist (because you don’t exist or you exist in some supernatural state where you don’t have credences) or doesn’t count (because I am after lifetime credences). Thus, the appropriate way to do a forward-looking scoring of your credence p is to score it s(pL) on L and 0 on D, where pL is the result of conditionalizing your credence on the evidence L (after all, if you are alive, you will conditionalize on being alive), and s is some proper scoring rule. In other words, your forward looking score is sL(p) = 1L ⋅ s(pL).

Is this score proper? Yes! For by propriety of s we have:

  • EpL(s(pL)) ≥ EpL(s(qL)).

But this is the same as:

  • (p(L))−1Ep(1Ls(pL)) ≥ (p(L))−1Ep(1Ls(qL)).

Multiplying both sides by p(L) (I am assuming a non-zero probability of survival), we get:

  • Ep(sL(p)) ≥ Ep(sL(q)).

We can combine this with a more complex set of future investigations as in the previous post, and things will still work.

It is crucial to the above argument that when you’re alive, you can tell you’re alive. I suppose that’s not always true. When you’re asleep, you are alive, but can’t tell you’re alive. So to generalize beyond the above toy example, replace death with unconsciousness or something like that.

Forward-looking scoring rules

An accuracy scoring rule assigns a score to a probability function representing an agent’s credences, ostensibly measuring how close that probability function is to the truth. The score s(p) of a probability function p is a random variable, because the value of the score depends on what is actually true, i.e., on where we are in the probability space.

A proper scoring rule (on probabilistic credences) satisfies the propriety inequality

  1. Eps(p) ≥ Eps(q)

which says that the expected score of your current lights—your current credences p—by your current lights is optimal: you won’t improve your expected score (by your current lights) by switching to a different credence q.

You can think of a proper scoring rule as representing the epistemic utility of having a credence p.

But now let’s think about things dynamically. In the future, you will receive additional evidence. As a good Bayesian agent, you will update on this evidence by conditionalization. Perhaps instead of thinking about maximizing your current score, you should think about maximizing your future score. Maybe your true epistemic utility is the score you will end up with after all the future evidence is in.

A simple model of this is as follows. There is some finite partition I = (I1,...,In) of your probability space Ω with each cell Ii of the partition representing a possibility for what you might learn given future evidence. Your current credence function is p, and p(Ii) > 0 for all i. There is then a random credence function pI where pI(ω) is the credence function you will have once the evidene is in if you are at ω ∈ Ω. In other words, pI(ω)(A) = p(AIi) where Ii is the member of the partition that contains ω. (Technically, the function that maps ω to pI(ω)(A) is equal to the conditional probability p(AG) where G is the algebra generated by I.)

Now, given a proper scoring rule s, define a new scoring rule sI as follows:

  1. sI(p)(ω) = s(pI(ω))(ω).

Your sI-score for p at ω then represents the score you will have at ω once you learn which cell of the partition I you are in.

Theorem: The scoring rule sI is proper if s is proper.

Note that sI won’t be strictly proper (i.e., (1) won’t always have strict inequality when p and q are distinct) if I has two or more cells, because pI and qI are going to be the same if p and q assign different probabilities to the cells, but have the same conditional probabilities on each cell. But it might still be the case that sI is strictly proper with respect to some relevant subfield of Ω—that needs some further investigation.

Suppose now you are a Bayesian agent who is guaranteed to consciously live for n moments. In each moment, new information comes in. Thus, we have a sequence J0, ..., Jn of finer and finer partitions, with J0 being the trivial partition, and with pJk representing the credence you will have at time k. Your overall epistemic lifetime score is then:

  1. sΣ(p) = ∑ksJk(p).

It follows from the Theorem that sΣ is a proper scoring rule if s is. And if J0 is the trivial partition, then sJ0 = s, and so if s is strictly proper, then the lifetime score sΣ is strictly proper, since the sum of a strictly proper rule and a proper rule is strictly proper. So, lifetime scores are strictly proper if they are constructed from an instantaneous score—in the above toy model.

Alas, the toy model is not fully adequate, because it is random when we will die, and so our lifespan doesn’t have a fixed sequence of moments. Once we take into account the randomness of when we will die, the overall epistemic lifetime score might stop being proper: this needs further investigation.

Proof of Theorem: By the Greaves and Wallace Theorem, an optimal method of updating credences with respect to expected proper score is by Bayesian conditionalization. Apply the Greaves and Wallace Theorem to the scoring rule s and the starting credence p with the following two strategies:

A. Bayesian conditionalization on the true cell of I.

B. Switch your credence from p to q, then apply Bayesian conditionalization on the true cell of I.

Saying that (A) is at least as good as (B) is equivalent to the the proper scoring rule inequality (1) for sI.

Thursday, February 5, 2026

More on strong open-mindedness

For the last couple of days I have been exploring what I like to call strongly open-minded accuracy scoring rules. It’s well known that every proper scoring rule is open-minded in the sense that it never requires you to reject free information: the expected epistemic utility of updating on the free information is always at least as good as your current expected epistemic utility. It’s strictly open-minded provided that in non-trivial cases (i.e., when the information has a non-zero probability of having statistical relevance to the credences you are scoring) you are required to accept the free information.

Now there are two reasons why one might accept free information about some proposition q. First, you might be wrong about q: your credence may be high while q is false or your credence might be low while q is true. Second, even if you are right about q, the free information may boost your credence in the right direction. I say that a scoring rule is strongly open-minded provided that it licenses you to accept and update on the free information even if you disregard the first consideration. We can then tack on “strictly” if it requires you to do so in non-trivial cases. In the case of a strongly open-minded scoring rule, your acceptance of free information is not a sign of doubt in your propositions—it is not a way of hedging your bets—and thus is arguably compatible with faith in the propositions being evaluated.

A strongly open-minded scoring rule can also be characterized in the following way. There is a more ordinary kind of epistemic paternalism where I might have reason to block another from receiving free information on the grounds that this information could mislead due to the fact that the other has different likelihoods from the ones I think are right. For instance, if too many people have an unjustified mistrust of Dr. Smith such that they are likely to believe the opposite of what Dr. Smith’s experiments reveal, there is reason to give a grant to someone else, because Dr. Smith’s experiments are likely to lead people away from the truth, for no fault of Dr. Smith’s. Call this likelihood-based paternalism. But there is another kind of motivation of the refusal of free information for another, which we might call pure-risk-based paternalism. Even if someone else has the same likelihoods as you do—trusts Dr. Smith just as you do—perhaps the risk that Dr. Smith’s experiments will, by pure chance, provide evidence away from the truth is enough to justify not funding these experiments.

I’ve been collecting results about these issues. Here’s what I seem t have so far, though I have to emphasize that sometimes the proofs are just in my head and I might be wrong. I will specialize on scoring rules for a single proposition, given as a pair of functions T and F, where T(x) is the value of having credence x when the proposition is true and F(x) is the value of having credence x when the proposition is false.

  1. A scoring rule sometimes calls for pure-risk-based paternalism if and only if it is not strongly open-minded.

  2. A scoring rule that’s strongly open-minded is open-minded.

  3. A scoring rule (T,F) is (strictly) strongly open-minded if and only if xT(x) and (1−x)F(1−x) are both (strictly) convex.

  4. The logarithmic scoring rule is strictly strongly open-minded. The Brier and spherical rules are not strongly open-minded.

  5. If a proper scoring rule is generated by the Schervisch-style integral representation T(x) = T(1/2) + ∫x1/2(1−t)b(t)dt and F(x) = F(1/2) + ∫1/2xtb(t)dt and b is sufficiently differentiable, then the scoring rule is strongly open-minded if and only if the derivative of log b(x) lies between (3x−2)/[x(1−x)] and (3x−1)/[x(1−x)].

  6. A strongly open-minded scoring rule whose logarithm is sufficiently differentiable is unbounded.

  7. [Item deleted as I discovered it to be false.]

  8. For any credences p and r such that 1/2 < r and p < r, there is a strictly proper scoring rule and a situation where the scoring rule calls for the individual with credence r to have purely-risk-based epistemic paternalism for that hypothesis.

Monday, September 29, 2025

Lying and epistemic utility

Epistemic utility is the value of one’s beliefs or credences matching the truth.

Suppose your and my credences differ. Then I am going to think that my credences better match the truth. This is automatic if I am measuring epistemic utilities using a proper scoring rule. But that means that benevolence with respect to epistemic utilities gives me a reason to shift your credences to be closer to mine.

At this point, there are honest and dishonest ways to proceed. The honest way is to share all my relevant evidence with you. Suppose I have done that. And you’ve reciprocated. And we still differ in credences. If we’re rational Bayesian agents, that’s presumably due to a difference in prior probabilities. What can I do, then, if the honest ways are exhausted?

I can lie! Suppose your credence that there was once life on Mars is 0.4 and mine is 0.5. So I tell you that I read that a recent experiment provided a little bit of evidence in favor of there once having been life on Mars, even though I read no such thing. That boosts your credence that there was once life on Mars. (Granted, it also boosts your credence in the falsehood that there was such a recent experiment. But, plausibly, getting right whether there was once life on Mars gets much more weight in a reasonable person’s epistemic utilities than getting right what recent experiments have found.)

We often think of lying as an offense against truth. But in these kinds of cases, the lies are aimed precisely at moving the other towards truth. And they’re still wrong.

Thus, it seems that striving to maximize others’ epistemic utility is the wrong way to think of our shared epistemic life.

Maximizing others’ epistemic utility seems to lead to a really bad picture of our shared epistemic life. Should we, then, think of striving to maximize our own epistemic utility as the right approach to one’s individual epistemic life? Perhaps. For maybe what is apt to go wrong in maximizing others’ epistemic utility is paternalism, and paternalism is rarely a problem in one’s own case.

Thursday, September 11, 2025

Why do we like being confident?

We like being more confident. We enjoy having credences closer to 0 or 1. Even if the proposition we are confident in is one that is such that it is a bad thing that it is true, the confidence itself, abstracted from the badness of the state of affairs reported by the proposition, is something we enjoy.

Here is a potential justification of this attitude in many cases. We can think of the epistemic utility of one’s credence r in a proposition p as measured by an accuracy scoring rule given by two functions T(r) and F(r), where T(r) gives the value of having credence r in p when p is actually true and F(r) gives the value when p is actually false. Most people thinking about scoring rules think they should satisfy the technical condition of being strictly proper. But strict propriety implies that the function V(r) = rT(r) + (1−r)F(r) is strictly convex. Now suppose the scoring rule is also symmetric, so that T(r) = F(1−r). Then V(r) is a strictly convex function that is symmetric about r = 1/2. Such a function has its minimum at r = 1/2, and is strictly decreasing on [0,1/2] and strictly increasing on [1/2,1]. But the function V(r) measures your expectation of your epistemic utility. How happy you are about your credence, perhaps, corresponds to your expectation of your epistemic utility. So you are most unhappy at credence 1/2, and you get happier that closer you are to 0 or 1.

OK, it’s surely not that?!

Monday, September 8, 2025

Epistemic utilities and decision theories

Warning: I worry there may be something wrong in the reasoning below.

Causal Decision Theory (CDT) and Epistemic Decision Theory (EDT) tend to disagree when the payoff of an option statistically depends on your propensity to go for that option. The most example of this phenomenon is Newcomb’s Problem (where money is literally put into a box or not depending on what your propensities are), and there is a large literature of other clever and mind-twisting examples. From the literature, one might get a feeling that these cases are all somehow weird, and normally there is no such dependence.

But here is a family of cases that happens literally almost all the time to us. Pretty much whenever we act we gain information relevant to facts about ourselves, and specifically to facts about our propensities to act. For instance, when you choose chocolate over vanilla ice cream you raise your credence for the hypothesis that you have a greater propensity to choose chocolate ice cream than to choose vanilla ice cream. But truth about oneself is valuable and falsehood about oneself is disvaluable. If in fact you have a greater propensity to choose chocolate ice cream, then by eating chocolate ice cream you gain credence in a truth, which is a good thing. If in fact your propensity for vanilla ice cream is at least as great as for chocolate ice cream, then by eating chocolate ice cream, you gain credence in a falsehood. The payoffs of your decision as to flavor of ice cream thus statistically depend on what your propensities actually are, and so this is exactly the kind of case where we would expect CDT and EDT to disagree.

Let’s be more precise. You have a choice between eating chocolate ice cream (C), eating vanilla ice cream (V) or not eating ice cream at all (N). Let H be the hypothesis that you have a greater propensity for eating chocolate ice cream than for eating vanilla ice cream. Then if you choose C, you will gain evidence for H. If you choose V, you will gain evidence for not-H. And if you choose N, you will (plausibly) gain no evidence for or against H. Your epistemic utility with respect to H is, let us suppose, measured by a single-proposition accuracy scoring rule, which we can think of as a pair of functions TH and FH, where TH(p) is the value of having credence p in H if in fact H is true and FH(p) is the value of having credence p in H if in fact H is false.

The expected evidential utilities of your three options are:

  • Ee(C) = P(H|C)TH(P(H|C)) + (1−P(H|C))FH(P(H|C))

  • Ee(V) = P(H|V)TH(P(H|V)) + (1−P(H|V))FH(P(H|V))

  • Ee(N) = P(H|N)TH(P(H|N)) + (1−P(H|N))FH(P(H|N)) = P(H)TH(P(H)) + (1−P(H))FH(P(H)).

The expected causal utilities are:

  • Ec(C) = P(H)TH(P(H|C)) + (1−P(H))FH(P(H|C))

  • Ec(V) = P(H)TH(P(H|V)) + (1−P(H))FH(P(H|V))

  • Ec(N) = P(H)TH(P(H|N)) + (1−P(H))FH(P(H|N)) = P(H)TH(P(H)) + (1−P(H))FH(P(H)).

We can make some quick observations in the case where the scoring rule is strictly proper, given that P(H|V) < P(H) < P(H|C):

  1. Ec(C) < Ec(N)

  2. Ec(V) < Ec(N)

  3. At least one of Ee(C) > Ee(N) and Ee(V) > Ee(N) is true.

Observations 1 and 2 follow immediately from strict propriety and the formulas for Ec. Observation 3 follows from the fact that the expected accuracy score after Bayesian update on evidence is better (in non-trivial cases where the scoring rule is strictly proper) than before update, and the expected accuracy score after update on what you’ve chosen is:

  • P(C)Ee(C) + P(V)Ee(V) + P(N)Ee(N)

while the expected accuracy score before update is equal to Ee(N). Since P(C) + P(V) + P(N) = 1, it follows from the superiority of the post-update expectation that at least one of Ee(C) and Ee(V) must be bigger than Ee(N).

The above results seem to be a black eye for CDT, which recommends that if what you care about is your epistemic utility with regard to your propensities regarding chocolate and vanilla ice cream, then you should always avoid eating ice cream!

(What about ratifiability? Some CDTers say that only ratifiable options should count. Is N ratifiable? Given that you’ve learned nothing about H from choosing N, I think N should be ratifiable. But I may be missing something. I find the epistemic utility case confusing.)

It also seems to me (I haven’t checked details) that on EDT there are cases where eating either flavor is good for you epistemically, but there are also cases where only one specific flavor is good for you.

Tuesday, June 3, 2025

Combining epistemic utilities

Suppose that the right way to combine epistemic utilities or scores across individuals is averaging, and I am an epistemic act expected-utility utilitarian—I act for the sake of expected overall epistemic utility. Now suppose I am considering two different hypotheses:

  • Many: There are many epistemic agents (e.g., because I live in a multiverse).

  • Few: There are few epistemic agents (e.g., because I live in a relatively small universe).

If Many is true, given averaging my credence makes very little difference to overall epistemic utility. On Few, my credence makes much more of a difference to overall epistemic utility. So I should have a high credence for Few. For while a high credence for Few will have an unfortunate impact on overall epistemic utility if Many is true, because the impact of my credence on overall epistemic utility will be small on Many, I can largely ignore the Many hypothesis.

In other words, given epistemic act utilitarianism and averaging as a way of combining epistemic utilities, we get a strong epistemic preference for hypotheses with fewer agents. (One can make this precise with strictly proper scoring rules.) This is weird, and does not match any of the standard methods (self-sampling, self-indication, etc.) for accounting for self-locating evidence.

(I should note that I once thought I had a serious objection to the above argument, but I can't remember what it was.)

Here’s another argument against averaging epistemic utilities. It is a live hypothesis that there are infinitely many people. But on averaging, my epistemic utility makes no difference to overall epistemic utility. So I might as well believe anything on that hypothesis.

One might toy with another option. Instead of averaging epistemic utilities, we could average credences across agents, and then calculate the overall epistemic utility by applying a proper scoring rule to the average credence. This has a different problematic result. Given that there are at least billions of agents, for any of the standard scoring rules, as long as the average credence of agents other than you is neither very near zero nor very near one, your own credence’s contribution to overall score will be approximately linear. But it’s not hard to see that then to maximize expected overall epistemic utility, you will typically make your credence extreme, which isn’t right.

If not averaging, then what? Summing is the main alternative.

Tuesday, February 25, 2025

Being known

The obvious analysis of “p is known” is:

  1. There is someone who knows p.

But this obvious analysis doesn’t seem correct, or at least there is an interesting use of “is known” that doesn’t fit (1). Imagine a mathematics paper that says: “The necessary and sufficient conditions for q are known (Smith, 1967).” But what if the conditions are long and complicated, so that no one can keep them all in mind? What if no one who read Smith’s 1967 paper remembers all the conditions? Then no one knows the conditions, even though it is still true that the conditions “are known”.

Thus, (1) is not necessary for a proposition to be known. Nor is this a rare case. I expect that more than half of the mathematics articles from half a century ago contain some theorem or at least lemma that is known but which no one knows any more.

I suspect that (1) is not sufficient either. Suppose Alice is dying of thirst on a desert island. Someone, namely Alice, knows that she is dying of thirst, but it doesn’t seem right to say that it is known that she is dying of thirst.

So if it is neither necessary nor sufficient for p to be known that someone knows p, what does it mean to say that p is known? Roughly, I think, it has something to do with accessibility. Very roughly:

  1. Somebody has known p, and the knowledge is accessible to anyone who has appropriate skill and time.

It’s really hard to specify the appropriateness condition, however.

Does all this matter?

I suspect so. There is a value to something being known. When we talk of scientists advancing “human knowledge”, it is something like this “being known” that we are talking about.

Imagine that a scientist discovers p. She presents p at a conference where 20 experts learn p from her. Then she publishes it in a journal when 100 more people learn it. Then a Youtuber picks it up and now a million people know it.

If we understand the value of knowledge as something like the sum of epistemic utilities across humankind, then the successive increments in value go like this: first, we have a move from zero to some positive value V when the scientist discovers p. Then at the conference, the value jumps from V to 21V. Then after publication it goes from 21V to 121V. Then given Youtube, it goes from 121V to 100121V. The jump at initial discovery is by far the smallest, and the biggest leap is when the discovery is publicized. This strikes me as wrong. The big leap in value is when p becomes known, which either happens when the scientist discovers it or when it is presented at the conference. The rest is valuable, but not so big in terms of the value of “human knowledge”.

Monday, February 24, 2025

More on averaging to combine epistemic utilities

Suppose that the right way to combine epistemic utilities across people is averaging: the overall epistemic utility of the human race is the average of the individual epistemic utilities. Suppose, further, that each individual epistemic utility is strictly proper, and you’re a “humanitarian” agent who wants to optimize overall epistemic utility.

Suppose you’re now thinking about two hypotheses about how many people exist: the two possible numbers are m and n, which are not equal. All things considered, you have credence 0 < p0 < 1 in the hypothesis Hm that there are m people and 1 − p0 in the hypothesis Hn that there are n people. You now want to optimize overall epistemic utility. On an averaging view, if Hm is true, if your credence is p1, your contribution to overall epistemic utility will be:

  • (1/m)T(p1)

and if Hm is false, your contribution will be:

  • (1/n)F(p1),

where your strictly proper scoring rule is given by T, P. Since your credence is p1, by your lights the expected value after your changing your credence to p0 will be:

  • p0(1/m)T(p1) + (1−p0)(1/n)F(p1) + Q(p0)

where Q(p0) is the contribution of other people’s credences, which I assume you do not affect with your choice of p1. If m ≠ n and T, F is strictly proper, the expected value will be maximized at

  • p1 = (p0/m)/(p0/m+(1−p0)/n) = np0/(np0+m(1−p0)).

If m > n, then p1 < p0 and if m < n, then p1 > p0. In other words, as long as n ≠ m, if you’re an epistemic humanitarian aiming to improve overall epistemic utility, any credence strictly between 0 and 1 will be unstable: you will need to change it. And indeed your credence will converge to 0 if m > n and to 1 if m < n. This is absurd.

I conclude that we shouldn’t combine epistemic utilities across people by averaging the utilities.

Idea: What about combining them by computing the epistemic utilities of the average credences, and then applying a strictly proper scoring rule, in effect imagining that humanity is one big committee and that a committee’s credence is the average of the individual credences?

This is even worse, because it leads to problems even without considering hypotheses on which the number of people varies. Suppose that you’ve just counted some large number nobody cares about, such as the number of cars crossing some intersection in New York City during a specific day. The number you got is even, but because the number is big, you might well have made a mistake, and so your credence that the number is even is still fairly low, say 0.7. The billions of other people on earth all have credence 0.5, and because nobody cares about your count, you won’t be able to inform them of your “study”, and their credences won’t change.

If combined epistemic utility is given by applying a proper scoring rule to the average credence, then by your lights the expected value of the combined epistemic utility will increase the bigger you can budge the average credence, as long as you don’t get it above your credence. Since you can really only affect your own credence, as an epistemic humanitarian your best bet is to set your credence to 1, thereby increasing overall human credence from 0.5 to around 0.5000000001, and making a tiny improvement in the expected value of the combined epistemic utility of humankind. In doing so, you sacrifice your own epistemic good for the epistemic good of the whole. This is absurd!

I think the idea of averaging to produce overall epistemic utilities is just wrong.

Friday, February 21, 2025

Adding or averaging epistemic utilities?

Suppose for simplicity that everyone is a good Bayesian and has the same priors for a hypothesis H, and also the same epistemic interests with respect to H. I now observe some evidence E relevant to H. My credence now diverges from everyone else’s, because I have new evidence. Suppose I could share this evidence with everyone. It seems obvious that if epistemic considerations are the only ones, I should share the evidence. (If the priors are not equal, then considerations in my previous post might lead me to withhold information, if I am willing to embrace epistemic paternalism.)

Besides the obvious value of revealing the truth, here are two ways to reason for this highly intuitive conclusion.

First, good Bayesians will always expect to benefit from more evidence. If my place and that of some other agent, say Alice, were switched, I’d want the information regarding E to be released. So by the Golden Rule, I should release the information.

Second, good Bayesians’ epistemic utilities are measured by a strictly proper scoring rule. But if Alice’s epistemic utilities for H are measured by a strictly proper (accuracy) scoring rule s that assigns an epistemic utility s(p,t) to a credence p when the actual truth value of H is t, which can be zero or one. By definition of strict propriety, the expectation by my lights of what Alice’s epistemic utility for a given credence should be is strictly maximized when that credence equals my credence. Since Alice shares the priors I had before I observed E, if I can make E evident to her, her new posteriors will match my current ones, and so revealing E to her will maximize my expectation of her epistemic utility.

So far so good. But now suppose that the hypothesis H = HN is that there exist N people other than me, and my priors assign probability 1/2 to there being N and 1/2 to its being n, where N is much larger than n. Suppose further that my evidence E ends up significantly supporting hypothesis Hn, so that my posterior p in HN is smaller than 1/2.

Now, my expectation of the total epistemic utility of other people if I reveal E is:

  • UR = pNs(p,1) + (1−p)ns(p,0).

And if I conceal E, my expectation is:

  • UC = pNs(1/2,1) + (1−p)ns(1/2,0).

If we had N = n, then it would be guaranteed by strict propriety that UR > UC, and so I should reveal. But we have N > n. Moreover, s(1/2,1) > s(p,1): if some hypothesis is true, a strictly proper accuracy scoring rule increases strictly monotonically with the credence. If N/n is sufficiently large, the first terms of UR and UC will dominate, and hence we will have UC > UR, and thus I should conceal.

The intuition behind this technical argument is this. If I reveal the evidence, I decrease people’s credence in HN. If it turns out that the number of people other than me actually is N, I have done a lot of harm, because I have decreased the credence of a very large number N of people. Since N is much larger than n, this consideration trumps considerations of what happens if the number of people is n.

I take it that this is the wrong conclusion. On epistemic grounds, if everyone’s priors are equal, we should release evidence. (See my previous post for what happens if priors are not equal.)

So what should we do? Well, one option is to opt for averaging rather than summing of epistemic utilities. But the problem reappears. For suppose that I can only communicate with members of my own local community, and we as a community have equal credence 1/2 for the hypothesis Hn that our local community of n people contains all agents, and credence 1/2 for the hypothesis Hn + N that there is also a number N of agents outside our community much greater than n. Suppose, further, that my priors are such that I am certain that all the agents outside our community know the truth about these hypotheses. I receive a piece of evidence E disfavoring Hn and leading to credence p < 1/2. Since my revelation of E only affects the members of my own commmunity, depending on which hypothesis is true, if p is my credence after updating on E, the relevant part of the expected contribution to the utility of revealing E with regard to hypothesis Hn is:

  • UR = p((n−1)/n)s(p,1) + (1−p)((n−1)/(n+N))s(p,0).

And if I conceal E, my expectation contribution is:

  • UC = p((n−1)/n)s(1/2,1) + (1−p)((n−1)/(n+N))s(p,0).

If N is sufficiently large, again UC will beat UR.

I take it that there is something wrong with epistemic utilitarianism.

Bayesianism and epistemic paternalism

Suppose that your priors for some hypothesis H are 3/4 while my priors for it are 1/2. I now find some piece of evidence E for H which raises my credence in H to 3/4 and would raise yours above 3/4. If my concern is for your epistemic good, should I reveal this evidence E?

Here is an interesting reason for a negative answer. For any strictly proper (accuracy) scoring rule, my expected value for the score of a credence is uniquely maximized when the credence is 3/4. I assume your epistemic utility is governed by a strictly proper scoring rule. So the expected epistemic utility, by my lights, of your credence is maximized when your credence is 3/4. But if I reveal E to you, your credence will go above 3/4. So I shouldn’t reveal it.

This is epistemic paternalism. So, it seems, expected epistemic utility maximization (which I take it has to employ a strictly proper scoring rule) forces one to adopt epistemic paternalism. This is not a happy conclusion for expected epistemic utility maximization.

Wednesday, January 29, 2025

More on experiments

We all perform experiments very often. When I hear a noise and deliberately turn my head, I perform an experiment to find out what I will see if I turn my head. If I ask a question not knowing what answer I will hear, I am engaging in (human!) experimentation. Roughly, experiments are actions done in order to generate observations as evidence.

There are typically differences in rigor between the experiments we perform in daily life and the experiments scientists perform in the lab, but only typically so. Sometimes we are rigorous in ordinary life and sometimes scientists are sloppy.

The epistemic value to one of an experiment depends on multiple factors in a Bayesian framework.

  1. The set of questions towards answers to which the experiment’s results are expected to contribute.

  2. Specifications of the value of different levels of credence regarding the answers to the questions in Factor 1.

  3. One’s prior levels of credence for the answers.

  4. The likelihoods of different experimental outcomes given different answers.

It is easiest to think of Factor 2 in practical terms. If I am thinking of going for a recreational swim but I am not sure whether my swim goggles have sprung a leak, it may be that if the probability of the goggles being sound is at least 50%, it’s worth going to the trouble of heading out for the pool, but otherwise it’s not. So an experiment that could only yield a 45% confidence in the goggles is useless to my decision whether to go to the pool, and there is no difference in value between an experiment that yields a 55% confidence and one that yields a 95% confidence. On the other hand, if I am an astronaut and am considering performing a non-essential extravehicular task, but I am worried that the only available spacesuit might have sprung a leak, an experiment that can only yield 95% confidence in the soundness of the spacesuit is pointless—if my credence in the spacesuit’s soundness is only 95%, I won’t use the spacesuit.

Factor 3 is relevant in combination with Factor 4, because these two factors tell us how likely I am to end up with different posterior probabilities for the answers to the Factor 1 questions after the experiment. For instance, if I saw that one of my goggles is missing its gasket, my prior credence in the goggle’s soundness is so low that even a positive experimental result (say, no water in my eye after submerging my head in the sink) would not give me 50% credence that the goggle is fine, and so the experiment is pointless.

In a series of posts over the last couple of days, I explored the idea of a somewhat interest-independent comparison between the values of experiments, where one still fixes a set of questions (Factor 1), but says that one experiment is at least as good as another provided that it has at least as good an expected epistemic utility as the other for every proper scoring rule (Factor 2). This comparison criterion is equivalent to one that goes back to the 1950s. This is somewhat interest-independent, because it is still relativized to a set of questions.

A somewhat interesting question that occurred to me yesterday is what effect Factor 3 has on this somewhat interest-independent comparison of experiments. If experiment E2 is at least as good as experiment E1 for every scoring rule on the question algebra, is this true regardless of which consistent and regular priors one has on the question algebra?

A bit of thought showed me a somewhat interesting fact. If there is only one binary (yes/no) question under Factor 1, then it turns out that the somewhat interest-independent comparison of experiments does not depend on the prior probability for the answer to this question (assuming it’s regular, i.e., neither 0 nor 1). But if the question algebra is any larger, this is no longer true. Now, whether an experiment is at least as good as another in this somewhat interest-independent way depends on the choice of priors in Factor 3.

We might now ask: Under what circumstances is an experiment at least as good as another for every proper scoring rule and every consistent and regular assignment of priors on the answers, assuming the question algebra has more than two non-trivial members? I suspect this is a non-trivial question.

Tuesday, January 28, 2025

And one more post on comparing experiments

In my last couple of posts, starting here, I’ve been thinking about comparing the epistemic quality of experiments for a set of questions. I gave a complete geometric characterization for the case where the experiments are binary—each experiment has only two possible outcomes.

Now I want to finally note that there is a literature for the relevant concepts, and it gives a characterization of the comparison of the epistemic quality of experiments, at least in the case of a finite probability space (and in some infinite cases).

Suppose that Ω is our probability space with a finite number of points, and that FQ is the algebra of subsets of Ω corresponding to the set of questions Q (a question partitions Ω into subsets and asks which partition we live in; the algebra FQ is generated by all these partitions). Let X be the space of all probability measures on FQ. This can be identified with an (n−1)-dimensional subset of Euclidean Rn consisting of the points with non-negative coordinates summing to one, where n is the number of atoms in FQ. An experiment E also corresponds to a partition of Ω—it answers the question where in that partition we live. The experiment has some finite number of possible outcomes A1, ..., Am, and in each outcome Ai our Bayesian agent will have a different posterior PAi = P(⋅∣Ai). The posteriors are members of X. The experiment defines an atomic measure μE on X where μE(ν) is the probability that E will generate an outcome whose posterior matches ν on FQ. Thus:

  • μE(ν) = P(⋃{Ai:PAi|FQ=ν}).

Given the correspondence between convex functions and proper scoring rules, we can see that experiment E2 is at least as good as E1 for Q just in case for every convex function c on X we have:

  • XcdμE2 ≥ ∫XcdμE1.

There is an accepted name for this relation: μE2 convexly dominates μE1. Thus, we have it that experiment E2 is at least as good as experiment E1 for Q provided that there is a convex domination relation between the distributions the experiments induce on the possible posteriors for the questions in Q. And it turns out that there is a known mathematical characterization of when this happens, and it includes some infinite cases as well.

In fact, the work on this epistemic comparison of experiments turns out to go back to a 1953 paper by Blackwell. The only difference is that Blackwell (following 1950 work by Bohnenblust, Karlin and Sherman) uses non-epistemic utility while my focus is on scoring rules and epistemic utility. But the mathematics is the same, given that non-epistemic decision problems correspond to proper scoring rules and vice versa.

Monday, January 27, 2025

Comparing binary experiments for binary questions

In my previous post I introduced the notion of an experiment being better than another experiment for a set of questions, and gave a definition in terms of strictly proper (or strictly open-minded, which yields the same definition) scoring rules. I gave a sufficient condition for E2 to be at least as good as E1: E2’s associated partition is essentially at least as fine as that of E1.

I then ended with an open question as to what the necessary and sufficient conditions for a binary (yes/no) experiment to be at least as good as another binary one for a binary question.

I think I now have an answer. For a binary experiment E and a hypothesis H, say that E’s posterior interval for H is the closed interval joining P(HE) with P(H∣∼E). Then, I think:

  • Given the binary question whether a hypothesis H is true, and binary experiments E1 and E2, experiment E2 is at least as good as E1 if and only if its posterior interval for H contains the E1’s posterior interval for H.

Let’s imagine that you want to be confident of H, because H is nice. Then the above condition says that an experiment that’s better than another will have at least as big potential benefit (i.e., confidence in H) and at least as big potential risk (i.e., confidence in  ∼ H). No benefits without risks in the epistemic game!

The proof (which I only have a sketch of) follows from expressing the expected score after an experiment using formula (4) here, and using convexity considerations.

The above answer doesn’t work for non-binary experiments. The natural analogue to the posterior interval is the convex hull of the set of possible posteriors. But now imagine two experiments to determine whether a coin is fair or double-headed. The first experiment just tosses the coin and looks at the answer. The second experiment tosses an auxiliary independent and fair coin, and if that one comes out heads, then the coin that we are interested in is tossed. The second experiment is worse, because there is probability 1/2 that the auxiliary coin is tails in which case we get no information. But the posterior interval is the same for both experiments.

I don’t know what to say about binary experiments and non-binary questions. A necessary condition is containment of posterior intervals for all possible answers to the question. I don’t know if that’s sufficient.

Comparing experiments

When you’re investigating reality as a scientist (and often as an ordinary person) you perform experiments. Epistemologists and philosophers of science have spent a lot of time thinking about how to evaluate what you should do with the results of the experiments—how they should affect your beliefs or credences—but relatively little on the important question of which experiments you should perform epistemologically speaking. (Of course, ethicists have spent a good deal of time thinking about which experiments you should not perform morally speaking.) Here I understand “experiment” in a broad sense that includes such things as pulling out a telescope and looking in a particular direction.

One might think there is not much to say. After all, it all depends on messy questions of research priorities and costs of time and material. But we can at least abstract from the costs and quantify over epistemically reasonable research priorities, and define:

  1. E2 is epistemically at least as good an experiment as E1 provided that for every epistemically reasonable research priority, E2 would serve the priority at least as well as E1 would.

That’s not quite right, however. For we don’t know how well an experiment would serve a research priority unless we know the result of the experiment. So a better version is:

  1. E2 is epistemically at least as good an experiment as E1 provided that for every epistemically reasonable research priority, the expected degree to which E2 would serve the priority is at least as high as the expected degree to which E1 would.

Now we have a question we can address formally.

Let’s try.

  1. A reasonable epistemic research priority is a strictly proper scoring rule or epistemic utility, and the expected degree to which an experiment would serve that priority is equal to the expected value of the score after Bayesian update on the result of the experiment.

(Since we’re only interested in expected values of scores, we can replace “strictly proper” with “strictly open-minded”.)

And we can identify an experiment with a partition of the probability space: the experiment tells us where we are in that partition. (E.g., if you are measuring some quantity to some number of significant digits, the cells of the partition are equivalence classes under equality of the quantity up to those many significant digits.) The following is then easy to prove:

Proposition 1: On definitions (2) and (3), an experiment E2 is epistemically at least as good as experiment E1 if and only if the partition associated with E2 is essentially at least as fine as the partition associated with E1.

A partition R2 is essentially at least as fine as a partition R1 provided that for every event A in R1 there is an event B in R2 such that with probability one B happens if and only if A happens. The definition is relative to the current credences which are assumed to be probabilistic. If the current credences are regular—all non-empty events have non-zero probability—then “essentially” can be dropped.

However, Proposition 1 suggests that our choice of definitions isn’t that helpful. Consider two experiments. On E1, all the faculty members from your Geology Department have their weight measured to the nearest hundred kilograms. On E2, a thousand randomly chosen individiduals around the world have their weight measured to the nearest kilogram. Intuitively, E1 is better. But Proposition 1 shows that in the above sense neither experiment is better than the other, since they generate partitions neither of which is essentially finer than the other (the event of there being a member of the Geology Department with weight at least 150 kilograms is in the partition of E2 but nothing coinciding with that event up to probability zero is in the partition of E1). And this is to be expected. For suppose that our research priority is to know whether any members of your Geology Department are at least than 150 kilograms in weight, because we need to know if for a departmental cave exploring trip the current selection of harnesses all of which are rated for users under 150 kilograms are sufficient. Then E1 is better. On the other hand, if our research priority is to know the average weight of a human being to the nearest ten kilograms, then E2 is better.

The problem with our definitions is that the range of possible research priorities is just too broad. Here is one interesting way to narrow it down. When we are talking about an experiment’s epistemic value, we mean the value of the experiment towards a set of questions. If the set of questions is a scientifically typical set of questions about human population weight distribution, then E1 seems better than E2. But if it is an atypical set of questions about the Geology Department members’ weight distribution, then E2 might be better. We can formalize this, too. We can identify a set Q of questions with a partition of probability space representing the possible answers. This partition then generates an algebra FQ on the probability space, which we can call the “question algebra”. Now we can relativize our definitions to a set of questions.

  1. E2 is epistemically at least as good an experiment as E1 for a set of questions Q provided that for every epistemically reasonable research priority on Q, the expected degree to which E2 would serve the priority is at least as high as the expected degree to which E1 would.

  2. A reasonable epistemic research priority on a set of questions Q is a strictly proper scoring rule or epistemic utility on FQ, and the expected degree to which an experiment would serve Q is equal to the expected value of the score after Bayesian update on the result of the experiment.

We recover the old definitions by being omnicurious, namely letting Q be all possible questions.

What about Proposition 1? Well, one direction remains: if E2’s partition is essentially at least as fine as E1’s, then E2 is better with regard any set of questions, an in particular better with regard to Q. But what about the other direction? Now the answer is negative. Suppose the question is what the average weight of the six members of the Geology Department is up to the nearest 100 kg. Consider two experiments: on the first, the members are ordered alphabetically by first name, and a fair die is rolled to choose one (if you roll 1, you choose the first, etc.), and their height is measured. On the second, the same is done but with the ordering being by last name. Assuming the two orderings are different, neither experiment’s partition is essentially at least as fine as the other’s, but the expected contributions of both experiments towards our question is equal.

Is there a nice characterization in terms of partitions of when E2 is at least as good as E1 with regard to a set of questions Q? I don’t know. It wouldn’t surprise me if there was something in the literature. A nice start would be to see if we can answer the question in the special case where Q is a single binary question and where E1 and E2 are binary experiments. But I need to go for a dental appointment now.

Friday, January 17, 2025

Knowledge and anti-knowledge

Suppose knowledge has a non-infinitesimal value. Now imagine that you continuously gain evidence for some true proposition p, until your evidence is sufficient for knowledge. If you’re rational, your credence will rise continuously with the evidence. But if knowledge has a non-infinitesimal value, your epistemic utility with respect to p will have a discontinuous jump precisely when you attain knowledge. Further, I will assume that the transition to knowledge happens at a credence strictly bigger than 1/2 (that’s obvious) and strictly less than 1 (Descartes will dispute this).

But this leads to an interesting and slightly implausible consequence. Let T(r) be the epistemic utility of assigning evidence-based credence r to p when p is true, and let F(r) be the epistemic utility of assigning evidence-based credence r to p when p is false. Plausibly, T is a strictly increasing function (being more confident in a truth is good) and F is a strictly decreasing function (being more confident in a falsehood is bad). Furthermore, the pair T and F plausibly yields a proper scoring rule: whatever one’s credence, one doesn’t have an expectation that some other credence would be epistemically better.

It is not difficult to see that these constraints imply that if T has a discontinuity at some point 1/2 < rK < 1, so does F. The discontinuity in F implies that as we become more and more confident in the falsehood p, suddenly we have a discontinuous downward jump in utility. That jump occurs precisely at rK, namely when we gain what we might call “anti-knowledge”: when one’s evidence for a falsehood becomes so strong that it would constitute knowledge if the proposition were true.

Now, there potentially are some points where we might plausibly think that epistemic utility of having a credence in a falsehood takes a discontinuous downward jump. These points are:

  • 1, where we become certain of the falsehood

  • rB, the threshold of belief, where the credence becomes so high that we count as believing the falsehood

  • 1/2, where we start to become more confident in the falsehood p than the truth not-p

  • 1 − rB, where we stop believing not-p, and

  • 0, where the falsehood p becomes an epistemic possibility.

But presumably rK is strictly between rB and 1, and hence rK is no one of these points. Is it plausible to think that there is a discontinuous downward jump in epistemic utility when we achieve anti-knowledge by crossing the threshold rK in a falsehood.

I am incline to say not. But that forces me to say that there is no discontinuous upward jump in epistemic utility once we gain knowledge.

On the other hand, one might think that the worst kind of ignorance is when you’re wrong but you think you have knowledge, and that’s kind of like the anti-knowledge point.

Monday, August 5, 2024

Natural reasoning vs. Bayesianism

A typical Bayesian update gets one closer to the truth in some respects and further from the truth in other respects. For instance, suppose that you toss a coin and get heads. That gets you much closer to the truth with respect to the hypothesis that you got heads. But it confirms the hypothesis that the coin is double-headed, and this likely takes you away from the truth. Moreover, it confirms the conjunctive hypothesis that you got heads and there are unicorns, which takes you away from the truth (assuming there are no unicorns; if there are unicorns, insert a “not” before “are”). Whether the Bayesian update is on the whole a plus or a minus depends on how important the various propositions are. If for some reason saving humanity hangs on you getting it right whether you got heads and there are unicorns, it may well be that the update is on the whole a harm.

(To see the point in the context of scoring rules, take a weighted Brier score which puts an astronomically higher weight on you got heads and there are unicorns than on all the other propositions taken together. As long as all the weights are positive, the scoring rule will be strictly proper.)

This means that there are logically possible update rules that do better than Bayesian update. (In my example, leaving the probability of the proposition you got heads and there are unicorns unchanged after learning that you got heads is superior, even though it results in inconsistent probabilities. By the domination theorem for strictly proper scoring rules, there is an even better method than that which results in consistent probabilities.)

Imagine that you are designing a robot that maneouvers intelligently around the world. You could make the robot a Bayesian. But you don’t have to. Depending on what the prioritizations among the propositions are, you might give the robot an update rule that’s superior to a Bayesian one. If you have no more information than you endow the robot with, you won’t be able to expect to be able to design such an update rule. (Bayesian update has optimal expected accuracy given the pre-update information.) But if you know a lot more than you tell the robot—and of course you do—you might well be able to.

Imagine now that the robot is smart enough to engage in self-reflection. It then notices an odd thing: sometimes it feels itself pulled to make inferences that do not fit with Bayesian update. It starts to hypothesize that by nature it’s a bad reasoner. Perhaps it tries to change its programming to be more Bayesian. Would it be rational to do that? Or would it be rational for it to stick to its programming, which in fact is superior to Bayesian update? This is a difficult epistemology question.

The same could be true for humans. God and/or evolution could have designed us to update on evidence differently from Bayesian update, and this could be epistemically superior (God certainly has superior knowledge; evolution can “draw on” a myriad of information not available to individual humans). In such a case, switching from our “natural update rule” to Bayesian update would be epistemically harmful—it would take us further from the truth. Moreover, it would be literally unnatural. But what does rationality call on us to do? Does it tell us to do Bayesian update or to go with our special human rational nature?

My “natural law epistemology” says that sticking with what’s natural to us is the rational thing to do. We shouldn’t redesign our nature.