Showing posts with label propriety. Show all posts
Showing posts with label propriety. Show all posts

Thursday, April 9, 2026

Epistemic utilities and death

In the previous post, I proved that we get a proper scoring rule if we compute epistemic utilities as follows. We start with our current credence assignment, consider what credence assignment we will have in the future after we update on some further evidence, and then score that. I then suggested that one could get a lifetime epistemic utility by adding up the epistemic utilities over all the moments of life, and as long as death wasn’t random—as long as the lifespan was fixed—this would generate a proper scoring rule. I then said that if death is random (as it is) it might be the case that you don’t get a proper scoring rule.

My conjecture was wrong. You still get a proper scoring rule despite random death. It’s easy to see this. The basic idea is this. Suppose that you might die the next moment. This partitions the probability space into two subsets, D and L, for death and life. Your current credence is p. Next moment, on D, your credence either doesn’t exist (because you don’t exist or you exist in some supernatural state where you don’t have credences) or doesn’t count (because I am after lifetime credences). Thus, the appropriate way to do a forward-looking scoring of your credence p is to score it s(pL) on L and 0 on D, where pL is the result of conditionalizing your credence on the evidence L (after all, if you are alive, you will conditionalize on being alive), and s is some proper scoring rule. In other words, your forward looking score is sL(p) = 1L ⋅ s(pL).

Is this score proper? Yes! For by propriety of s we have:

  • EpL(s(pL)) ≥ EpL(s(qL)).

But this is the same as:

  • (p(L))−1Ep(1Ls(pL)) ≥ (p(L))−1Ep(1Ls(qL)).

Multiplying both sides by p(L) (I am assuming a non-zero probability of survival), we get:

  • Ep(sL(p)) ≥ Ep(sL(q)).

We can combine this with a more complex set of future investigations as in the previous post, and things will still work.

It is crucial to the above argument that when you’re alive, you can tell you’re alive. I suppose that’s not always true. When you’re asleep, you are alive, but can’t tell you’re alive. So to generalize beyond the above toy example, replace death with unconsciousness or something like that.

Forward-looking scoring rules

An accuracy scoring rule assigns a score to a probability function representing an agent’s credences, ostensibly measuring how close that probability function is to the truth. The score s(p) of a probability function p is a random variable, because the value of the score depends on what is actually true, i.e., on where we are in the probability space.

A proper scoring rule (on probabilistic credences) satisfies the propriety inequality

  1. Eps(p) ≥ Eps(q)

which says that the expected score of your current lights—your current credences p—by your current lights is optimal: you won’t improve your expected score (by your current lights) by switching to a different credence q.

You can think of a proper scoring rule as representing the epistemic utility of having a credence p.

But now let’s think about things dynamically. In the future, you will receive additional evidence. As a good Bayesian agent, you will update on this evidence by conditionalization. Perhaps instead of thinking about maximizing your current score, you should think about maximizing your future score. Maybe your true epistemic utility is the score you will end up with after all the future evidence is in.

A simple model of this is as follows. There is some finite partition I = (I1,...,In) of your probability space Ω with each cell Ii of the partition representing a possibility for what you might learn given future evidence. Your current credence function is p, and p(Ii) > 0 for all i. There is then a random credence function pI where pI(ω) is the credence function you will have once the evidene is in if you are at ω ∈ Ω. In other words, pI(ω)(A) = p(AIi) where Ii is the member of the partition that contains ω. (Technically, the function that maps ω to pI(ω)(A) is equal to the conditional probability p(AG) where G is the algebra generated by I.)

Now, given a proper scoring rule s, define a new scoring rule sI as follows:

  1. sI(p)(ω) = s(pI(ω))(ω).

Your sI-score for p at ω then represents the score you will have at ω once you learn which cell of the partition I you are in.

Theorem: The scoring rule sI is proper if s is proper.

Note that sI won’t be strictly proper (i.e., (1) won’t always have strict inequality when p and q are distinct) if I has two or more cells, because pI and qI are going to be the same if p and q assign different probabilities to the cells, but have the same conditional probabilities on each cell. But it might still be the case that sI is strictly proper with respect to some relevant subfield of Ω—that needs some further investigation.

Suppose now you are a Bayesian agent who is guaranteed to consciously live for n moments. In each moment, new information comes in. Thus, we have a sequence J0, ..., Jn of finer and finer partitions, with J0 being the trivial partition, and with pJk representing the credence you will have at time k. Your overall epistemic lifetime score is then:

  1. sΣ(p) = ∑ksJk(p).

It follows from the Theorem that sΣ is a proper scoring rule if s is. And if J0 is the trivial partition, then sJ0 = s, and so if s is strictly proper, then the lifetime score sΣ is strictly proper, since the sum of a strictly proper rule and a proper rule is strictly proper. So, lifetime scores are strictly proper if they are constructed from an instantaneous score—in the above toy model.

Alas, the toy model is not fully adequate, because it is random when we will die, and so our lifespan doesn’t have a fixed sequence of moments. Once we take into account the randomness of when we will die, the overall epistemic lifetime score might stop being proper: this needs further investigation.

Proof of Theorem: By the Greaves and Wallace Theorem, an optimal method of updating credences with respect to expected proper score is by Bayesian conditionalization. Apply the Greaves and Wallace Theorem to the scoring rule s and the starting credence p with the following two strategies:

A. Bayesian conditionalization on the true cell of I.

B. Switch your credence from p to q, then apply Bayesian conditionalization on the true cell of I.

Saying that (A) is at least as good as (B) is equivalent to the the proper scoring rule inequality (1) for sI.

Monday, February 24, 2025

More on averaging to combine epistemic utilities

Suppose that the right way to combine epistemic utilities across people is averaging: the overall epistemic utility of the human race is the average of the individual epistemic utilities. Suppose, further, that each individual epistemic utility is strictly proper, and you’re a “humanitarian” agent who wants to optimize overall epistemic utility.

Suppose you’re now thinking about two hypotheses about how many people exist: the two possible numbers are m and n, which are not equal. All things considered, you have credence 0 < p0 < 1 in the hypothesis Hm that there are m people and 1 − p0 in the hypothesis Hn that there are n people. You now want to optimize overall epistemic utility. On an averaging view, if Hm is true, if your credence is p1, your contribution to overall epistemic utility will be:

  • (1/m)T(p1)

and if Hm is false, your contribution will be:

  • (1/n)F(p1),

where your strictly proper scoring rule is given by T, P. Since your credence is p1, by your lights the expected value after your changing your credence to p0 will be:

  • p0(1/m)T(p1) + (1−p0)(1/n)F(p1) + Q(p0)

where Q(p0) is the contribution of other people’s credences, which I assume you do not affect with your choice of p1. If m ≠ n and T, F is strictly proper, the expected value will be maximized at

  • p1 = (p0/m)/(p0/m+(1−p0)/n) = np0/(np0+m(1−p0)).

If m > n, then p1 < p0 and if m < n, then p1 > p0. In other words, as long as n ≠ m, if you’re an epistemic humanitarian aiming to improve overall epistemic utility, any credence strictly between 0 and 1 will be unstable: you will need to change it. And indeed your credence will converge to 0 if m > n and to 1 if m < n. This is absurd.

I conclude that we shouldn’t combine epistemic utilities across people by averaging the utilities.

Idea: What about combining them by computing the epistemic utilities of the average credences, and then applying a strictly proper scoring rule, in effect imagining that humanity is one big committee and that a committee’s credence is the average of the individual credences?

This is even worse, because it leads to problems even without considering hypotheses on which the number of people varies. Suppose that you’ve just counted some large number nobody cares about, such as the number of cars crossing some intersection in New York City during a specific day. The number you got is even, but because the number is big, you might well have made a mistake, and so your credence that the number is even is still fairly low, say 0.7. The billions of other people on earth all have credence 0.5, and because nobody cares about your count, you won’t be able to inform them of your “study”, and their credences won’t change.

If combined epistemic utility is given by applying a proper scoring rule to the average credence, then by your lights the expected value of the combined epistemic utility will increase the bigger you can budge the average credence, as long as you don’t get it above your credence. Since you can really only affect your own credence, as an epistemic humanitarian your best bet is to set your credence to 1, thereby increasing overall human credence from 0.5 to around 0.5000000001, and making a tiny improvement in the expected value of the combined epistemic utility of humankind. In doing so, you sacrifice your own epistemic good for the epistemic good of the whole. This is absurd!

I think the idea of averaging to produce overall epistemic utilities is just wrong.

Monday, January 20, 2025

Open-mindedness and epistemic thresholds

Fix a proposition p, and let T(r) and F(r) be the utilities of assigning credence r to p when p is true and false, respectively. The utilities here might be epistemic or of some other sort, like prudential, overall human, etc. We can call the pair T and F the score for p.

Say that the score T and F is open-minded provided that expected utility calculations based on T and F can never require you to ignore evidence, assuming that evidence is updated on in a Bayesian way. Assuming the technical condition that there is another logically independent event (else it doesn’t make sense to talk about updating on evidence), this turns out to be equivalent to saying that the function G(r) = rT(r) + (1−r)F(r) is convex. The function G(r) represents your expected value for your utility when your credence is r.

If G is a convex function, then it is continuous on the open interval (0,1). This implies that if one of the functions T or F has a discontinuity somewhere in (0,1), then the other function has a discontinuity at the same location. In particular, the points I made in yesterday’s post about the value of knowledge and anti-knowledge carry through for open-minded and not just proper scoring rules, assuming our technical condition.

Moreover, we can quantify this discontinuity. Given open-mindedness and our technical condiiton, if T has a jump of size δ at credence r (e.g., in the sense that the one-sided limits exist and differ by y), then F has a jump of size rδ/(1−r) at the same point. In particular, if r > 1/2, then if T has a jump of a given size at r, F has a larger jump at r.

I think this gives one some reason to deny that there are epistemically important thresholds strictly between 1/2 and 1, such as the threshold between non-belief and belief, or between non-knowledge and knowledge, even if the location of the thresholds depends on the proposition in question. For if there are such thresholds, then now imagine cases of propositions p with the property that it is very important to reach a threshold if p is true while one’s credence matters very little if p is false. In such a case, T will have a larger jump at the threshold than F, and so we will have a violation of open-mindedness.

Here are three examples of such propositions:

  • There are objective norms

  • God exists

  • I am not a Boltzmann brain.

There are two directions to move from here. The first is to conclude that because open-mindedness is so plausible, we should deny that there are epistemically important thresholds. The second is to say that in the case of such special propositions, open-mindedness is not a requirement.

I wondered initially whether a similar argument doesn’t apply in the absence of discontinuities. Could one have T and F be openminded even though T continuously increases a lot faster than F decreases? The answer is positive. For instance the pair T(r) = e10r and F(r) =  − r is open-minded (though not proper), even though T increases a lot faster than F decreases. (Of course, there are other things to be said against this pair. If that pair is your utility, and you find yourself with credence 1/2, you will increase your expected utility by switching your credence to 1 without any evidence.)

Friday, January 17, 2025

Knowledge and anti-knowledge

Suppose knowledge has a non-infinitesimal value. Now imagine that you continuously gain evidence for some true proposition p, until your evidence is sufficient for knowledge. If you’re rational, your credence will rise continuously with the evidence. But if knowledge has a non-infinitesimal value, your epistemic utility with respect to p will have a discontinuous jump precisely when you attain knowledge. Further, I will assume that the transition to knowledge happens at a credence strictly bigger than 1/2 (that’s obvious) and strictly less than 1 (Descartes will dispute this).

But this leads to an interesting and slightly implausible consequence. Let T(r) be the epistemic utility of assigning evidence-based credence r to p when p is true, and let F(r) be the epistemic utility of assigning evidence-based credence r to p when p is false. Plausibly, T is a strictly increasing function (being more confident in a truth is good) and F is a strictly decreasing function (being more confident in a falsehood is bad). Furthermore, the pair T and F plausibly yields a proper scoring rule: whatever one’s credence, one doesn’t have an expectation that some other credence would be epistemically better.

It is not difficult to see that these constraints imply that if T has a discontinuity at some point 1/2 < rK < 1, so does F. The discontinuity in F implies that as we become more and more confident in the falsehood p, suddenly we have a discontinuous downward jump in utility. That jump occurs precisely at rK, namely when we gain what we might call “anti-knowledge”: when one’s evidence for a falsehood becomes so strong that it would constitute knowledge if the proposition were true.

Now, there potentially are some points where we might plausibly think that epistemic utility of having a credence in a falsehood takes a discontinuous downward jump. These points are:

  • 1, where we become certain of the falsehood

  • rB, the threshold of belief, where the credence becomes so high that we count as believing the falsehood

  • 1/2, where we start to become more confident in the falsehood p than the truth not-p

  • 1 − rB, where we stop believing not-p, and

  • 0, where the falsehood p becomes an epistemic possibility.

But presumably rK is strictly between rB and 1, and hence rK is no one of these points. Is it plausible to think that there is a discontinuous downward jump in epistemic utility when we achieve anti-knowledge by crossing the threshold rK in a falsehood.

I am incline to say not. But that forces me to say that there is no discontinuous upward jump in epistemic utility once we gain knowledge.

On the other hand, one might think that the worst kind of ignorance is when you’re wrong but you think you have knowledge, and that’s kind of like the anti-knowledge point.

Monday, February 13, 2023

Strictly proper scoring rules in an infinite context

This is mostly just a note to self.

In a recent paper, I prove that there is no strictly proper score s on the countably additive probabilities on the product space 2κ where we require s(p) to be measurable for every probability p and where κ is a cardinal such that 2κ > κω (e.g., the continuum). But here I am taking strictly proper scores to be real- or extended-real valued. What if we relax this requirement? The answer is negative.

Theorem: Assume Choice. If 2κ > κω and V is a topological space with the T1 property and ≤ is a total pre-order on V. Let s be a function s from the extreme probabilities on Ω = 2κ, with the product σ-algebra, concentrated on points to the set of measurable functions from Ω to V. Then there exist extreme probabilities r and q concentrated at distinct points α and β respectively such that s(r)(α) ≤ s(q)(α).

A space is T1 iff singletons are closed.
Extreme probabilities only have 0 and 1 values. An extreme probability p is concentrated at α provided that p(A) = 1 if α ∈ A and p(A) = 0 otherwise. (We can identify extreme probabilities with the points at which they are concentrated, as long as the σ-algebra separates points.) Note that in the product σ-lagebra on 2κ singletons are not measurable.

Corollary: In the setting of the Theorem, if Ep is a V-valued prevision from V-valued measurable functions to V such that (a) Ep(c⋅1A) = cp(A) always, (b) if f and g are equal outside of a set with p-measure zero, then Epf = Epg, and (c) Eps(q) is always defined for extreme singleton-concentrated p and q, then the strict propriety inequality Ers(r) > Ers(q) fails for some distinct extreme singleton-concentrated p and q.

Proof of Corollary: Suppose p is concentrated at α and Epf is defined. Let Z = {ω : f(ω) = f(α)}. Then p(Z) = 1 and p(Zc) = 0. Hence f is p-almost surely equal to f(α) ⋅ 1Z. Thus, Epf = f(α). Thus if r and q are as in the Theorem, Ers(r) = s(r)(α) and Ers(q) = s(q)(α), and our result follows from the Theorem.

Note: Towards the end of my paper, I suggested that the unavailability of proper scores in certain contexts is due to the limited size of the set of values—normally assumed to be real. But the above shows that sometimes it’s due to measurability constraints instead, as we still won’t have proper scores even if we let the values be some gigantic collection of surreals.

Proof of Theorem: Given an extreme probability p and z ∈ V, the inverse image of z under s(p), namely Lp, z = {ω : s(p) = z}, is measurable. A measurable set A depends only on a countable number of factors of 2κ: i.e., there is a countable set C ⊆ κ such that if ω, ω′ ∈ κ agree on C, and one is in A, so is the other. Let Cp ⊂ κ be a countable set such that Lp, s(p)(ω) depends only on Cp, where ω is the point p is concentrated on (to get all the Cp we need the Axiom of Choice). The set of all the countable subsets of κ has cardinality at most κω, and hence so does the set of all the Cp. Let Zp = {q : Cq = Cp}. The union of the Zp is all the extreme point-concentrated probabilities and hence has cardinality 2κ. There are at most κω different Zp. Thus for some p, the set Zp has cardinality bigger than κω (here we used the assmption that 2ω > κω).

There are q and r in Zp concentrated at α and β, respectively, such that αCp = βCp (i.e., α and β agree on Cp). For the function (⋅)∣Cp on the concentration points of the probabilities in Zp has an image of cardinality at most 2ω ≤ κω but Zp has cardinality bigger than κω. Since αCp = βCp, and Lq, s(q)(α) and Lr, s(r)(β) both depend only on the factors in Cp, it follows that s(q)(β) = s(q)(α) and s(r)(α) = s(r)(β). Now either s(q)(β) ≥ s(r)(α) or s(q)(β) ≤ s(r)(α) by totality. Swapping (q,α) with (r,β) if necessary, assume s(q)(α) = s(q)(β) ≤ s(r)(α).

Wednesday, February 1, 2023

Open-mindedness and propriety

Suppose we have a probability space Ω with algebra F of events, and a distinguished subalgebra H of events on Ω. My interest here is in accuracy H-scoring rules, which take a (finitely-additive) probability assignment p on H and assigns to it an H-measurable score function s(p) on Ω, with values in [−∞,M] for some finite M, subject to the constraint that s(p) is H-measurable. I will take the score of a probability assignment to represent the epistemic utility or accuracy of p.

For a probability p on F, I will take the score of p to be the score of the restriction of p to H. (Note that any finitely-additive probability on H extends to a finitely-additive probability on F by Hahn-Banach theorem, assuming Choice.)

The scoring rule s is proper provided that Eps(q) ≤ Eps(p) for all p and q, and strictly so if the inequality is strict whenever p ≠ q. Propriety says that one never expects a different probability from one’s own to have a better score (if one did, wouldn’t one have switched to it?).

Say that the scoring rule s is open-minded provided that for any probability p on F and any finite partition V of Ω into events in F with non-zero p-probability, the p-expected score of finding out where in V we are and conditionalizing on that is at least as big as the current p-expected score. If the scoring rule is open-minded, then a Bayesian conditionalizer is never precluded from accepting free information. Say that the scoring rule s is strictly open-minded provided that the p-expected score increases of finding out where in V we are and conditionalizing increases whenever there is at least one event E in V such that p(⋅|E) differs from p on H and p(E) > 0.

Given a scoring rule s, let the expected score function Gs on the probabilities on H be defined by Gs(p) = Eps(p), with the same extension to probabilities on F as scores had.

It is well-known that:

  1. The (strict) propriety of s entails the (strict) convexity of Gs.

It is easy to see that:

  1. The (strict) convexity of Gs implies the (strict) open-mindedness of s.

Neither implication can be reversed. To see this, consider the single-proposition case, where Ω has two points, say 0 and 1, and H and F are the powerset of Ω, and we are interested in the proposition that one of these point, say 1, is the actual truth. The scoring rule s is then equivalent to a pair of functions T and F on [0,1] where T(x) = s(px)(1) and F(x) = s(px)(0) where px is the probability that assigns x to the point 1. Then Gs corresponds to the function xT(x) + (1−x)F(x), and each is convex if and only if the other is.

To see that the non-strict version of (1) cannot be reversed, suppose (T,F) is a non-trivial proper scoring rule with the limit of F(x)/x as x goes to 0 finite. Now form a new scoring rule by letting T * (x) = T(x) + (1−x)F(x)/x. Consider the scoring rule (T*,0). The corresponding function xT * (x) is going to be convex, but (T*,0) isn’t going to be proper unless T* is constant, which isn’t going to be true in general. The strict version is similar.

To see that (2) cannot be reversed, note that the only non-trivial partition is {{0}, {1}}. If our current probability for 1 is x, the expected score upon learning where we are is xT(1) + (1−x)F(0). Strict open-mindedness thus requires precisely that xT(x) + (1−x)F(x) < xT(1) + (1−x)F(0) whenever x is neither 0 nor 1. It is clear that this is not enough for convexity—we can have wild oscillations of T and F on (0,1) as long as T(1) and F(1) are large enough.

Nonetheless, (2) can be reversed (both in the strict and non-strict versions) on the following technical assumption:

  1. There is an event Z in F such that Z ∩ A is a non-empty proper subset of A for every non-empty member of H.

This technical assumption basically says that there is a non-trivial event that is logically independent of everything in H. In real life, the technical assumption is always satisfied, because there will always be something independent of the algebra H of events we are evaluating probability assignments to (e.g., in many cases Z can be the event that the next coin toss by the investigator’s niece will be heads). I will prove that (2) can be reversed in the Appendix.

It is easy to see that adding (3) to our assumptions doesn’t help reverse (1).

Since open-mindedness is pretty plausible to people of a Bayesian persuasion, this means that convexity of Gs can be motivated independently of propriety. Perhaps instead of focusing on propriety of s as much as the literature has done, we should focus on the convexity of Gs?

Let’s think about this suggestion. One of the most important uses of scoring rules could be to evaluate the expected value of an experiment prior to doing the experiment, and hence decide which experiment we should do. If we think of an experiment as a finite partition V of the probability space with each cell having non-zero probability by one’s current lights p, then the expected value of the experiment is:

  1. A ∈ Vp(A)EpAs(pA) = ∑A ∈ Vp(A)Gs(pA),

where pA is the result of conditionalizing p on A. In other words, to evaluate the expected values of experiments, all we care about is Gs, not s itself, and so the convexity of Gs is a very natural condition: we are never oligated to refuse to know the results of free experiments.

However, at least in the case where Ω is finite, it is known that any (strictly) convex function (maybe subject to some growth conditions?) is equal to Gu for a some (strictly) proper scoring rule u. So we don’t really gain much generality by moving from propriety of s to convexity of Gs. Indeed, the above observations show that for finite Ω, a (strictly) open-minded way of evaluating the expected epistemic values of experiments in a setting rich enough to satisfy (3) is always generatable by a (strictly) proper scoring rule.

In other words, if we have a scoring rule that is open-minded but not proper, we can find a proper scoring rule that generates the same prospective evaluations of the value of experiments (assuming no special growth conditions are needed).

Appendix: We now prove the converse of (2) assuming (3).

Assume open-mindedness. Let p1 and p2 be two distinct probabilities on H and let t ∈ (0,1). We must show that if p = tp1 + (1−t)p2, then

  1. Gs(p) ≤ tGs(p1) + (1−t)Gs(p2)

with the inequality strict if the open-mindedness is strict. Let Z be as in (3). Define

  1. p′(AZ) = tp1(A)

  2. p′(AZc) = (1−t)p2(A)

  3. p′(A) = p(A)

for any A ∈ H. Then p′ is a probability on the algebra generated by H and Z extending p. Extend it to a probability on F by Hahn-Banach. By open-mindedness:

  1. Gs(p′) ≤ p′(Z)EpZs(pZ) + p′(Zc)EpZcs(pZc).

But p′(Z) = p(ΩZ) = t and p′(Zc) = 1 − t. Moreover, pZ = p1 on H and pZc = p2 on H. Since H-scores don’t care what the probabilities are doing outside of H, we have s(pZ) = s(p1) and s(pZc) = s(p2) and Gs(p′) = Gs(p). Moreover our scores are H-measurable, so EpZs(p1) = Ep1s(p1) and EpZcs(p2) = Ep2s(p2). Thus (9) becomes:

  1. Gs(p) ≤ tGs(p1) + (1−t)Gs(p2).

Hence we have convexity. And given strict open-mindedness, the inequality will be strict, and we get strict convexity.

Sunday, September 25, 2022

A strict propriety argument for probabilism without any continuity assumptions

Here’s an accuracy-theoretic argument for probabilism (the thesis that only probabilities are rationally admissible credences) on finite spaces that does not make any continuity assumptions on the scoring rule. I will assume all credence functions take values on [0,1].

  1. All probabilities are rationally admissible credences.

  2. If any non-probabilities are rationally admissible, then all non-probabilities satisfying Normalization (whole space has credence 1) and Subadditivity (P(A) + P(B) ≤ P(AB) when A and B are disjoint) are rationally admissible with the appropriate prevision being given by a level set integral [correction: actually, I need LSI, not the version of LSI in the earlier blog post].

  3. A rationally appropriate scoring rule s satisfies strict propriety for all rationally admissible credences with an appropriate prevision: if V is an appropriate prevision then Vus(u) is better than Vus(v) whenever u and v are different rationally admissible credences.

  4. There is a rationally appropriate scoring rule.

But now we have a cute theorem:

  • On any finite space Ω with at least two points, no scoring rule satisfies strict propriety for the credences with Normalization and Subadditivity and level set integral prevision.

It follows no non-probabilities are rationally admissible.

Is this a good argument? I find (2) somewhat plausible—it’s hard to think of a less problematic weakening of the axioms of probability than from Additivity to Subadditivity, and I have not been able to find a better prevision than the level set integral one. Standard arguments for probabilism assume strict propriety for all probabilities. But it seems to me that a non-probabilist will find strict propriety for all probabilities plausible only insofar as they find strict propriety for all admissible credences plausible. Thus (3) is dialectically as good as the usual strict propriety assumption.

I think the non-probabilist’s best way out is to deny strict propriety or to deny that there is a rationally appropriate scoring rule. Both of these ways out work just as well against more standard arguments for probabilism, and I think both are good ways out.

Technically speaking, the advantage of this argument over standard arguments for probabilism is that it makes no assumptions of continuity.

Monday, May 2, 2022

Truth-directedness and propriety of scoring rules does not imply strict propriety

A scoring rule assigns a score to a credence assignment (which can but need not satisfy the axioms of probability), where a score is a random variable measuring how close the credence assignment is to the truth.

A scoring rule is strictly truth-directed provided that if c is a credence assignment that is closer to the truth than c is at ω, then c gets a better a score at ω. A scoring rule is proper provided that for all probabilities p, the p-expected value of the score of a probability p is at least as good as the p-expected value of the score of any other credence, and is strictly proper.

Propriety for a scoring rule is a pretty plausible condition, but it’s a bit harder to argue philosophically for strict propriety. But scoring-rule based philosophical arguments for probabilism—the doctrine that credences ought to be probabilities—require strict propriety.

In a clever move, Campbell-Moore and Levinstein showed that propriety plus strict truth-directedness and additivity (the idea that the score can be decomposed into a sum of single-event scores) implies strict propriety.

Here’s an interesting fact I will show: propriety plus strict truth-directedness do not imply strict propriety in the absence of additivity. Further, my counterexample will be bounded, infinitely differentiable and strictly proper on the probabilities. Personally don’t find additivity all that plausible, so I conclude the Campbell-Moore and Levinstein move does not move the discussion of strict propriety and probabilism ahead much.

Let Ω = {0, 1}. Given a credence function c (with values in [0,1]) on the powerset of Ω, define the credence function c* which has the same value as c on the empty set and on Ω, but where c*({0}) is the number z in [0,1] that minimizes (c({0})−z)2 + (c({1})−(1−z))2, and where c*({1}) = 1 − c*({0}). In other words, c* is the credence function closest to c in the Euclidean metric such that c*({0}) + c*({1}) = 1.

Now let b*(c) = b(c*). Then b* agrees with b score on the probabilities, and hence is strictly proper on them. Further, every value of b* is a Brier score of some credence, and hence b* is proper.

We now check that it is strictly truth-directed. Brier scores are strictly truth-directed. Thus, replacing a credence function with one that is closer to the truth on Ω or on the empty set will improve the b* score. Moreover, it is easy to check that c*({0}) = (1+c({0})−c({1}))/2. It’s easy to check that if we tweak c({0}) to move us closer to the truth at some fixed ω ∈ {0, 1}, then c* will be closer to the truth at ω as well, and similarly if we tweak c({1}) to be closer to the truth at ω, and in both cases we will improve the score by the strict truth-directedness of Brier scores.

Finally, however, note that b* is not strictly proper and does not have a domination theorem of the sort used in arguments for probabilism, since the b*-score of any credence c that fails to be a probability due to its being the case c({0}) + c({1}) ≠ 1 but that gets the right values on the empty set and Ω (zero and one, respectively) is equal to the b*-score of c*, and c* will be a probability in that case.

Note that in the example above we don't have quasi-strict propriety either.

Wednesday, December 1, 2021

Investigative scoring rules

Let s be an accuracy scoring rule on a finite probability space. Thus, s(P) is a random variable measuring how close a probability assignment P is to the truth. Here are two reasonable conditions on the rule (the name for the second is made up):

  1. Propriety: EPs(P)≤EPs(Q) for any distinct probability assignments P and Q.

  2. Investigativeness: EPs(P)≤P(A)EPAs(PA)+P(Ac)EPAcs(PAc) whenever 0 < P(A)<1.

where EP is expected value with respect to P, PA is short for P(⋅|A), and Ac is the complement of A. Propriety says that if we are trying to maximize expected accuracy, we will never have reason to evidencelessly switch to a different credence. Investigativeness says that expected accuracy maximization never requires one to close one’s eyes to evidence because the expected accuracy after conditionalizing on learning whether A holds is at least as good as the currently expected accuracy. And we have strict versions of the two conditions provided the inequalities are always strict.

It is well-known that propriety implies investigativeness, and ditto for the strict variants.

One might guess that the other direction holds as well: that investigativeness implies propriety. But (perhaps surprisingly) not! In fact, strict investigativeness does not imply propriety.

Let s(P) be the following score: s(P)(w)=|{A : P(A)=1 and w ∈ A}|. In other words, s(P) measures how many true propositions P assigns probability 1 to. It is easy to see that s(PA)≥s(P) everywhere on A, and ditto for Ac in place of A, so the right-hand side in (2) is at least as big as P(A)EPAs(P)+P(Ac)EPAcs(P)=EPs(P).

But propriety does not hold as long as our probability space has at least two points. For let P be any regular probability—one that assigns a non-zero value to every non-empty set—and let Q be any probability concentrated at one point w0. Then s(P)=1 everywhere (the only subset P assigns probability 1 to is the whole space) while EPs(Q)≥1 + P({w0}) > 1 (since Q assigns probability 1 to {w} and to the whole space), and so we don’t have propriety.

If we want strict investigativeness, just replace s with s + ϵs′ where s′ is a Brier score and ϵ is small and positive. Then we will have strict investigativeness for s′, and hence for s + ϵs′ as well, but if ϵ is sufficiently small, we won’t have propriety.

It is interesting to think if investigativeness plus some additional plausible condition might imply propriety. A very plausible further condition is that if P is at least as close to the truth as Q for every event, then P gets a no-worse score. Another plausible condition is additivity. But my examples satisfy both conditions. I don’t see other plausible conditions to add, besides propriety as such.