Showing posts with label comparative probability. Show all posts
Showing posts with label comparative probability. Show all posts

Saturday, October 19, 2024

There is no canonical way to define a regular comparative probability in terms of a full conditional probability

I claim that there is no general, straightforward and satisfactory way to define a total comparative probability with the standard axioms using full conditional probabilities. By a “straightforward” way, I mean something like:

  1. A ≲ B iff P(AB|AΔB) ≤ P(BA|AΔB)

or:

  1. A ≲ B iff P(A|AB) ≤ P(A|AB) (Pruss).

The standard axioms of comparative probability are:

  1. Transitivity, reflexivity and totality.

  2. Non-negativity: ⌀ ≤ A for all A

  3. Additivity: If A ∪ B is disjoint from C, then A ≲ B iff A ∪ C ≲ B ∪ C.

A “straightforward” definition is one where the right-hand-side is some expression involving conditional probabilities of events definable in a boolean way in terms of A and B.

To be “satisfactory”, I mean that it satisfies some plausible assumptions, and the one that I will specifically want is:

  1. If P(A|C) < P(B|C) where A ∪ B ⊆ C, then A < B.

Definitions (1) and (2) are straightforward and satisfactory in the above-defined senses, but (1) does not satisfy transitivity while (2) does not satisfy the right-to-left direction of additivity.

Here is a proof of my claim. If the definition is straightforward, then if A ≲ B, and A′ and B are events such that there is a boolean algebra isomorphism ψ from the algebra of events generated by A and B to the algebra of events generated by A′ and B such that ψ(A) = A, ψ(B) = B and P(C|D) = P(ψ(C)|ψ(D)) for all C and D in the algebra generated by A and B, then A′ ≲ B.

Now consider a full conditional probability P on the interval [0,1] such that P(A|[0,1]) is equal to the Lebesgue measure of A when A is an interval. Let A = (0,1/4) and suppose B is either (1/4,1/2) or (1/4, 1/2]. Then there is an isomorphism ψ from the algebra generated by A and B to the same algebra that swaps A and B around and preserves all conditional probabilities. For the algebra consists of the eight possible unions of sets taken from among A, B and [0,1] − (AB), and it is easy to define a natural map between these eight sets that swaps A and B, and this will preserve all conditional probabilities. It follows from my definition of straightforwardness that we have A ≲ B if and only if we have B ≲ B. Since the totality axiom for comparative probabilities implies that either A ≲ B or B ≲ A, so we must have both A ≲ B and B ≲ A. Thus A ∼ B. Since this is true for both choices of B, we have

  1. (0,1/4) ∼ (1/4,1/2) ∼ (1/4, 1/2].

But now note that ⌀ < {1/2} by (3) (just let A = ⌀, B = {1/2} and C = {1/2}). The additivity axiom then implies that (1/4,1/2) < (1/4, 1/2], a contradiction.

I think that if we want to define a probability comparison in terms of conditional probabilities, what we need to do is to weaken the axioms of comparative probabilities. My current best suggestion is to replace Additivity with this pair of axioms:

  1. One-Sided Additivity: If A ∪ B is disjoint from C and A ≲ B, then A ∪ C ≲ B ∪ C.

  2. Weak Parthood Principle: If A and B are disjoint, then A < A ∪ B or B < A ∪ B.

Definition (2) satisfies the axioms of comparable probabilities with this replacement.

Here is something else going for this. In this paper, I studied the possibility of defining non-classical probabilities (full conditional, hyperreal or comparative) that are invariant under a group G of transformations. Theorem 1 in the paper characterizes when there are full conditional probabilities that are strongly invariant. Interesting, we can now extend Theorem 1 to include this additional clause:

  1. There is a transitive, reflexive and total relation satisfying (4), (8) and (9) as well as the regularity assumption that ⌀ < A whenever A is non-empty and that is invariant under G in the sense that gA ∼ A whenever both A and gA are subsets of Ω.

To see this, note that if there is are strongly invariant full conditional probabilities, then (2) will define in a way that satisfies (vi). For the converse, suppose (vi) is true. We show that condition (ii) of the original theorem is true, namely that there is no nonempty paradoxical subset. For to obtain a contradiction suppose there is a non-empty paradoxical subset E. Then E can be written as the disjoint union of A1, ..., An, and there are g1, ..., gn in G and 1 ≤ m < n such that g1A1, ..., gmAm and gm + 1Am + 1, ..., gnAn are each a partition of E.

A standard result for additive comparative probabilities in Krantz et al.’s measurement book is that if B1, ..., Bn are disjoint, and C1, ..., Cn are disjoint, with Bi ≲ Ci for all i, then B1 ∪ ... ∪ Bn ≲ C1 ∪ ... ∪ Cn. One can check that the proof only uses One-Sided Additivity, so it holds in our case. It follows from G-invariance that A1 ∪ ... ∪ Am ∼ E ∼ Am + 1 ∪ ... ∪ An. Since E is the disjoint union of A1 ∪ ... ∪ Am with Am + 1 ∪ ... ∪ An, this violates the Weak Parthood Principle.

Tuesday, October 15, 2024

More on full conditional probabilities and comparative probabilities

I claim that there is no general, straightforward and satisfactory way to define a total comparative probability with the standard axioms using full conditional probabilities. By a “straightforward” way, I mean something like:

  1. A ≲ B iff P(AB|AΔB) ≤ P(BA|AΔB)

or:

  1. A ≲ B iff P(A|AB) ≤ P(B|AB).

The standard axioms of comparative probability are:

  1. Transitivity, reflexivity and totality.

  2. Non-negativity: ⌀ ≤ A for all A

  3. Additivity: If A ∪ B is disjoint from C, then A ≲ B iff A ∪ C ≲ B ∪ C.

A “straightforward” definition is one where the right-hand-side is some expression involving conditional probabilities of events definable in a boolean way in terms of A and B.

To be “satisfactory”, I mean that it satisfies some plausible assumptions, and the one that I will specifically want is:

  1. If P(A|C) < P(B|C) where A ∪ B ⊆ C, then A < B.

Definitions (1) and (2) are straightforward and satisfactory in the above-defined senses, but (1) does not satisfy transitivity while (2) does not satisfy the right-to-left direction of additivity.

Here is a proof of my claim. If the definition is straightforward, then if A ≲ B, and A′ and B are events such that there is a boolean algebra isomorphism ψ from the algebra of events generated by A and B to the algebra of events generated by A′ and B such that ψ(A) = A, ψ(B) = B and P(C|D) = P(ψ(C)|ψ(D)) for all C and D in the algebra generated by A and B, then A′ ≲ B.

Now consider a full conditional probability P on the interval [0,1] such that P(A|[0,1]) is equal to the Lebesgue measure of A when A is an interval. Let A = (0,1/4) and suppose B is either (1/4,1/2) or (1/4, 1/2]. Then there is an isomorphism ψ from the algebra generated by A and B to the same algebra that swaps A and B around and preserves all conditional probabilities. For the algebra consists of the eight possible unions of sets taken from among A, B and [0,1] − (AB), and it is easy to define a natural map between these eight sets that swaps A and B, and this will preserve all conditional probabilities. It follows from my definition of straightforwardness that we have A ≲ B if and only if we have B ≲ B. Since the totality axiom for comparative probabilities implies that either A ≲ B or B ≲ A, so we must have both A ≲ B and B ≲ A. Thus A ∼ B. Since this is true for both choices of B, we have

  1. (0,1/4) ∼ (1/4,1/2) ∼ (1/4, 1/2].

But now note that ⌀ < {1/2} by (3) (just let A = ⌀, B = {1/2} and C = {1/2}). The additivity axiom then implies that (1/4,1/2) < (1/4, 1/2], a contradiction.

I think that if we want to define a probability comparison in terms of conditional probabilities, what we need to do is to weaken the axioms of comparative probabilities. My current best suggestion is to replace Additivity with this pair of axioms:

  1. One-Sided Additivity: If A ∪ B is disjoint from C and A ≲ B, then A ∪ C ≲ B ∪ C.

  2. Weak Parthood Principle: If A and B are disjoint, then A < A ∪ B or B < A ∪ B.

Definition (2) satisfies the axioms of comparable probabilities with this replacement.

Here is something else going for this. In this paper, I studied the possibility of defining non-classical probabilities (full conditional, hyperreal or comparative) that are invariant under a group G of transformations. Theorem 1 in the paper characterizes when there are full conditional probabilities that are strongly invariant. Interesting, we can now extend Theorem 1 to include this additional clause:

  1. There is a transitive, reflexive and total relation satisfying (4), (8) and (9) as well as the regularity assumption that ⌀ < A whenever A is non-empty and that is invariant under G in the sense that gA ∼ A whenever both A and gA are subsets of Ω.

To see this, note that if there is are strongly invariant full conditional probabilities, then (2) will define in a way that satisfies (vi). For the converse, suppose (vi) is true. We show that condition (ii) of the original theorem is true, namely that there is no nonempty paradoxical subset. For to obtain a contradiction suppose there is a non-empty paradoxical subset E. Then E can be written as the disjoint union of A1, ..., An, and there are g1, ..., gn in G and 1 ≤ m < n such that g1A1, ..., gmAm and gm + 1Am + 1, ..., gnAn are each a partition of E.

A standard result for additive comparative probabilities in Krantz et al.’s measurement book is that if B1, ..., Bn are disjoint, and C1, ..., Cn are disjoint, with Bi ≲ Ci for all i, then B1 ∪ ... ∪ Bn ≲ C1 ∪ ... ∪ Cn. One can check that the proof only uses One-Sided Additivity, so it holds in our case. It follows from G-invariance that A1 ∪ ... ∪ Am ∼ E ∼ Am + 1 ∪ ... ∪ An. Since E is the disjoint union of A1 ∪ ... ∪ Am with Am + 1 ∪ ... ∪ An, this violates the Weak Parthood Principle.

Monday, October 14, 2024

Defining comparative probabilities in terms of conditional probabilities

Suppose we have a full conditional probability P(AB) defined for all pairs of events (stipulating that P(A∣⌀) = 1 if we wish). I've proposed two methods for defining a probability comparison using conditional probabilities:

  1. A ⪅ B iff P(AAB) ≤ P(BAB).

  2. A ⪅ B iff P(ABAΔB) ≤ P(BAAΔB), where AΔB = (AB) ∪ (BA) is the symmetric difference.

In a footnote in a paper, I wrote about the second ordering, which I incorrectly attributed to de Finetti: “This ordering has the advantage that if A is a proper subset of B, then A < B, but it is somewhat harder to prove transitivity”.

Well, that was an understatement! It’s not just harder to prove transitivity: it’s impossible.

Define:

  • Ω: all integers

  • E0: non-negative even integers

  • E: positive even integers

  • D: positive odd integers.

Let P be a full conditional probability such that:

  1. P(E0E0D) = 1/2 = P(DE0D)

  2. P(DDE) = 1/2 = P(E|DE).

It is easy to see from (3) and (4) that because E0 and D are disjoint, and so are D and E, then by definition (2) we have E0 ⪅ D and D ⪅ E. (For disjoint A and B, the definition (2) of A ⪅ B is equivalent to thedefinition (1).) However, E0 − E = {0}, E − E0 = ⌀, and E0ΔE = {0}, so P(EE0E0ΔE) = 0 while P(E0EE0ΔE) = 1, and thus we cannot have E0 ⪅ E.

The only question is whether there actually is a full conditional probability satisfying (3) and (4). If there is, then (2) is not transitive in our case.

There is such a full conditional probability. Let Qn(A) =  ∣ A ∩ [−n,n] ∣ /(2n+1), where $B$ is the cardinality of a set B. Then Qn is a probability. Let Q be a limit of the Qn along an ultrafilter. This is a finitely additive hyperreal probability which is non-zero for all non-empty sets. Define P(AB) as the standard part of Q(AB)/Q(B) for B non-empty. This is a full conditional probability. Moreover, P(AB) = lim Qn(AB)/Qn(B) whenever the latter limit is defined. That limit is defined in the cases of the events involved in (3) and (4), and it is easy to evaluate the limits and see that (3) and (4) are true.

I don’t know what I was thinking when I wrote that footnote. My guess is that I had in my mind a proof sketch that doesn’t work (I have some idea what that might have been).

Whew! I noticed this afternoon that Theorem 1 of this paper of mine was incompatible with the transitivity of comparison (2). This made me really worried that my Theorem 1 was false. But since the comparison (2) isn’t transitive, I can relax.

This raises an interesting potential research problem. The Pruss definition of ⪅ does not satisfy the additivity axiom for comparative probabilities, namely that if C is disjoint from A ∪ B, then A ≲ B if and only if A ∪ C ≲ B ∪ C (it only preserves the left-to-right implication). Definition (2) does satisfy the additivity axiom, which is what I liked about it.

I suspect there is no good definition of comparative probabilities in terms of full conditional probabilities that satisfies the additivity axiom. (One reason for this intuition has to do with the fact that in Figure 1 here there are entries with a Yes in column 3 and a No in column 5.)

So, I now wonder: Is there some good combination of a definition of comparative probabilities in terms of full conditional probabilities with some weakened version of the additivity axiom?

Thursday, October 28, 2021

Symmetric qualitative (and other) probabilities

I recently worked out the precise conditions under which one can have Popper functions, hyperreal probabilities or qualitative probabilities that are invariant under some group of symmetries and are regular in the sense that they assign a bigger probability to non-empty sets than to the empty set.

But what if we don’t require regularity? Then the following is mainly a matter of putting together known theorems:

Proposition. Suppose G is a group acting on set Ω* ⊇ Ω where Ω is non-empty. Then the following are equivalent:

  1. There is a finitely additive G-invariant real-valued probability measure on the powerset of Ω

  2. There is a finitely additive G-invariant hyperreal probability measure on the powerset of Ω

  3. There is a finitely additive approximately G-invariant hyperreal probability measure on the powerset of Ω

  4. There is a strongly G-invariant total qualitative probability ⪅ on the powerset of Ω such that ⌀ < Ω

  5. There is a strongly G-invariant partial qualitative probability ⪅ on the powerset of Ω such that ⌀ < Ω

  6. The set Ω is not G-paradoxical.

The definitions are in the paper I linked to at the top, except that approximate G-invariance only requires that P(A)−P(gA) be infinitesimal rather than requiring that it be zero.

Proof: Trivially, (a) implies (b) which implies (c). The standard part of a finitely additive approximately G-invariant hyperreal probability measure will be a finitely additive G-invariant real-valued probability measure, so (c) implies (a). Thus, (a)–(c) are equivalent.

Condition (a) implies condition (d): just define A ⪅ B iff P(A)≤P(B) where P is the measure in (a). And (d) implies (e) trivially.

Now we show that not-(f) implies not-(e). Suppose Ω is G-paradoxical, so Ω has disjoint subsets A and B with partitions A1, ..., Am and B1, ..., Bn respectively, and there are elements g1, ..., gm and h1, ..., hn of G such that g1A1, ..., gmAm and h1B1, ..., hnBn are each a partition of Ω. Then by a standard result on qualitative probabilities (use the proof of Krantz, et al., Lemma 5.3.1.2):

  1. A = A1 ∪ ... ∪ Am ≈ g1A1 ∪ ... ∪ gmAm = Ω

  2. B = B1 ∪ ... ∪ Bn ≈ h1B1 ∪ ... ∪ hnBn = Ω.

Since ⌀ < Ω, we have ⌀ < A by (1). By the proof of Corollary 5.3.1.2 in Krantz, et al., we have B < Ω iff ⌀ < Ω − B. But A ⊆ Ω − B, and ⌀ < A, so indeed we must have B < Ω, which contradicts (2).

Finally, Tarski’s Theorem says that (f) implies (a). □

Note 1: The two results from Krantz et al. are given for total qualitative probabilities, but the proofs do not use totality. (In the linked paper, I didn’t notice that Krantz et al. are working with total qualitative probabilities, but fortunately all works out.)

Note 2: There is a pleasing direct construction of a partial qualitative probability satisfying (d). For each A ⊆ Ω, let [A] be the corresponding member of the equidecomposability type semigroup. Then define A ⪅ B providing there is a c in the semigroup such that [A]+c ≤ [B]+c. It turns out that the condition ⌀ < Ω is then equivalent to 2[Ω]≠[Ω], i.e., is equivalent to the non-paradoxicality of Ω under G.

Thursday, October 22, 2020

Preprint: Conditional, Regular Hyperreal and Regular Qualitative Probabilities Invariant Under Symmetries

Abstract: Classical countably additive real-valued probabilities come at a philosophical cost: in many infinite situations, they assign the same probability value---namely, zero---to cases that are impossible as well as to cases that are possible. There are three non-classical approaches to probability that can avoid this drawback: full conditional probabilities, qualitative probabilities and hyperreal probabilities. These approaches have been criticized for failing to preserve intuitive symmetries that can easily be preserved by the classical probability framework, but there has not been a systematic study of the conditions under which these symmetries can and cannot be preserved. This paper fills that gap by giving complete characterizations under which symmetries understood in a certain "strong" way can be preserved by these non-classical probabilities, as well as by offering some results to make it plausible that the strong notion of symmetry here may be the right one. Philosophical implications are briefly discussed, but the main purpose of the paper is to offer technical results to inform more sophisticated further philosophical discussion.

Preprint here.

Tuesday, September 8, 2020

The fairness of infinite lotteries and qualitative probabilities

Suppose that we wish to model an infinite fair lottery with tickets numbered by integers by means of qualitative probabilities, i.e., a reflexive and transitive relation ≲ between sets of tickets that satisfies the non-negativity constraint that ∅ ≲ A for all A and the additivity constraint that A ≲ B iff A − B ≲ B − A. Suppose, further, that we want to have the regularity constraint that ∅ < A if A is not empty.

At this point, we want to ask what “fairness” is. One proposal is that fairness is strong translation invariance: if A is a set of integers and n + A is the set {n + m : m ∈ A} of all the members of A shifted over by m, then A and n + A are equally probable. Unfortunately, if we require strong translation invariance, then we violate the regularity constraint, since we will have to assign the same probability to the winning ticket being in {1, 2, 3, ...} as to the winning ticket being in {2, ...}, which (given additivity) violates the constraint that ∅ < {1}.

One possible option that I’ve been thinking about is is to require weak translation invariance. Weak translation invariance says that A ≲ B iff n + A ≲ n + B. Thus, a set might not have the same probability as a shift of itself, but comparisons between sets are not changed by shifts. I’ve spent a good chunk of the last week or two trying to figure out whether (given the other constraints) it is coherent to require weak translation invariance. Last night, Harry West gave an elegant affirmative proof on MathOverflow. So, yes, one can require weak translation invariance.

However, weak translation invariance does not capture the concept of fairness. Here is one reason why.

Say that a set B of integers is right-to-left (RTL) bigger than a set A of integers provided that there is an integer n such that:

  1. n ∈ B but not n ∈ A, and

  2. for every m > n, if m ∈ A, then m ∈ B.

RTL comparison of sets of integers thus always favors sets with larger integers. Thus, the set {2, 3} is RTL bigger than the infinite set {..., − 3, −2, −1, 0, 1, 3}, because the former set has 2 in it while the latter does not.

It looks to me that West’s proof straightforwardly adapts to show that that there is a weakly translation invariant qualitative probability that coheres with RTL ordering: if B is RTL bigger than A, then B is strictly more likely than A. But a probability comparison that coheres with RTL ordering is about as far from fairness as we can imagine: a bigger ticket number is always more likely than a smaller one, and indeed each ticket number is more likely to be the winner than the disjunction of all the smaller ticket numbers!

So, weak translation invariance doesn’t capture the concept of fairness.

Here is a natural suggestion. Let’s add to weak translation invariance the following constraint: any two tickets are equally likely.

I think—but here I need to check more details—that a variant of West’s proof again shows that this won’t do. Say that a set B of integers is right-skewed (RS) at least as big as a set A of integers provided that one or more of the following holds:

  1. A is finite and B has at least as many members than A, or

  2. B has infinitely many positive integers and A does not, or

  3. A is a subset of B.

Intuitively, a probability ordering that coheres with RS ordering fails to be fair, because, for instance, it makes it more likely that the winning ticket will be, say, a power of two than that it be a negative number. But at the same time, a probability ordering that coheres with RS ordering makes all individual tickets be equally likely by (1).

To make this work with West’s proof, replace his C0 with the set of bounded functions that have a well-defined and non-negative sum or whose positive part has an infinite sum.

Monday, July 27, 2020

Loading infinitely many coins

Suppose I have infinitely many coins and I toss them. I then load each of them equally in favor of heads and toss them all again. All the tosses are independent.

Trick question: Was it more likely that they all landed tails on the second toss or on the first?

Although it seems more likely on the first, the correct answer has to be neither. Here’s one way to see it. Let’s imagine for simplicity that the second set of coins is physically distinct from the first, and that all the coins—the loaded and the unloaded—are tossed simultaneously. That shouldn’t make any difference to the comparison. But now let’s suppose that the loaded coins are such that the probability of heads is 3/4 for each coin. Now, line up the coins in such a way that two fair coins are put beside each loaded one. The probability that both of the fair coins land tails is 1/4. The probability that the loaded coin lands tails is 1/4. To get all the fair coins landing tails thus shouldn’t be any more likely than to get all the loaded coins landing tailings, or vice versa. And if the loading is different, we just tweak the arrangement.

So, either the event of all the fair coins landing tails is equal in probability with the event of all the loaded coins landing tails, or the two events cannot be compared in probability.

Tuesday, July 21, 2020

Do Popper functions carry enough information?

Let P be a Popper function on some algebra F of events on Ω. There is a natural way to define a probability comparison given P, namely A ⪅ B iff P(A|A ∪ B)≤P(B|A ∪ B), which I think I’ve used before. Unfortunately, this can violate a common axiom of probability comparisons, namely that A ⪅ B iff Ω − B ⪅ Ω − A.

For instance, consider the Popper function generated by a regular hyperreal probability on [0, 1] whose standard part agrees with Lebesgue measure on intervals. Then [0, 1]⪅[0, 1), since both have conditional probability 1 on [0, 1]∪[0, 1)=[0, 1]. But their complements are ∅ and {1}, respectively, and {1}⪅∅ is false.

This little fact may have some significance. There is a theorem in the literature that shows a correlation between Popper function and hyperreal probabilities: every Popper function with all non-empty sets regular can be generated from a normal hyperreal probability using the conditional probability formula, and given a Popper function with all non-empty sets regular there is a normal hyperreal probability from which the Popper function can be generated. This has led to a debate whether it’s better to work with Popper functions or hyperreal probabilities. One argument for working with Popper functions—an argument I’ve approved of in print—is that the hyperreal probabilities carry more information than reality does. I still think that this problem is there for hyperreal probabilities. But the above argument suggests that going to a Popper function may have discarded too much information. Popper functions allow for fine-grained comparisons of “small” events, such as singletons, but not for the complements of these events.

Friday, July 17, 2020

Arbitrariness, regularity and comparative probabilities

In “Underdetermination of infinitesimal probabilities”, following a referee’s suggestion, I allow that qualitative (i.e., comparative) probabilities may escape arbitrariness problems for infinitesimal probabilities. But I now think this may be wrong.

Consider an infinite line of independent fair coins, numbered with the integers ..., −2, −1, 0, 1, 2, .... Let Rn be the hypothesis that coins n, n + 1, ... are all heads. Let Ln be the hypothesis that coins ..., n − 2, n − 1, n are all heads.

Suppose “being less likely or equally likely” is transitive, reflexive and total. Write A < B for A being less likely than B, A ≈ B for A and B being equally likely, and A ⪅ B for A being less likely or equally like than B.

We have strong regularity provided that if event A is a proper subset of event B, then A is strictly less likely than B.

Given strong regularity, the events Ln are strictly decreasing in probability: ... > L−2 > L−1 > L0 > L1 > ... and the events Rn are strictly increasing in probability: ... < R−2 < R1 < R0 < R1 < ....

Theorem: Given strong regularity, exactly one of the following options is true:

  1. For all n and m, Ln < Rm (heads-runs are right-biased)

  2. For all n and m, Rm < Ln (heads-runs are left-biased)

  3. There is a unique n such that for all m ≤ n we have Rm ⪅ Lm and for all m > n we have Lm < Rm (there is a switch-over point at n).

But if (1) or (2) is true, it is difficult to see what objective reality could possibly ground whether heads-runs are right-biased or left-biased in our infinite sequence of coin tosses. And if (3) is true, it is difficult to see what objective reality could possibly make n be a switch-over point. The choice between left- and right-bias seems completely arbitrary, and the choice of a switch-over point is also arbitrary.

The Theorem follows from the following lemma:

Lemma: Let (S, ⪅) be a totally preordered set, and let Ln and Rn be sequences of members of S as n ranges over the integers. Suppose Ln is strictly decreasing and Rn is strictly increasing. Then exactly one of the following is true:

  1. For all n and m, Ln < Rm

  2. For all n and m, Rm < Ln

  3. There is a unique n such that for all m ≤ n we have Rm ≤ Lm and for all m > n we have Lm < Rm.

Proof of Lemma: Suppose that for all n we have Ln < Rn. I claim that (1) is true. For suppose that (1) is false and hence (by totality) we have Rm ⪅ Ln for some m and n. Then m ≠ n, since Ln < Rn. Either m < n or n < m. If m < n, then Rm ⪅ Ln < Lm, and if m > n, then Rn < Rm ⪅ Ln. In either case we have a violation of the fact that Ln < Rn for all n′.

Now suppose that for all n we have Rn < Ln. I now claim that (2) is true. For if (2) is false, we have Ln ⪅ Rm for some m ≠ n. Suppose m < n. Then Ln ⪅ Rm < Rn, a contradiction. And if n < m, then Lm < Ln ⪅ Rm, also a contradiction.

Let A(n) be the statement that for all m ≤ n we have Rm ⪅ Lm and for all m > n we have Lm < Rm. There is at most one n such that A(n). For suppose that we had A(n) and A(n′) and that n < n′. Then by A(n) we have Ln < Rn (since n′>n), and by A(n′) we have Rn ⪅ Ln, resulting in a contradiction.

Assume (1) and (2) are false. By what we have shown earlier and totality, if (2) is not true, there is an n such that Ln ⪅ Rn. Note that if Ln ⪅ Rn, then Lm < Ln ⪅ Rn < Rm for all m > n. Hence, either Ln ⪅ Rn is true for all n or else there is a smallest n for which it’s true. If it’s true for all n, then for all n we have Ln + 1 < Ln ⪅ Rn < Rn + 1, and hence for all n we have Ln < Rn, which we saw would imply (1), which we assumed to be false.

So suppose that n is the smallest integer for which Ln ⪅ Rn. Thus, Rm < Lm whenever m < n. Moreover, since Li is strictly decreasing and Ri is strictly increasing, we have Lm < Rn whenever m > n. There are now two possibilities. Either Ln < Rn or Ln ≈ Rn. If Ln ≈ Rn, then we have A(n) and the proof is complete. Suppose now that Ln < Rn. Then we have A(n − 1) and the proof is also complete.

Symmetry, regularity and qualitative probability

Let ⪅ be a qualitative probability comparison for some collection F of subsets of a space Ω. Say that A ≈ B iff A ⪅ B and B ⪅ A, and that A < B provided that A ⪅ B but not B ⪅ A. Minimally suppose that ⪅ is a partial preorder (i.e., transitive and reflexive). Say it’s total provided that for all A and B either A ⪅ B or B ⪅ A. Suppose that G is a group of symmetries acting on Ω, and that F is G-invariant in the sense that gA ∈ F for all g ∈ G. Then we can define:

  1. ⪅ is strongly G-invariant provided that for all A in G and all g in G we have A ≈ gA, and

  2. ⪅ is weakly G-invariant provided that for all A and B in G and all g in G we have A ⪅ B iff gA ⪅ gB.

There is some reason to be suspicious of strong G-invariance. For in some interesting cases, say where Ω is a circle that G is the set of all rotations, there will be cases where gA is a proper subset of A, and by regularity we would expect to have gA < A rather than gA ≈ A. But weak G-invariance seems harder to question.

Say that g ∈ G is of order n provided that gn = e, where e is identity. However, we also have:

Lemma 1. If ⪅ is total, g is of order 2, and A ⪅ B implies gA ⪅ gB for all A and B, then A ≈ gA for all A.

Proof: Since ⪅ is a total order, either A ⪅ gA or gA ⪅ A. Suppose A ⪅ gA. Then gA ⪅ g2A. But g2A = A. Hence gA ≈ A. Similarly, if gA ⪅ A, then g2A ≈ gA. But g2A = A, so gA ≈ A.

So, we have:

Proposition 1. If ⪅ is total and G is a group generated by elements of order 2, then weak G-invariance entails strong G-invariance.

Say that ⪅ is strongly regular provided that if A is a proper subset of B, then A < B. Weak regularity would say that if B is non-empty then ∅ < B. Weak regularity together with an appropriate additivity condition will imply strong regularity (details left to the reader).

Proposition 2. If G is generated by elements of order 2, and ⪅ is total and weakly G-invariant, then if there exists a g ∈ G and A ∈ F such that gA is a proper subset of A, then G is not strongly regular.

Proof: Strong regularity would require that gA < A, but that would contradict strong G-invariance which we have by Proposition 1.

Corollary 1. If F is a collection of subsets of the unit circle containing all countable sets and invariant under all reflections, and ⪅ is a total qualitative probability comparison weakly invariant under all reflections, then ⪅ is not strongly regular.

Proof: The group generated by all reflections includes all rotations. But there is a subset A of the circle and a rotation g such that gA is a proper subset of A. For instance, let A be the set of points at angles r, 2r, 3r, ... in degrees, where r is irrational, and let g be rotation by r. Then rA is the set of points at angles 2r, 3r, 4r, ... in degrees.

Now, imagine an infinite line on which there are infinitely many evenly spaced people, stretching out in both directions, each of whom flips a fair coin. Let Ω be the probability space describing these flips. Let Hn be the event that all the flips starting with person number n (i.e., n, n + 1, n + 2, ...) land heads. Suppose that F contains all the Hn and is invariant under all reflections of the situation (where we reflect the setup either about the point at which some person stands or at a point half-way between two neighboring people).

Corollary 2. If ⪅ is a total qualitative probability comparison weakly invariant under all reflections, then ⪅ is not strongly regular.

Proof: Let G be the group generated by the reflections. This group contains all translations. A non-trivial translation of Hn will either be a proper subset or a proper superset of Hn, depending on the direction of Hn. So by Proposition 2 we cannot have regularity.

Bibliographic note: Lemma 1 and Corollary 2 are analogous to Lemma 2 and Theorem 4 of this paper.

Thursday, July 16, 2020

The choice between qualitative probabilities and generalized quantitative probabilities is illusory

There are two approaches to generalizing probabilities beyond the classical real-valued numerical approach.

  1. Switch the values of the probability function from real numbers to values in some other ordered algebraic entity (e.g., the hyperreals, the surreals, an arbitrary totally ordered monoid).

  2. Switch to qualitative probabilities, where instead of assigning values to events, one just compares the events probabilistically (this one is at least as likely than that one).

In the literature (including stuff I wrote myself!), the qualitative probability approach is treated as more general. But in fact, it’s not, at least not if we assume that the “at least as likely as” relation is transitive and reflexive, i.e., is a partial preorder. For suppose that we have a collection F of events and a partial preorder ⪅ on them. Say that A ≈ B iff A ⪅ B and B ⪅ A. Let V be the set of equivalence classes of events under the relation ≈, and define the partial order ≤ on V by [A]≤[B] iff A ⪅ B, where [A] is the equivalence class of A. (All this makes sense if ⪅ is a partial preorder.) Define P(A)=[A].

And that’s it! You now have a probability function P on F whose values are a partially ordered set V. So, what you do with qualitative comparisons you can do with values. Now that I’ve said it, it’s obvious to me and I’m kicking myself for not noticing it earlier. Perhaps some people in the field have noticed it and found it so obvious that it’s not worth saying.

And rather than thinking of there being a debate between probability functions with values and qualitative probabilities, we can now just stick to the probability function approach, and say that there is a serious debate as to what kind of a structure the values have: are they a real closed field, a totally ordered field, a totally ordered monoid (and if so, does its operation correspond to addition or to multiplication in the classical case), a mere partially ordered set, etc.?

The point here applies in other contexts, such as qualitative utilities, moral value comparisons, etc. Using the apparatus of set theory, we can replace a comparison relation by a value-assignment, which makes it more convenient to apply all the apparatus of ordered algebraic entities of different sorts.

As an application, in a recent post I said that given two very plausible axioms on probability functions, one cannot have a regular rotationally-invariant probability function on the measurable subsets of the circle. Using the above remarks, that point immediately extends to qualitative probabilities that satisfy the following two axioms:

  1. If A and B are disjoint, A and C are disjoint, and B and C are equally likely, then A ∪ B and A ∪ C are equally likely, and

  2. Ω − A and Ω − B are equally likely iff A and B are equally likely,

with rotational invariance being understood as saying that A is at least as likely as B if and only if ρA is at least as likely as B if and only if A is at least as likely as ρB for every rotation ρ, and with regularity understood as saying that every non-empty event is more likely than the empty event.

Saturday, August 10, 2013

Normal Popper functions and comparative probability

Let Ω be a non-empty set and F be a field of subsets (i.e., set of subsets closed under finite unions and complements). A Popper function on F is a real-valued function P defined for pairs of members of F such that:

  1. 0≤P(X|Y)≤P(Y|Y)=1
  2. if P(Ω−Y|Y)<1, then P(−|Y) is a finitely additive probability on F
  3. P(XY|Z)=P(X|Z)P(Y|XZ)
  4. if P(X|Y)=P(Y|X)=1, then P(Z|X)=P(Z|Y).
(This is van Fraassen's axiomatization.) A member B of F is normal provided that P(Ω−B|B)<1. (This condition is equivalent to saying that P(∅|B)<1.) A Popper function is normal provided that every non-empty member of F is normal.

On the other hand, a comparative probability function on F (Paul Bartha has talked about things like this) is a function from pairs of members of F to [0,∞] such that:

  1. C(X,X)=1
  2. C(X,Y)C(Y,Z)=C(X,Z) provided the left-hand side is defined (0 times ∞ and ∞ times zero are the undefined cases)
  3. C(−,Y) is a finitely-additive measure if Y is non-empty.

Alan Hajek, I think, has suggested that one can define a comparative probability function in terms of a Popper function. We can do it as follows. If P is a normal Popper function, then let CP(X,Y)=P(X|XY)/P(Y|XY). And given a comparative probability function, we can define a normal Popper function by PC(X|Y)=CP(XY|Y).

I haven't written out the details, but it looks like we then have:

Proposition. If P is a normal Popper function, then CP is a comparative probability function, and if C is a comparative probability function, then PC is a normal Popper function. Moreover, if P is a normal Popper function and C=CP, then PC=P, while if C is a comparative probability function and P=PC, then CP=C.

Thus, there is a nice one-to-one correspondence between comparative probabilities and normal Popper functions.

Personally, I find comparative probabilities to be easier to prove theorems about than Popper functions, because I find (6) much easier to remember than (3).