Showing posts with label qualitative probability. Show all posts
Showing posts with label qualitative probability. Show all posts

Monday, October 14, 2024

Defining comparative probabilities in terms of conditional probabilities

Suppose we have a full conditional probability P(AB) defined for all pairs of events (stipulating that P(A∣⌀) = 1 if we wish). I've proposed two methods for defining a probability comparison using conditional probabilities:

  1. A ⪅ B iff P(AAB) ≤ P(BAB).

  2. A ⪅ B iff P(ABAΔB) ≤ P(BAAΔB), where AΔB = (AB) ∪ (BA) is the symmetric difference.

In a footnote in a paper, I wrote about the second ordering, which I incorrectly attributed to de Finetti: “This ordering has the advantage that if A is a proper subset of B, then A < B, but it is somewhat harder to prove transitivity”.

Well, that was an understatement! It’s not just harder to prove transitivity: it’s impossible.

Define:

  • Ω: all integers

  • E0: non-negative even integers

  • E: positive even integers

  • D: positive odd integers.

Let P be a full conditional probability such that:

  1. P(E0E0D) = 1/2 = P(DE0D)

  2. P(DDE) = 1/2 = P(E|DE).

It is easy to see from (3) and (4) that because E0 and D are disjoint, and so are D and E, then by definition (2) we have E0 ⪅ D and D ⪅ E. (For disjoint A and B, the definition (2) of A ⪅ B is equivalent to thedefinition (1).) However, E0 − E = {0}, E − E0 = ⌀, and E0ΔE = {0}, so P(EE0E0ΔE) = 0 while P(E0EE0ΔE) = 1, and thus we cannot have E0 ⪅ E.

The only question is whether there actually is a full conditional probability satisfying (3) and (4). If there is, then (2) is not transitive in our case.

There is such a full conditional probability. Let Qn(A) =  ∣ A ∩ [−n,n] ∣ /(2n+1), where $B$ is the cardinality of a set B. Then Qn is a probability. Let Q be a limit of the Qn along an ultrafilter. This is a finitely additive hyperreal probability which is non-zero for all non-empty sets. Define P(AB) as the standard part of Q(AB)/Q(B) for B non-empty. This is a full conditional probability. Moreover, P(AB) = lim Qn(AB)/Qn(B) whenever the latter limit is defined. That limit is defined in the cases of the events involved in (3) and (4), and it is easy to evaluate the limits and see that (3) and (4) are true.

I don’t know what I was thinking when I wrote that footnote. My guess is that I had in my mind a proof sketch that doesn’t work (I have some idea what that might have been).

Whew! I noticed this afternoon that Theorem 1 of this paper of mine was incompatible with the transitivity of comparison (2). This made me really worried that my Theorem 1 was false. But since the comparison (2) isn’t transitive, I can relax.

This raises an interesting potential research problem. The Pruss definition of ⪅ does not satisfy the additivity axiom for comparative probabilities, namely that if C is disjoint from A ∪ B, then A ≲ B if and only if A ∪ C ≲ B ∪ C (it only preserves the left-to-right implication). Definition (2) does satisfy the additivity axiom, which is what I liked about it.

I suspect there is no good definition of comparative probabilities in terms of full conditional probabilities that satisfies the additivity axiom. (One reason for this intuition has to do with the fact that in Figure 1 here there are entries with a Yes in column 3 and a No in column 5.)

So, I now wonder: Is there some good combination of a definition of comparative probabilities in terms of full conditional probabilities with some weakened version of the additivity axiom?

Wednesday, April 12, 2023

Generating qualitative probabilities

A (partial) qualitative probability is a relation defined on an algebra of sets and satisfying the following axioms:

  1. Preorder: ≾ is reflexive and transitive

  2. Zero: ⌀ ≾ A

  3. Additivity: if A ∪ B and C are disjoint, then A ≾ B if and only A ∪ C ≾ B ∪ C.

It’s just occurred to me that there is a nice way to construct qualitative probabilities out of a family of finitely additive probabilities. Fix an algebra F of subsets of Ω. Let Q be a non-empty set of finitely additive probability functions on F taking values in totally ordered fields. Then say that AQB if and only if p(A) ≤ p(B) for all p in Q. This clearly satisfies the three axioms above.

This may seem like a major extension of the concept of a total qualitative probability p generated by a single probability function p, but it’s not as much an extension as it may seem.

Remark 1: Not every qualitative probability, not even every total qualitative probability, can be constructed as Q for some Q.

At least one 1959 Kraft, Pratt and Seidenberg counterexample to the thesis that every total qualitative probability is generated by a single real-valued probability function trivially extends to show this.

Now let ≺ be the strict order generated by ≾ (and similarly with subscripts): A ≺ B iff A ≾ B but not B ≾ A.

Theorem 1: Assume Choice. Given a non-empty set Q of finitely additive probability functions on F there is a probability function p taking values in some ordered field such that ≾p extends ≾Q and ≺p extends Q.

Sketch of proof: All the ordered fields that the members of Q take values in can be embedded in the surreals. There will be a set-sized field that contains the ranges of all the embeddings. So we can assume that all the members of Q take values in a single field. The rest follows by an ultrafilter construction. Let H be the set of non-empty finite subsets of Q ordered by inclusion. Given K in H, let pK be the sum of the members of K. Then let p be the ultraproduct of the pK with respect to some ultrafilter on H. Verifying finite additivity and the fact that ≾p extends ≾Q is trivial. Verifying that p extends Q is only slightly harder. Suppose AQB. Then AQB. For some q in Q we must have q(A) < q(B). Then for any K in H containing {q}, we have pK(A) < pK(B), and so p(A) < p(B).

Theorem 2: Assume Choice. Given a non-empty set Q of finitely additive real-valued probability functions on F there is a real-valued probability function p such that p extends Q.

Sketch of proof: Let H be the set of non-empty finite subsets of Q ordered by inclusion. For K in H, let pK be the sum of the members of K. By compactness, pK has a limit point p.

Remark 2: One cannot require that p extend Q in Theorem 2. For let Ω have cardinality greater than the continuum. Then there is no regular real-valued finitely additive probability on the powerset of Ω (a probability is regular if P(A) > 0 for every non-empty A), since if there were, then each Ωz = {x ∈ Ω : x < z} would have different probabilities, and so the probability would have more values than the continuum. Let Q be all finitely additive real-valued probabilities on the powerset of Ω. Then ⌀≺QA for any non-empty A (since 0 < q(A) for q concentrated on some point of A). But if we had ⌀≺pA, then p would be regular. I am not sure what to say if Ω has continuum cardinality.

Friday, November 18, 2022

Social choice principles and invariance under symmetries

A comment by a referee of a recent paper of mine that one of my results in decision theory didn’t actually depend on numerical probabilities and hence could extend to social choice principles made me realize that this may be true for some other things I’ve done.

For instance, in the past I’ve proved theorems on qualitative probabilities. A qualitative probability is a relation on the subsets of some sample space Ω such that:

  1. ≼ is transitive and reflexive.

  2. ⌀ ≼ A

  3. if A ∩ C = B ∩ C = ⌀, then A ≼ B iff A ∩ C ≼ B ∩ C (additivity).

But need not think of Ω as a space of possibilities and of ≼ as a probability comparison. We could instead think of it as a set of people who are candidates for getting some good thing, with A ≼ B meaning that it’s at least as good for the good thing to be distributed to the members of B as to the members of A. Axioms (1) and (2) are then obvious. And axiom (3) is an independence axiom: whether it is at least as good to give the good thing to the members of B as to the members of A doesn’t depend on whether we give it to the members of a disjoint set C at the same time.

Of course, for a general social choice principle we need more than just a decision whether to give one and the same good to the members of some set. But we can still formalize those questions in terms of something pretty close to qualitative probabilities. For a general framework, suppose a population set X (a set of people or places in spacetime or some other sites of value) and a set of values V (this could be a set of types of good, or the set of real numbers representing values). We will suppose that V comes with a transitive and reflexive (preorder) preference relation . Now let Ω = X × V. A value distribution is a function f from X to V, where f(x) = v means that x gets something of value v.

We want to generate a reflexive and transitive preference ordering ≼ on the set VX of value distributions.

Write f ≈ g when f ≼ g and g ≼ f, and f ≺ g when f ≼ g but not g ≼ f. Similarly for values v and w, write v < w if v ≤ w but not w ≤ v.

Here is a plausible axiom on value distributions:

  1. Sameness independence: if f1, f2, g1, g2 are value distributions and A ⊆ X is such that (a) f1 ≼ f2, (b) f1(x) = f2(x) and g1(x) = g2(x) if x ∉ A, (c) f1(x) = g1(x) and f2(x) = g2(x) if x ∈ A.

In other words, the mutual ranking between two value distributions does not depend on what the two distributions do to the people on whom the distributions agree. If it’s better to give $4 to Jones than to give $2 to Smith when Kowalski is getting $7, it’s still better to give $4 to Jones than to give $2 to Smith when Kowalski is getting $3. There is probably some other name in the literature for this property, but I know next to nothing about social choice literature.

Finally, we want to have some sort of symmetries on the population. The most radical would be that the value distributions don’t care about permutations of people, but more moderate symmetries may be required. For this we need a group G of permutations acting on X.

  1. Strong G-invariance: if g ∈ G and f is a value distribution, then f ∘ g ≈ f.

Here, f ∘ g is the value distribution where site x gets f(g(x)).

Additionally, the following is plausible:

  1. Pareto: If f(x) ≤ g(x) for all x with f(x) < g(x) for some x, then f ≺ g.

Theorem: Assume the Axiom of Choice. Suppose on V is reflexive, transitive and non-trivial in the sense that it contains two values v and w such that v < w. There exists a reflexive, transitive preference ordering on the value distributions satisfying (4)–(6) if and only if there is such an ordering that is total if and only if G has locally finite action on X.

A group of symmetries G has locally finite action a set X provided that for each finite subset H of G and each x ∈ X, applying finite combinations of members of G to x generates only a finite subset of X. (More precisely, if ⟨H⟩ is the subgroup generated by G, then Hx is finite.)

If X is finite, then local finiteness of action is trivial. If X is infinite, then it will be satisfies in some cases but not others. For instance, it will be satisfied if G is permutations that only move a finite number of members of X at a time. It will on the other hand fail if X is a infinite bunch of people regularly spaced in a line and G is shifts.

The trick to the proof of the Theorem is to reduce preferences between distributions to comparisons of subsets of X × V and to reduce comparisons of subsets of X to preferences between binary distributions.

Proof of Therem: Suppose that G has locally finite action. Define Ω = X × V. By Theorem 2 of my invariance of non-classical probabilities paper, there is a strongly G-invariant regular (i.e., ⌀ ≺ A if A is non-empty) qualitative probability ≼ on Ω. Given a value distribution f, let f* = {(x,v) : v ≤ f(x)} be a subset of Ω. Define f ≼ g iff f* ≼ g.

Totality, reflexivity, transitivity and strong G-invariance for value distributions follows from the same conditions for subsets of Ω. Regularity of on the subsets of Ω and additivity implies that if A ⊂ B then A ≺ B. The Pareto condition for ≼ on the value distributions follows since if f and g satisfy are such that f(x) ≤ g(x) for all x with strict inequality for some x, then f* ⊂ g*. Finally, the complicated sameness independence condition follows from additivity.

Now suppose there is a (not necessarily total) strongly G-invariant reflexive and transitive preference ordering ≼ on the value distributions satisfying (4)–(6). Given a subset A of X, define A to be the value distribution that gives w to all the members of A and v to all the non-members, where v < w. Define A ≼ B iff A ≼ B. This will be a strongly G-invariant reflexive and transitive relation on the subsets of X. It will be regular by the Pareto condition. Finally, additivity follows from the sameness independence condition. Local finiteness of action of G then follows from Theorem 2 of my paper. ⋄

Note that while it is natural to think of X has just a set of people or of locations, inspired by Kenny Easwaran one can also think of it as a set Q × Ω where Ω is a probability space and Q is a population, so that f(x,ω) represents the value x gets at location ω. In that case, G might be defined by symmetries of the population and/or symmetries of the probability space. In such a setting, we might want a weaker Pareto principle that supposes additionally that f(x,ω) < g(x,ω) for some x and all ω. With that weaker Pareto principle, the proof that the existence of a G-invariant preference of the right sort on the distributions implies local finiteness of action does not work. However, I think we can still prove local finiteness of action in that case if the symmetries in G act only on the population (i.e., for all x and ω there is an y such that g(x,ω) = (y,ω)). In that case, given a subset A of the population Q, we define A to be the distribution that gives w to all the persons in A with certainty (i.e., everywhere on Ω) and gives v to everyone else, and the rest of the proof should go through, but I haven’t checked the details.