Showing posts with label Popper functions. Show all posts
Showing posts with label Popper functions. Show all posts

Tuesday, April 14, 2026

A problem with perfectly rational agents and decision theory

Suppose I am perfectly rational in the decision theoretic sense. A coin is about to be tossed, and I will get five dollars on heads (H) and one dollar on tails (T). I have a choice whether to leave the coin fair (F) or load it (L) in favor of tails so that the probability of tails is 3/4.

It is obvious what I do. I calculate the expected utilities of my options F and L as follows.

  • EU(F) = P(H|F) ⋅ $5 + P(T|F) ⋅ $1 = (1/2) ⋅ $5 + (1/2) ⋅ $1 = $3

and

  • EU(L) = P(H|L) ⋅ $5 + P(T|L) ⋅ $1 = (1/4) ⋅ $5 + (3/4) ⋅ $1 = $2.

And then I choose F.

Except it’s not so simple. For I am perfectly rational. But since, as we just saw, the perfectly rational agent has to choose F, it follows that P(L) = 0, and so P(H|L) and P(T|L) are undefined. So I can’t decide! So now there is no guarantee how I will act, and P(H|L) and P(T|L) once again make sense. And then again they don’t. Oops!

What can be done? Causal decision theorists will note that I reasoned like an evidential decision theorist above. But this makes no difference in this case. The causalist’s story will be a bit more complicated but will end up with the same problem.

We might want to introduce primitive conditional probabilities like Popper functions that let you conditionalize on events with zero probability, and then have P(H|L) = 1/4 and P(T|L) = 3/4, even though P(L) = 0. But that is introducing a lot of complications. Primitive conditional probabilities are not unproblematic.

What should we do? Maybe we should suppose something like primitive suppositional decision theory, where what we are primitively given are the suppositional probabilities PF and PL, without them being defined in terms of conditional and unconditional credences as in evidential and causal decision theories. But this seems problematic. Do we have to suppose that in addition to conditional and unconditional credences, we have suppositional credences? Maybe.

Or perhaps decision theory only applies to agents that have non-zero credences of going for all the options.

Thursday, August 18, 2022

Non-uniqueness of "uniform" full conditional probabilities

Consider a fair spinner that uniformly chooses an angle between 0 and 360. Intuitively, I’ve just fully described a probabilistic situation. In classical probability theory, there is indeed a very natural model of this: Lebesgue probability measure on the unit circle. This model’s probability measure can be proved to be the unique function λ on the subsets of the unit circle that satisfies these conditions:

  1. Kolmogorov axioms with countable additivity

  2. completeness: if λ(B) is zero and A ⊆ B, then λ is defined for A

  3. rotational invariance

  4. at least one arc on the circle of length greater than zero and less than 360 has an assigned probability

  5. minimality: any other function that satisfies 1-4 agrees with λ on the sets where λ is defined.

In that sense “uniformly chooses” can be given a precise and unique meaning.

But we may be philosophically unhappy with λ as our probabilistic model of the spinner for one of two reasons. First, but less importantly, we may want to have meaningful probabilities for all subsets of the unit circle, while λ famously has “non-measurable sets” where it is not defined. Second, we may want to do justice to such intuitions as that it is more likely that the spinner will land exactly at 0 or 180 than that it will land exactly at 0. But λ as applied to any finite (in fact, any countable) set of positions yields zero: there is no chance of the spinner landing there. Moreover, we want to be able to update our probabilities on learning, say, that the spinner landed on 0 or 180—presumably, after learning that disjunction, we want 0 and 180 to have probability 1/2—but λ provides no guidance how to do that.

One way to solve this is to move to probabilities whose values are in some field extending the reals, say the hyperreals. Then we can assign a non-zero (but in some cases infinitesimal) probability to every subset of the circle. But this comes with two serious costs. First, we lose rotational invariance: it is easy to prove that we cannot have rotational invariance in such a context. Second, we lose uniqueness: there are many ways of assigning non-zero probabilities, and we know of no plausible set of conditions that makes the assignment unique. Both costs put in serious question whether we have captured the notion of “uniform distribution”, because uniformity sure sounds like it should involve rotational invariance and be the kind of property that should uniquely determine the probability model given some plausible assumptions like (1)–(5).

There is another approach for which one might have hope: use Popper functions, i.e., take conditional probabilities to be primitive. It follows from results of Armstrong and the supramenability of the group of rotations on the circle that there is a rotation-invariant (and, if we like, rotation and reflection invariant) finitely-additive full conditional probability on the circle, which assigns a meaningful real number to P(A|B) for any subsets A and B with B non-empty. Moreover, if Ω is the whole circle, then we can further require that P(A|Ω) = λ(A) if λ(A) is defined. And now we can compare the probability of two points and the probability of one point. For although P({x,y}|Ω) = λ({x,y}) = 0 = λ({x}) = P({x}|Ω) when x ≠ y, there is a natural sense in which {x, y} is more likely than {x} because P({x}|{x,y}) = 1/2.

Unfortunately, the conditional probability approach still doesn’t have uniqueness, and this is the point of this post. Let’s say that what we require of our conditional probability assignment P is this:

  1. standard axioms of finitely-additive full conditional probabilities

  2. (strong) rotational and reflection invariance

  3. being defined for all pairs of subsets of the circle with the second one non-empty

  4. P(A|Ω) = λ(A) for any Lebesgue-measurable A.

Unfortunately, these conditions fail to uniquely define P. In fact, they fail to uniquely define P(A|B) for countably infinite B.

Here’s why. Let E be a countably infinite subset of the circle with the following property: for any non-identity isometry ρ of the circle (combination of rotations and reflections), E ∩ ρE is finite. (One way to generate E is this. Let E0 be any singleton. Given En, let Gn be the set of isometries ρ such that ρx = y for some x, y in E. Then Gn is finite. Let z be any point not in {ρx : ρ ∈ Gn, x ∈ E}. Let En + 1 = En ∪ {z} (since z is not unique, we’re using the Axiom of Dependent Choice, but a lot of other stuff depends on stronger versions of Choice anyway). Let E be the union of the En. Then it’s easy to see that E ∩ ρE contains at most one point for any non-identity isometry ρ.)

Let μ be any finitely additive probability on E that assigns zero to finite subsets. Note that μ is not unique: there are many such μ. Now define a finitely additive measure ν on Ω as follows. If A is uncountable, let ν(A) = ∞. Otherwise, let ν(A) = ∑ρμ(EρA), where the sum is taken over all isometries ρ. The condition that E ∩ ρE is finite for non-identity ρ and that μ is zero for finite sets ensures that if A ⊆ E, then ν(A) = μ(A). It is clear that ν is isometrically invariant.

Let λ* be any invariant extension of Lebesgue measure to a finitely additive measure on all subsets of the circle. By Armstrong’s results (most relevantly Proposition 1.7), there is a full conditional probability P satisfying (6)–(8) and such that P(A|E) = μ(AE) and P(A|Ω) = λ*(A) (here we use the fact that ν(A) = ∞ whenever λ*(A) > 0, since λ*(A) > 0 only for uncountable A). Since μ wasn’t unique and E is countable, conditions (6)–(9) fail to uniquely define P for countably additive conditions.

Wednesday, October 7, 2020

Weak invariance of full conditional probabilities

In two papers (here and here), I explored two different concepts of symmetry for conditional probabilities. The concept of strong invariance says that P(gA|B)=P(A|B) for a symmetry g as long as A and gA are subsets of B. The concept of weak invariance says that P(gA|gB)=P(A|B) for a symmetry g. In some special cases, the weak concept implies the strong concept.

Anyway, here’s an interesting thing: the weak concept does not capture our symmetry intuitions. Take perhaps the simplest case, a lottery on the set of integers Z, and say that the symmetries are shifts. It turns out that there is a weakly shift-invariant full conditional probability P such that:

  1. P({m}|{m, n}) = P({n}|{m, n}) (singleton fairness)

  2. P(A|A ∪ B)=0 and P(B|A ∪ B)=1 whenever B has infinitely many positive integers and A has finitely many positive integers.

Condition (2) implies that it is more likely that the winning ticket is a power of two than that that is a negative integer. So weak shift invariance is very far from strong invariance.

(And in fact one can have strong invariance for the lottery on Z if one wants. One can even have have strong invariance under shifts and reflections if one wants.)

The proof is a modification of West's proof of a result for qualitative probabilities.

Wednesday, August 19, 2020

Product spaces for hyperreal and full conditional probabilities

I think the following is a consequence of a hyperreal variant of the Horn-Tarski extension theorem for measures on boolean algebras:

Claim: Suppose that <Ωi, Fi, Pi> for i ∈ I is a finitely additive probability space with values in some field R* of hyperreals. Then, assuming the Axiom of Choice, there is a hyperreal-valued finitely additive probability space <Ω, 2Ω, P> where Ω = ∏i ∈ IΩi and where the Ωi-valued random variables πi given by the natural projections of Ω to Ωi are independent and have the distributions given by the Pi.

Note that the values of P might be in a hyperreal field larger than R*.

Given the Claim, and given the well-known correspondences between hyperreal-valued probabilities and full conditional real-valued probabilities, it follows that we can define meaningful product-space conditional real-valued probabilities.

It would be really nice if the product-space conditional probabilities were unique in the special case where Fi is the power set of Ωi, or at least if they were close enough to uniqueness to define the same real-valued conditional probabilities.

For a particularly interesting case, consider the case where X and Y are generated by uniform throws of a dart at the interval [0, 1], and we have a regular finitely additive hyperreal-valued probability on [0, 1] (regular meaning that all non-empty sets have positive measure). Let Z be the point (X, Y) in the unit square.

Looking at how the proof of the Horn-Tarski extension theorem works, it seems to me that for any positive real number r, and any non-trivial line segment L along the x = y diagonal in the square [0, 1]2, there is a product measure P satisfying the conditions of the Claim (where P1 and P2 are the uniform measures on [0, 1]) such that P(L)=rP(H), where H is the horizontal line segment {(x, 0):x ∈ [0, 1]}. For instance, if L is the full diagonal, we would intuitively expect P(L)=21/2P(H), but in fact we can make P(L)=100000P(H) or P(L)=P(H)/100000 if we like. It is clear that such a discrepancy will generate different conditional probabilities.

I haven’t checked all the details yet, so this could be all wrong.

But if it is right, here is a philosophical upshot. We would expect there to be a unique canonical product probability for independent random variables. However, if we insist on probabilities that are so fine-grained as to tell infinitesimal differences apart, then we do not at present have any such unique canonical product probability. If we are to have one, we need some condition going beyond independence.

This is part of a larger set of claims, namely that we do not at present have a clear notion of what “uniform probability” means once we make our probabilities more finegrained than classical real-valued probability.

Putative Sketch of Proof of Claim: Embedding R* in a larger field if necessary, we may assume that R* is |2Ω|-saturated. Define a product measure on the cylinder subsets of Ω as usual. The proof of the Horn-Tarski extension theorem for measures on boolean algebras looks to me like it works for |B|-saturated hyperreal-valued probability measures where B is the boolean algebra, and completes the proof of our claim.

Tuesday, July 21, 2020

Do Popper functions carry enough information?

Let P be a Popper function on some algebra F of events on Ω. There is a natural way to define a probability comparison given P, namely A ⪅ B iff P(A|A ∪ B)≤P(B|A ∪ B), which I think I’ve used before. Unfortunately, this can violate a common axiom of probability comparisons, namely that A ⪅ B iff Ω − B ⪅ Ω − A.

For instance, consider the Popper function generated by a regular hyperreal probability on [0, 1] whose standard part agrees with Lebesgue measure on intervals. Then [0, 1]⪅[0, 1), since both have conditional probability 1 on [0, 1]∪[0, 1)=[0, 1]. But their complements are ∅ and {1}, respectively, and {1}⪅∅ is false.

This little fact may have some significance. There is a theorem in the literature that shows a correlation between Popper function and hyperreal probabilities: every Popper function with all non-empty sets regular can be generated from a normal hyperreal probability using the conditional probability formula, and given a Popper function with all non-empty sets regular there is a normal hyperreal probability from which the Popper function can be generated. This has led to a debate whether it’s better to work with Popper functions or hyperreal probabilities. One argument for working with Popper functions—an argument I’ve approved of in print—is that the hyperreal probabilities carry more information than reality does. I still think that this problem is there for hyperreal probabilities. But the above argument suggests that going to a Popper function may have discarded too much information. Popper functions allow for fine-grained comparisons of “small” events, such as singletons, but not for the complements of these events.

Monday, July 20, 2020

Complete conditional probabilities, infinitesimal probabilities and two easy Frankenstein facts

Suppose we have a complete finitely-additive conditional probability P(⋅|⋅) on some algebra F of events on a space Ω (i.e., P is a Popper function with all non-empty sets regular), so that P(A|B) is defined for every non-empty B.

Here’s a curious thing: there is a very large disconnect between how P behaves when conditioning on sets of zero measure and when conditioning on sets of non-zero measure.

Here’s one way to see the disconnect. Consider any other complete finitely additive conditional probability Q(⋅|⋅) on the same algebra F, and suppose that Q assigns unconditional measure zero to everything that P does: i.e., if P(A|Ω)=0, then Q(A|Ω)=0.

Then, Frankenstein-fashion, we can sew P and Q into a new conditional probability R where R(A|B)=P(A|B) if P(A|Ω)>0 and R(A|B)=Q(A|B) if P(A|Ω)=0. In other words, R behaves exactly like P when conditioning on non-zero measure sets and exactly like Q when conditioning on zero measure sets.

[To check that R is a conditional probability, the one non-trivial condition to check here is that R(A ∩ B|C)=R(A|C)R(B|A ∩ C). Suppose P(C|Ω)=0. Then the equality follows from the corresponding equality for Q. Suppose P(C|Ω)>0. If P(A ∩ C|Ω)>0 as well, then our equality follows from the corresponding equality for P. Suppose now that P(A ∩ C|Ω)=0. Then the equality to be demonstrated is equivalent to P(A ∩ B|C)=P(A|C)Q(B|A ∩ C). But if P(A ∩ C|Ω)=0, then P(A ∩ B|C)=0 and P(A|C)=0.]

There is an analogous Frankenstein fact for infinitesimal probabilities. Let P and Q be any two finitely-additive probabilities with values in some hyperreal field, and suppose that Q is tiny whenever P is tiny, where a hyperreal is “tiny” provided that it is zero or infinitesimal. Then there is a Frankenstein probability R = Std P + Inf Q, where Std x and Inf x are the standard and infinitesimal parts of a finite hyperreal. (The fact that Q is tiny when P is tiny is used to show that R is non-negative.) This R then has the same large-scale (i.e., standard scale) behavior as P and small-scale behavior as Q.

In other words, as we depart from classical probability, we get a nearly complete disconnect between small-scale and large-scale behavior.

Wednesday, June 17, 2015

Non-conglomerability

This result is probably known, and probably not optimal. A conditional probability function P is conglomerable provided that for any partition {Hi} (perhaps infinite and maybe even uncountable) of the state space if P(A|Hi)≥r for all i, then P(A)≥r.

Theorem. Assume the Axiom of Choice. Suppose P is a full conditional probability function (i.e., Popper function) P on an uncountable space such that:

  1. all singletons are measurable
  2. the function satisfies this regularity condition for all elements x and y: P({x}|{x,y})>0
  3. there is a partition of the probability space into two disjoint subsets A and B with the same cardinality such that P(A)>0 and P(B)>0
Then P is not conglomerable.

Conditions (2) and (3) are going to be intuitively satisfied for plausible continuous probabilities, like uniform and Gaussian ones. So in those cases there is no hope for a conglomerable conditional probability.

Sketch of proof: Let Q be a hyperreal-valued unconditional probability corresponding to P, so that P(X|Y)=Q(XY)/Q(Y). The regularity condition (2) implies that that there is a hyperreal α such that Q(F)/α is finite, non-zero and non-infinitesimal for each finite set F. (Just let α=Q({x0}) for any fixed x0.) Let R(F) be the standard part of Q(F)/α for any finite set α. Then P(F|G)=R(FG)/R(G) for any finite sets F and G. Moreover, R is finitely additive and non-zero on every singleton.

Since A2 has the same cardinality as A, there is a function f from B to the subsets of A with the property that f(b) and f(c) are disjoint if b and c are distinct and every f(b) is uncountable. Choose a finite number c such that P(A)<c/(1+c). For each b in B, choose a finite subset Fb of f(b) such that R(Fb)>cR({b}). Such a finite subset exists as R is finitely additive and the sum of an uncountable number of non-zero positive numbers is always infinity. Let H be the union of the Fb as b ranges over B. Then AH has at most the cardinality of B. Let h be a one-to-one function from AH to B. For each b in B, let Gb=Fb if there is no a in AH such that h(a)=b; otherwise, let Gb=Fb∪{a} for such an a. Let Hb={b}∪Gb. Then R(Gb)>cR({b}) and so R(Gb)/R(Hb)>c/(1+c). Then P(A|Hb)=P(Gb|Hb)=R(Gb)/R(Hb)>c/(1+c). But the Hb are a partition of our probability space, and P(A)<c/(1+c), so we have a violation of conglomerability.

Tuesday, June 16, 2015

Mixing and matching conditional probabilities

Given an unconditional probability function P, one can always (at least given the Axiom of Choice) extend to a full conditional probability function, or a Popper function, that allows one to assign values to P(A|B) even when P(B)=0. Typically, the extension is not unique. In fact, it turns out that there is no logical connection between the conditional probabilities P(A|B) for B a null set (a set of zero probability) and the unconditional probabilities.

What do I mean by saying there is no logical connection? Well, it turns out we can mix and match and the null-probability-condition parts of Popper functions with the other parts. Suppose that P and Q are two Popper functions defined on the same sets. Then we can define a frankenfunction by letting R(A|B)=P(A|B) when B is not a null set and R(A|B)=Q(A|B) when it is a null set. And this frankenfunction is a perfectly fine Popper function.

This is a problem. There is a complete disconnect between the null-probability-condition of the Popper function and the non-null-probability condition parts. (Another manifestation of this problem is the fact that in many cases we lack conglomerability.)

Wednesday, September 4, 2013

Something positive about Bayesian regularity

The brunt of a lot of my recent posts has been that there is no hope for Bayesian regularity if one requires natural invariance conditions. But here is a positive result. For this result, we will need the values of the probabilities to be taken in a very special space which is a variant of a space defined by Dos Santos. We now define this space. Let I be a totally ordered set under ≤. Let R(I) be the set of monotone non-increasing functions f from A to [0,∞] with the property that either f(x)=0 for all x or there is a unique (!) member i of I such that 0<f(i)<∞. Note that R(I) is itself a totally ordered set under pointwise comparison and it has a natural pointwise addition operation that respects the ordering. You can think of R(I) as very much like a set of non-negative hyperreals where numbers whose ratio is infinitesimally close to 1 are identified.

It is fairly easy to see that it follows from Proposition 1.7 of Armstrong that if G is a supramenable group and X is any space acted on by G, then there is an I and a finitely additive measure P on all subsets of X with values in A that is strictly positive in the sense that P(B)=0 if and only if B is the empty set. Moreover, we can normalize P into something like a probability by supposing that I has a final element, call it 1, and P(X) is the member f of R(I) such that f(1)=1.

In particular, there will be a strictly positive R(I)-valued finitely additive measure on the circle and the line, invariant under isometries. But not in dimensions greater than one due to Banach-Tarski related stuff.

Fact: There is a natural correspondence between real-valued Popper functions on X that make every non-empty subset normal and strictly-positive finitely-additive R(I)-valued measures. It's easy to see how this correspondence goes in one direction. Suppose we have such a strictly positive measure P. We want to define P(A|B) for some non-empty B. Choose the unique i in I such that P(B)(i) is in (0,∞) and then define P(A|B)=P(AB)(i)/P(B)(i). Moreover, the Popper function will be strongly-invariant (P(gA|B)=P(A|B) if gA and A are subsets of B)) if and only if the corresponding R(I)-valued measure is invariant.

For epistemological purposes, this is a move in the happy direction, but the fact that nothing like this can work in Euclidean settings in higher dimensions is a problem.

Note that P as above will be regular in the weak sense that 0<P(A) if A is non-empty but typically not in the strong sense that if A is a proper subset of B, then P(A)<P(B).

Friday, August 23, 2013

Progress report: Positive results for invariant Popper functions

This is a very technical note. I've spent a fair amount of time this week thinking about invariant Popper functions. Say that a group G is neatly supramenable if there is a Popper function P on G with every non-empty subset normal and satisfying the strong invariance condition P(gA|B)=P(A|B) whenever gAAB. Neatly supramenable groups are supramenable: every non-empty subset A has an invariant finitely additive measure m such that m(A)=1. Anyway, I think I can prove—I now have two proofs drafted, so that makes me more confident—that every exponentially bounded group is neatly supramenable. Thus, every elementary supramenable group is neatly supramenable.

One philosophically interesting upshot of all this is that n-dimensional Euclidean space supports a Popper function with all non-empty subsets normal that is invariant under all translations as well as under single-coordinate reflections ((x1,...,xi,...,xn) going to (x1,...,−xi,...,xn)). But when one adds rotations into the mix, this is false for n≥2. So there is something philosophically problematic about rotations for the notion of uniform probability.

Don't quote the result yet as the proofs use mathematics that I am not very familiar with (ultrafilters, non-standard analysis, etc.).

Oh, and all of this uses the Axiom of Choice.

Saturday, August 10, 2013

Normal Popper functions and comparative probability

Let Ω be a non-empty set and F be a field of subsets (i.e., set of subsets closed under finite unions and complements). A Popper function on F is a real-valued function P defined for pairs of members of F such that:

  1. 0≤P(X|Y)≤P(Y|Y)=1
  2. if P(Ω−Y|Y)<1, then P(−|Y) is a finitely additive probability on F
  3. P(XY|Z)=P(X|Z)P(Y|XZ)
  4. if P(X|Y)=P(Y|X)=1, then P(Z|X)=P(Z|Y).
(This is van Fraassen's axiomatization.) A member B of F is normal provided that P(Ω−B|B)<1. (This condition is equivalent to saying that P(∅|B)<1.) A Popper function is normal provided that every non-empty member of F is normal.

On the other hand, a comparative probability function on F (Paul Bartha has talked about things like this) is a function from pairs of members of F to [0,∞] such that:

  1. C(X,X)=1
  2. C(X,Y)C(Y,Z)=C(X,Z) provided the left-hand side is defined (0 times ∞ and ∞ times zero are the undefined cases)
  3. C(−,Y) is a finitely-additive measure if Y is non-empty.

Alan Hajek, I think, has suggested that one can define a comparative probability function in terms of a Popper function. We can do it as follows. If P is a normal Popper function, then let CP(X,Y)=P(X|XY)/P(Y|XY). And given a comparative probability function, we can define a normal Popper function by PC(X|Y)=CP(XY|Y).

I haven't written out the details, but it looks like we then have:

Proposition. If P is a normal Popper function, then CP is a comparative probability function, and if C is a comparative probability function, then PC is a normal Popper function. Moreover, if P is a normal Popper function and C=CP, then PC=P, while if C is a comparative probability function and P=PC, then CP=C.

Thus, there is a nice one-to-one correspondence between comparative probabilities and normal Popper functions.

Personally, I find comparative probabilities to be easier to prove theorems about than Popper functions, because I find (6) much easier to remember than (3).

Friday, August 9, 2013

Problems for isometrically invariant Popper functions in the plane

Popper functions encode finitely-additive conditional probabilities. The hope of Popper functions is to allow conditionalizations on sets that classically have null probability, such as finite sets in a continuous context. A Popper function allows for non-trivial conditionalization on a set A only if A is normal, i.e., P(∅|A)=0. A Popper function P is weakly invariant under isometries (the group generated by rotations, translations and reflections, in the case of the plane) provided that P(gA|gB)=P(A|B) for any isometry g and A and B such that A,B,gA,gB are subsets of our probability space Ω. (Strong invariance would also say that P(gA|B)=P(A|B) and P(A|gB)=P(A|B) under appropriate circumstances.)

In an earlier post I basically proved (without the Axiom of Choice) that any weakly isometrically invariant Popper function on a solid three-dimensional ball that makes all countable sets measurable also makes all finite sets abnormal. This triviality result of course generalizes to higher dimensions.

I am now able to prove that if P is a weakly isometrically invariant Popper function on a subset Ω of the plane that contains a solid disc, with P defined at least for all countable sets, then there is a countable abnormal set. Again, since subsets of an abormal set are abnormal, it follows that some singleton is abnormal, and hence all singletons are abnormal by invariance, and hence all finite sets are abnormal as finite unions of abnormal sets are abnormal.

A proof is very roughly sketched here. It uses Just's construction of a bounded paradoxical subset of the plane (cf. this paper which gives another construction that could be used).

It also follows that there is no weakly isometrically invariant comparative probability function defined for all nonempty pairs of subsets of Ω (for definition of comparative probability, see my answer here).

On the other hand, Parikh and Parnes have shown that there is a Popper function defined for all pairs of subsets of the unit interval on the line making all nonempty subsets normal. Moreover I think their proof can be used to show that there is a translation-invariant Popper function defined for all pairs of subsets of n-dimensional Euclidean space making all non-empty subsets normal. We can get invariance under reflection in any finite set of orthogonal hyperplanes for free just by averaging combinations.

So the problem really is with rotations. Something important changes when one goes from dimension one to higher dimensions: rotations become available.

Thursday, June 13, 2013

Popper functions and null sets

Let's go back to the problem that I keep on thinking about: How to distinguish possibilities that are classically of null probability. For instance, given a uniform choice of a point on some nice set (say, a ball) in Euclidean space, we want to say something like P({x,y})>P({x}), when x and y are distinct: it's more likely that one would hit one of two points than that one would hit a particular point. A series of my blog posts (and at least one article) showed that infinitesimals aren't the way. What about conditional probabilities?

For various reasons, instead of taking unconditional probabilities to be fundamental and defining conditional probabilities in terms of them, one may want to take conditional probabilities as fundamental. The standard method is to use Popper functions (I'll assume the linked axioms below). One might hope to use Popper functions to do things like make sense of the difference between the probability of two points in a continuous case (say, where a point is uniformly chosen in some nice subset of a Euclidean space) and the probability of a single point. For instance, one might hope that P({x,y}|{x,y})>P({x}|{x,y}) whenever x and y are distinct.

This won't work, however. Instead of working with propositions, I will work with sets—the definitions of Popper functions neatly adapt. Let Ω be a solid three-dimensional ball. Assume that P(A|B) is defined for all A and B in some algebra F of subsets of Ω.

Say that a set A in F is P-trivial provided that P(C|A)=1 for all C (including for the empty set). The empty set is trivial, of course. In order for us to have any hope of saying things like P({x,y}|{x,y})>P({x}|{x,y}), we better have sets with two points be non-trivial. Now, it's not hard to show that a finite union of trivial sets is trivial and that the subsets of a trivial set are trivial, so in order for all the finite sets to be non-trivial, it's necessary and sufficient that singletons be non-trivial.

Moreover, we want rotational symmetry. Say that F and P are rotationally symmetric provided that for any rotation r around the origin and A and B in F, rA and rB are in F, and P(rA|rB)=P(A|B).

Theorem. If P is rotationally symmetric and F includes all countable subsets of Ω, then for any sphere of positive radius around the origin lying in Ω there is at least one P-trivial countably infinite set on the surface of
that sphere. Thus, all finite sets not containing the origin are trivial.

If we are to be extending something like classical probabilities, we do want countable subsets of Ω to be in F. The triviality of finite sets follows from the fact that some singleton on every sphere of non-zero radius about the origin is trivial (since any subset of a trivial set is trivial) and if one singleton on that sphere is trivial, then by rotational invariance, they all are, while finite unions of trivial sets are trivial. So all we need to prove is the existence of that trivial countably infinite set.

The proof is easily based on standard ideas from the proof of the Banach-Tarski Paradox.

Proof. Let SO(3) be the rotation group around the origin. Famously, there is a subgroup G that is isomorphic to the free group F2 on two elements. Choose a point ω in Ω such that ρ(ω)=ω only if ρ is the identity rotation. (This is a counting argument: G is a countable group, and each non-identity member of G has two fixed points on any given sphere around the center, so for any fixed sphere there will be only countably many fixed points of non-identity members of G on it, and hence there will be a point on the sphere that isn't a fixed point of any non-identity member.) Let H={ρ(ω):ρ∈G}. Then H is a countable subset of Ω.

For a reductio ad absurdum, suppose H is non-trivial. Then P(−|H) is a finitely additive probability measure on H. Moreover, HH for any ρ∈G, so by rotational invariance PK|H)=PKH)=P(K|H) for any ρ∈G, and so P(−|H) is a finitely additive G-invariant measure on H. Using the bijection between G and H given by f(ρ)=ρ(ω) (we use the choice of ω to see that f is one-to-one), we can then get a finitely additive G-left-invariant measure on G. But G is isomorphic to F2 and hence is not amenable and hence has no such invariant measure. (One could also neatly demonstrate a paradoxical decomposition here.) That's a contradiction, so H must be trivial. QED

Even though this argument uses ideas from the proof of Banach-Tarski, and famously the latter uses the Axiom of Choice, this argument does not use the Axiom of Choice.

One can perhaps get out of this by having a more onerous requirement on the sets A and B that are on either side of the bar in "P(A|B)" than that they are fit in a single algebra F. We want any countable set to be an acceptable A, but perhaps we don't want to allow every countable set as an acceptable B. I don't know what natural requirement could be put here.