Showing posts with label nonmeasurable sets. Show all posts
Showing posts with label nonmeasurable sets. Show all posts

Thursday, February 1, 2024

Fusion and the Axiom of Choice

Assume classical mereology. Then for any formula that has a satisfier, there is a fusion of all of its satisfiers. More precisely, if ϕ is a formula with z not a free variable in ϕ, then the universal closure of the following under all free variables is true:

  1. xϕ → ∃zFϕ, x(z)

where Fϕ, x(z) says that z is a fusion of the satisfiers of ϕ with respect to the variable x (there is more than one account of what exactly the “fusion” is). This is the fusion axiom schema.

Stipulate that a region of physical space is a fusion of points.

Question: Is there a nonmeasurable region of (physical) space?

Assuming the language for formulas in our classical mereology is sufficiently rich, the answer is positive. For simplicity, suppose that physical space is Euclidean (the non-Euclidean case is handled by working in a small neighborhood which is diffeomorphic to a neighborhood of a Euclidean space). Let ψ be the isomorphism between the points of physical space and the mathematical space R3. Let ϕ be the formula ψ(x) ∈ y. Applying (1), we conclude that for any subset a of R3, there is a set of points of physical space that correspond to a under ψ. If we let a be one of the standard nonmeasurable subsets of R3, we get an affirmative answer to our question.

But now we have an interesting question:

  1. What grounding or explanatory relation is there between the existence of a nonmeasurable region of physical space and the existence of a nonmeasurable subset of mathematical space?

The two simplest options are that one is explanatorily prior to the other. Let’s explore these.

Suppose the existence of a nonmeasurable physical region depends on the existence of the nonmeasurable set. Well, it is a bit strange to think of a concrete object—a region of physical space—as partly grounded in the existence of a set. This doesn’t sound quite right to me.

What about the other way around? This challenges the fairly popular doctrine that complex things entities are a free lunch given simples. For if the existence of the nonmeasurable region is prior to the existence of an abstract set, it seems that we actually have quite a significant metaphysical “effect” of this complex object.

Moreover, if the existence of the nonmeasurable region is not grounded in the existence of nonmeasurable set, whether or not there is grounding running the other way, we have a difficult question of why there is in fact a nonmeasurable region. Without relying on nonmeasurable sets, it doesn’t seem we can get the nonmeasurable region out of the axioms of classical mereology. It seems we need some sort of a mereological Axiom of Choice. How exactly to formulate that is difficult to say, but one version that is enough for our purposes would be that given any formula ρ(x,y) that expresses a non-empty equivalence relation on the simples satisfying ϕ, there is an object z such that if ϕ(x) then there is exactly one simple x′ such that ρ(x,x′) and x is a part of z, and every object that meets z meets some simple satsifying ϕ.

But my intuition is that a mereological Axiom of Choice would badly violate the doctrine that complex objects are a free lunch. If all we had in the way of complex-object-forming axioms were reflexivity, transitivity and fusion, then it would not be crazy to say that complex objects are a fancy way of talking about simples. But the “indeterministic” nature of the Axiom of Choice does not, I think, allow one to say that.

Wednesday, January 31, 2024

Modality and the Axiom of Choice

Suppose that the set theory of our world is a Solovay model, where we don’t have the Axiom of Choice (AC), and where every subset of the reals is Lebesgue measurable. Now imagine that God picks out a line in space, and defines the Vitali equivalence relation for points on that line (where two points are equivalent if and only if the distance between them is a rational number). It is then surely within God’s power to create a particle of some unexemplified type T at exactly one point in every equivalence class. There is nothing incoherent about that! But if God did that, then there would surely be a set of the points containing a particle of type T. And that set would be a nonmeasurable Vitali set.

So what?

Well, prima facie, there are three possibilities about the existence of nonmeasurable sets:

  1. Necessarily, there are no nonmeasurable sets.

  2. Necessarily, there are nonmeasurable sets.

  3. It is contingent whether there are nonmeasurable sets.

My argument strongly suggests that if there are no nonmeasurable sets, it is nonetheless possible that there are nonmeasurable sets. Hence, (1) is ruled out.

So we have an argument for the disjunction of (2) and (3).

Now, I think a lot of people have the intuition that mathematical facts are necessary. If so, then (3) is ruled out. They will see this as an argument for (2).

I don’t see it that way myself: I am quite open to contingent mathematical truths.

More generally, the argument shows that:

  1. For any set of disjoint nonempty subsets of the reals, it is possible that there is a choice function.

Again, if the existence of pure sets is not a contingent matter, we conclude AC is true for all subsets of the reals.

Tuesday, December 6, 2022

Dividing up reasons

One might think that reasons for action are exhaustively and exclusively divided into the moral and the prudential. Here is a problem with this. Suppose that you have a spinner divided into red and green areas. If you spin it and it lands into red, something nice happens to you; if it lands on green, something nice happens to a deserving stranger. You clearly have reason to spin the spinner. But, assuming the division of reasons, your reason for spinning it is neither moral nor prudential.

So what should we say? One possibility is to say that there are only reasons of one type, say the moral. I find that attractive. Then benefits to yourself also give you moral reason to act, and so you simply have a moral reason to spin the spinner. Another possibility is to say that in addition to moral and prudential reasons there is some third class of “mixed” or “combination” reasons.

Objection: The chance p of the spinner landing on red is a prudential reason and the chance 1 − p of its landing on green is a moral reason. So you have two reasons, one moral and one prudential.

Response: That may be right in the simple case. But now imagine that the “red” set is a saturated nonmeasurable subset of the spinner edge, and the “green” set is also such. A saturated nonmeasurable subset has no reasonable probability assignment, not even a non-trivial range of probabilities like from 1/3 to 1/2 (at best we can assign it the full range from 0 to 1). Now the reason-giving strength of a chancy outcome is proportionate to the probability. But in the saturated nonmeasurable case, there is no probability, and hence no meaningful strength for the red-based reason or for the green-based reason. But there is a meaningful strength for the red-or-green moral-cum-prudential reason. The red-or-green-based reason hence does not reduce to two separate reasons, one moral and one prudential.

Now, one might have technical worries about saturated nonmeasurable sets figuring in decisions. I do. (E.g., see the Axiom of Choice chapter in my infinity book.) But now instead of supposing saturated nonmeasurable sets, suppose a case where an agent subjectively has literally no idea whether some event E will happen—has no probability assignment for E whatsoever, not even a ranged one (except for the full range from 0 to 1). The spinner landing on a set believed to be saturated nonmeasurable might be an example of such a case, but the case could be more humdrum—it’s just a case of extreme agnosticism. And now suppose that the agent is told that if they so opt, then they will get something nice on E and a deserving stranger will get something nice otherwise.

Final remark: The argument applies to any exclusive and exhaustive division of reasons into “simple” (i.e., non-combination) types.

Friday, February 25, 2022

Inscrutable probabilities and the replay argument

Suppose that we have two hypotheses about a long sequence X of zeroes and/or ones:

  • S: The sequence has the underlying structure of a sequence of independent and identically distributed (iid) probabilistic phenomena with possible outcomes zero or one.

  • U: The sequence is completely probabilistically unstructured, being completely probabilistically inscrutable.

What can we say about the evidential import of X on S and U?

Here is one thought. We have a picture of how a repeated sequence of independent and identically distributed binary events should look. Much as van Inwagen says in his discussion of the replay argument, we expect to have the proportion of ones to oscillate but then converge to some value. So if X looks like that, then we have evidential support for S, and if X doesn’t look like that—maybe there is wilder oscillation and no convergence—then we do not have evidential support for S.

But does the fact that the sequence X has a “nice gradual convergence” in frequencies support the structure hypothesis S? Van Inwagen in his replay argument against indeterministic free will seems to assume so.

Actually, convergence of frequencies does not support S (or U, for that matter). Here is why not. Let’s model hypothesis S as follows. There is some unknown real dispositional probability p between 0 and 1, and then our sequence X is taken by observing a sequence of iid results with the probability of the result being 1 being equal to p. Given p, any sequence of n events whose proportion of ones is m/n has probability pm(1−p)n − m, regardless of whether the sequence shows any “nice gradual convergence” of the frequency of ones to m/n or is just a seqence of n − m zeroes followed by m ones.

Of course, hypothesis S doesn’t say what the probability p of getting a one on a given experiment is. So, we need to suppose some probability distribution Q over the possible values of p in the interval [0,1], and then our probability of the sequence X of m ones and n − m zeroes will be pm(1−p)n − mdQ(p). But it’s still true that this probability does not depend on the order between the ones and zeroes. The sequence could “look random”, or it could look as fishy as you like, and the probability of it on S would be exactly the same if the number of ones is the same.

On the other hand, on the hypothesis U, all possible sequences X of length n are probabilistically inscrutable, and hence we cannot say anything about any sequence being more likely than another—we might as well represent the probability of any sequence X of observations as just an interval-valued probability of [0,1].

So, no facts about the order of events in X make any difference between the hypotheses S and U. In particular, van Inwagen’s idea that an appearance of convergence is evidence for S is false.

Now, let’s say a little more about the Bayesian import of the observation X. Let’s suppose that S and I are our only two hypotheses, and that both have serious prior probabilities, say somewhere between 0.1 and 0.9. Further, let’s suppose that that the sequence X is long and has a decent amount of variation in it—for instance, let’s suppose that it has at least 50 zeroes and at least 50 ones. Because of this, P(X|S) will be some astronomically small positive number α. Indeed, we can prove that α is at most 1/2100 ≈ 10−30, for any distribution of the unknown probability p of getting a one.

On the other hand, P(X|U) will be completely probabilistically inscrutable, and hence can be represented as the full interval [0,1].

Here’s what follows from Bayes’s theorem combined with a natural way of handling inscrutable probabilities as a range of probability assignments: The posterior probability of S will be a range of probabilities that starts with something within an order of magnitude of α and ends with 1. Hence, our observation of X does not support S: upon observing X, we move from S having some moderate probability between 0.1 and 0.9, to a nearly completely inscrutable probability in a range starting with something astronomically small and ending at one. The upper end of the range is higher than what S started with but the lower end of the range is lower.

Thus, if we were to actually do the experiment that van Inwagen describes, and get the results that he thinks we would get (namely, a sequence converging to some probability), that would not support the hypothesis S that van Inwagen thinks it would support.

Monday, September 30, 2019

Classical probability theory is not enough

Here’s a quick argument that classical probability cannot capture all probabilistic phenomena even if we restrict our attention to phenomena where numbers should be assigned. Consider a nonmeasurable event E, maybe a dart hitting a nonmeasurable subset of the target, and consider a fair coin flip that is causally isolated from E. Let H and T be the heads and tails results of the flip. Then let A be this disjunctive event:

  • (E and H) or (not-E and T).

Intuitively, event A clearly has probability 1. If E happens, the probability of A is 1/2 (heads) and if E doesn’t happen, it’s also 1/2 (tails). (The argument uses finite conglomerability, but it is also highly intuitive.)

So a precise number should be assigned to A, namely 1/2. And ditto to H. But we cannot have these assignments in classical probability theory. For if we did that, then we would also have to assign a probability to the conjunction of H and A, which is equivalent to the conjunction of E and H. But we cannot assign a probability to the conjunction of E and H, because E and H are independent, and so we would have a precise probability for E, namely P(E)P(H)/P(H)=P(E&H)/P(H), contrary to the nonmeasurability of E.

Monday, April 8, 2019

The probability of the universe popping into existence

Consider the hypothesis that contingent reality popped into existence uncaused.

Now, either popping into existence uncaused is astronomically unlikely or not astronomically unlikely.

If it is astronomically unlikely, then we have a very strong Bayesian argument for theism. For then P(contingent reality | no God) is astronomically small while P(contingent reality | God) is at least moderately high.

If uncaused popping into existence is not astronomically unlikely, then there are two main options. The first option is that there is no meaningful probability of such an event. In that case, there is no meaningful probability of Maxwell’s Demon popping into existence for no cause at all in one’s lab. But if Maxwell’s Demon were to pop into existence in one’s lab, then one wouldn’t expect to get the predicted observations. Thus, if there is no meaningful probability of things popping into existence for no cause at all, then there is no meaningful probability of our scientific predictions, and science falls apart. That’s not acceptable.

The other option is that there is a probability, and it’s not astronomically small. But then at every moment of time, it is not astronomically unlikely that an object would causelessly pop into existence. Since there are astronomically many moments of time during a second (perhaps infinitely many, but at least equal to the number of Planck times in a second, i.e., of the order of 1043), it seems we should expect to see lots of objects pop into existence causelessly. And we don’t observe that.

There is lots of technical detail to fix in this argument.

Wednesday, May 16, 2018

Possibly giving a finite description of a nonmeasurable set

It is often assumed that one couldn’t finitely specify a nonmeasurable set. In this post I will argue for two theses:

  1. It is possible that someone finitely specifies a nonmeasurable set.

  2. It is possible that someone finitely specifies a nonmeasurable set and reasonably believes—and maybe even knows—that she is doing so.

Here’s the argument for (1).

Imagine we live an uncountable multiverse where the universes differ with respect to some parameter V such that every possible value of V corresponds to exactly one universe in the multiverse. (Perhaps there is some branching process which generates a universe for every possible value of V.)

Suppose that there is a non-trivial interval L of possible values of V such that all and only the universes with V in L have intelligent life. Suppose that within each universe with V in L there runs a random evolutionary process, and that the evolutionary processes in different universes are causally isolated of each other.

Finally, suppose that for each universe with V in L, the chance that the first instance of intelligent life will be warm-blooded is 1/2.

Now, I claim that for every subset W of L, the following statement is possible:

  1. The set W is in fact the set of all the values of V corresponding to universes in which the first instance of intelligent life is warm-blooded.

The reason is that if some subset W of L were not a possible option for the set of all V-values corresponding to the first instance of intelligent life being warm-blooded, then that would require some sort of an interaction or dependency between the evolutionary processes in the different universes that rules out W. But the evolutionary procesess in the different universes are causally isolated.

Now, let W be any nonmeasurable subset of L (I am assuming that there are nonmeasurable sets, say because of the Axiom of Choice). Then since (3) is possible, it follows that it is possible that the finite description “The set of values of V corresponding to universes in which the first instance of intelligent life is warm blooded” describes W, and hence describes a nonmeasurable set. It is also plainly compossible with everything above that somebody in this multiverse in fact makes use of this finite description, and hence (1) is true.

The argument for (2) is more contentious. Enrich the above assumptions with the added possibility that the people in one of the universes have figured out that they live in a multiverse such as above: one parametrized by values of V, with an interval L of intelligent-life-permitting values of V, with random and isolated evolutionary processes, and with the chance of intelligent life being warm-blooded being 1/2 conditionally on V being in L. For instance, the above claims might follow from particularly elegant and well-confirmed laws of nature.

Given that they have figured this out, they can then let “Q” be an abbreviation for “The set of all values of V corresponding to universes wehre the first instance of intelligent life is warm-blooded.” And they can ask themselves: Is Q likely to be measurable or not?

The set Q is a randomly chosen subset of L. On the standard (product measure) understanding of how to probabilistically make sense of this “random choice” of subset, the event of Q being nonmeasurable is itself nonmeasurable (see the Sawin answer here). However, intuitively we would expect Q to be nonmeasurable. Terence Tao shares this intuition (see the paragraph starting “Intuitively”). His reason for the intuition is that if Q were measurable, then by something like the Law of Large Numbers, we would expect the intersection of Q with a subinterval I of L to have a measure equal to half of the measure of I, which would be in tension with the Lebesgue Density Theorem. This reasoning may not be precisifiable mathematically, but it is intuitively compelling. One might also just have a reasonable and direct intuition that the nonmeasurability is the default among subsets, and so a “random subset” is going to be nonmeasurable.

So, the denizens of our multiverse can use these intuitions to reasonably conclude that Q is nonmeasurable. Hence, (2) is true. Can they leverage these intuitions into knowledge? That’s less clear to me, but I can’t rule it out.

Friday, January 26, 2018

Open theism and technical formal epistemology

If open theism is true and there is an infinite future afterlife full of free choices, then some of the puzzling cases involving non-measurable sets and probability that I like to discuss on this blog are faced by God. For any set A of sets of natural numbers, there is the proposition pA that the set of days in heaven on which David will dance a jig is a member of A. But it seems likely that some sets A will be non-measurable relative to the relevant probability measure.

So, open theists should have motivation to work on highly technical formal epistemology. The more working on that, the merrier. :-)

Thursday, November 9, 2017

A simple "construction" of non-measurable sets from coin-toss sequences

Here’s a simple “construction” of a non-measurable set out of coin-toss sequences, i.e., of an event that doesn’t have a well-defined probability, going back to Blackwell and Diaconis, but simplified by me not to use ultrafilters. I’m grateful to John Norton for drawing my attention to this.

Let Ω be the set of all countably infinite coin-toss sequences. If a and b are two such sequences, say that a ∼ b if and only if a and b differ only in finitely many places. Clearly ∼ is an equivalence relation (it is reflexive, symmetric and transitive).

For any infinite coin-toss sequence a, let ra be the reversed sequence: the one that is heads wherever a is tails and vice-versa. For any set A of sequences, let rA be the set of the corresponding sequences. Observe that we never have a ∼ ra, and that U is an equivalence class under ∼ (i.e., a maximal set all of whose members are ∼-equivalent) if and only if rU is an equivalence class. Also, if U is an equivalence class, then rU ≠ U.

Let C be the set of all unordered pairs {U, rU} where U is an equivalence class under ∼. (Note that every equivalence class lies in exactly one such unordered pair.) By the Axiom of Choice (for collections of two-membered sets), choose one member of each pair in C. Call the chosen member “selected”. Then let N be the union of all the selected sets.

Here are two cool properties of N:

  1. Every coin-toss sequence is in exactly one of N and rN.

  2. If a and b are coin-toss sequences that differ in only finitely many places, then a is in N if and only if b is in N.

We can now prove that N is not measurable. Suppose N is measurable. Then by symmetry P(rN)=P(N). By (1) and additivity, 1 = P(N)+P(rN), so P(N)=1/2. But by (2), N is a tail set, i.e., an event independent of any finite subset of the tosses. The Kolmogorov Zero-One Law says that every (measurable) tail set has probability 0 or 1. But that contradicts the fact that P(N)=1/2, so N cannot be measurable.

An interesting property of N is that intuitively we would think that P(N)=1/2, given that for every sequence a, exactly one of a and ra is in N. But if we do say that P(N)=1/2, then no finite number of observations of coin tosses provides any Bayesian information on whether the whole infinite sequence is in N, because no finite subsequence has any bearing on whether the whole sequence is in N by (2). Thus, if we were to assign the intuitive probability 1/2 to P(N), then no matter what finite number of observations we made of coin tosses, our posterior probability that the sequence is in N would still have to be 1/2—we would not be getting any Bayesian convergence. This is another way to see that N is non-measurable—if it were measurable, it would violate Bayesian convergence theorems.

And this is another way of highlighting how non-measurability vitiates Bayesian reasoning (see also this).

We can now use Bayesian convergence to sketch a proof that N is saturated non-measurable, i.e., that if A ⊆ N is measurable, then P(A)=0 and if A ⊇ N is measurable, then P(A)=1. For suppose A ⊆ N is measurable. Suppose that we are sequentially observing coin tosses and forming posteriors for A. These posteriors cannot ever exceed 1/2. Here is why. For a coin toss sequence a, let rna be the sequence obtained by keeping the first n tosses fixed and reversing the rest of the tosses. For any any finite sequence o1, ..., on of observations, and any infinite sequence a of coin-tosses compatible with these observations, at most one of a and rna is a member of N (this follows from (1) and the fact that ra ∈ N if and only if rna ∈ N by (2)). By symmetry P(A ∣ o1...on)=P(rnA ∣ rn(o1...on)) (where rnA is the result of applying rn to every member of A). But rn(o1...on) is the same as o1...on, so P(A ∣ o1...on)=P(rnA ∣ o1...on). But A and rnA are disjoint, so P(A ∣ o1...on)+P(rnA ∣ o1...on)≤1 by additivity, and hence P(A ∣ o1...on)≤1/2. Thus, the posteriors for A are always at most 1/2. By Bayesian convergence, however, almost surely the posteriors will converge to 1 or 0, respectively, depending on whether the sequence being observed is actually in A. They cannot converge to 1, so the probability that the sequence is in A must be equal to 0. Thus, P(A)=0. The claim that if A ⊇ N is measurable then P(A)=1 is proved by noting that then N − A ⊇ rN (as rN is the complement of N), and so by the above argument with rN in place of N, we have P(N − A)=0 and thus P(A)=1.

Wednesday, September 13, 2017

Probabilities and Boolean operations

When people question the axioms of probability, they may omit to question the assumptions that if A and B have a probability, so do A-or-B and A-and-B. (Maybe this is because in the textbooks those assumptions are often not enumerated in the neat lists of the “three Kolmogorov axioms”, but are given in a block of text in a preamble.)

First note that as long as one keeps the assumption that if A has a probability, so does not-A, then by De Morgan’s, any counterexample to conjunctions having a probability will yield a counterexample to disjunctions having a probability. So I’ll focus on conjunctions.

I’m thinking that there is reason to question these axioms, in fact two reasons. The first reason, one that I am a bit less impressed with, is that limiting frequency frequentism can easily violate these two axioms. It is easy to come up with cases where A-type events have a limiting frequency, B-type ones do, too, but (A-and-B)-type ones don’t. I’ve argued before that so much the worse for frequentism, but now I am not so sure in light of the second reason.

The second reason is cases like this. You have an event C that has no probability whatsoever–maybe it’s an event of a dart hitting a nonmeasurable set–and a fair indeterministic coin flip causally independent of C. Let H and T be the events of the coin flip being heads or tails. Then let A be the event:

  • (H and C) or (T and not C).

Here’s an argument that P(A)=1/2. Imagine a coin with erasable heads and tails images, and imagine that a trickster prior to flipping a coin is going to decide, using some procedure or other, whether to erase the heads and tails images on the coin and draw them on the other side. “Clearly” (as we philosophers say when we have no further argument!) as long as the trickster has no way of seeing the future, the trickster’s trick will not affect the probabilities of heads or tails. She can’t make the coin be any less or more likely to land heads by changing which side heads lies on. But that’s basically what’s going on in A: we are asking what the probability of heads is, with the convention that if C doesn’t happen, then we’ll have relabeled the two sides.

Another argument that P(A)=1/2 is this (due to a comment by Ian). Either C happens or it doesn’t. No matter which is the case, A has a chance 1/2 of happening.

So A has probability 1/2. But now what is the probability of A-and-H? It is the same as the probability of C-and-H, which by independence is half of the probability of C, and the latter probabilit is undefined. Half of something undefined is still undefined, so A-and-H has an undefined probability, even though A has a perfectly reasonable probability of 1/2.

A lot of this is nicely handled by interval-valued theories of probability. For we can assign to C the interval [0, 1], and assign to H the sharp probability [1/2, 1/2], and off to the races we go: A has a sharp probability as does H, but their conjunction does not. This is good motivation for interval-valued theories of probability.

Monday, September 11, 2017

Supertasks and empirical verification of non-measurability

I have this obsession with probability and non-measurable events—events to which a probability cannot be attached. A Bayesian might think that this obsession is silly, because non-measurable events are just too wild and crazy to come up in practice in any reasonably imaginable situation.

Of course, a lot depends on what “reasonably imaginable” means. But here is something I can imagine, though only by denying one of my favorite philosophical doctrines, causal finitism. I have a Thomson’s Lamp, i.e., a lamp with a toggle switch that can survive infinitely many togglings. I have access to it every day at the following times: 10:30, 10:45, 10:52.5, and so on. Each day, at 10:00 the lamp is off, and nobody else has access to the machine. At each time when I have access to the lamp, I can either toggle or not toggle its switch.

I now experiment with the lamp by trying out various supertasks (perhaps by programming a supertask machine), during which various combinations of toggling and not toggling happen. For instance, I observe that if I don’t ever toggle the switch, the lamps stays off. If I toggle it a finite number of times, it’s on when that number is odd and off when that number is even. I also notice the following regularities about cases where an infinite number of togglings happens:

  1. The same sequence (e.g., toggle at 10:30, don’t toggle at 10:45, toggle at 10:52.5, etc.) always produces the same result.

  2. Reversing a finite number of decisions in a sequence produces the same outcome when an even number of decisions is reversed, and the opposite outcome when an odd number of decisions is reversed.

(Of course, 1 is a special case of 2.) How fun! I conclude that 1 and 2 are always going to be true.

Now I set up a supertask machine. It will toss a fair coin just prior to each of my lamp access times, and it will toggle the switch if the coin is heads and not toggle it if it is tails.

Question: What is the probability that the lamp will be on at 11?

“Answer:” Given 1 and 2, the event that the lamp will be on at 11 is not measurable with respect to the standard (completed) product measure on a countable infinity of coin tosses. (See note 1 here.)

So, given supertasks (and hence the falsity of causal finitism), we could find ourselves in a position where we would have to deal with a non-measurable set.

Non-measurable sets and intuition

Here’s an interesting reason to accept the existence of non-measurable sets (and hence of whatever weak version of the Axiom of Choice that it depends on). A basic family of mathematical results in analysis says that most measurable real-valued functions on the real line are “close to” being continuous, i.e., that they can be approximated by continuous functions in some appropriate sense. But it is intuitive to think that there “should” be real-valued functions on the real line that are not close to being continuous—there “should” be functions that are very, very messy. So, intuitively, there should be non-measurable functions, and hence non-measurable sets.

Thursday, September 7, 2017

Two kinds of non-measurable events

Non-measurable events are ones to which the probability function in the situation assigns no probability. Philosophically speaking, non-measurable events come in two varieties:

  1. Non-measurable events that should not have any probability assignment.

  2. Non-measurable events that should have a probability assignment.

Type (1) non-measurable events are the kinds of weird events that can be constructed from the Hausdorff and Banach-Tarski paradoxes, as well as perhaps (this is less clear) the Vitali non-measurable sets.

But I think there are also type (2) non-measurable events relative to standard choices of probability functions. For instance, suppose that in each universe of an infinite multiverse a fair coin is tossed countably infinitely often.

How likely is it that in at least one universe all the coin tosses are heads? If the universes form a countable infinity, classical probability theory gives an answer: zero. But if the universes form an uncountable infinity, classical probability theory gives no answer at all—the standard completed product measure makes the event be non-measurable. However, intuitively, there should be an answer in at least some cases. If the number of universes is much larger than the number of possible countable sequences of coin tosses (i.e., is much larger than 2ω), we would expect the probability to be 1 or close to it. We can coherently extend the standard probability function to give that answer. But we can also coherently extend it to give a different answer, including the answer that the probability of an all-heads universe is zero, even if the number of universes is a gigantic infinite cardinality.

We don’t want to just make up an answer here. We want the answer to be derivable in some way resembling the proof of the theorem that if you toss a coin infinitely many times, you’ve got probability 1 of getting heads at least once.

I suppose we could take it to be a metaphysical axiom that if you have K disjoint collections each with M coin tosses, then if K and M are infinite and K > M, then with probability one at least one collection yields all heads. But it would be nice to have more than just intuition here, and in similar problems.

Thursday, August 10, 2017

Uncountable independent trials

Suppose that I am throwing a perfectly sharp dart uniformly randomly at a continuous target. The chance that I will hit the center is zero.

What if I throw an infinite number of independent darts at the target? Do I improve my chances of hitting the center at least once?

Things depend on what size of infinity of darts I throw. Suppose I throw a countable infinity of darts. Then I don’t improve my chances: classical probability says that the union of countably many zero-probability events has zero probability.

What if I throw an uncountable infinity of darts? The answer is that the usual way of modeling independent events does not assign any meaningful probabilities to whether I hit the center at least once. Indeed, the event that I hit the center at least once is “saturated nonmeasurable”, i.e., it is nonmeasurable and every measurable subset of it has probability zero and every measurable superset of it has probability one.

Proposition: Assume the Axiom of Choice. Let P be any probability measure on a set Ω and let N be any non-empty event with P(N)=0. Let I be any uncountable index set. Let H be the subset of the product space ΩI consisting of those sequences ω that hit N, i.e., ones such that for some i we have ω(i)∈N. Then H is saturated nonmeasurable with respect to the I-fold product measure PI (and hence with respect to its completion).

One conclusion to draw is that the event H of hitting the center at least once in our uncountable number of throws in fact has a weird “nonmeasurable chance” of happening, one perhaps that can be expressed as the interval [0, 1]. But I think there is a different philosophical conclusion to be drawn: the usual “product measure” model of independent trials does not capture the phenomenon it is meant to capture in the case of an uncountable number of trials. The model needs to be enriched with further information that will then give us a genuine chance for H. Saturated nonmeasurability is a way of capturing the fact that the product measure can be extended to a measure that assigns any numerical probability between 0 and 1 (inclusive) one wishes. And one requires further data about the system in order to assign that numerical probability.

Let me illustrate this as follows. Consider the original single-case dart throwing system. Normally one describes the outcome of the system’s trials by the position z of the tip of the dart, so that the sample space Ω equals the set of possible positions. But we can also take a richer sample space Ω* which includes all the possible tip positions plus one more outcome, α, the event of the whole system ceasing to exist, in violation of the conservation of mass-energy. Of course, to be physically correct, we assign chance zero to outcome α.

Now, let O be the center of the target. Here are two intuitions:

  1. If the number of trials has a cardinality much greater than that of the continuum, it is very likely that O will result on some trial.

  2. No matter how many trials—even a large infinity—have been performed, α will not occur.

But the original single-case system based on the sample space Ω* does not distinguish O and α probabilistically in any way. Let ψ be a bijection of Ω* to itself that swaps O and α but keeps everything else fixed. Then P(ψ[A]) = P(A) for any measurable subset A of Ω* (this follows from the fact that the probability of O is equal to the probability of α, both being zero), and so with respect to the standard probability measure on Ω*, there is no probabilistic difference between O and α.

If I am right about (1) and (2), then what happens in a sufficiently large number of trials is not captured by the classical chances in the single-case situation. That classical probabilities do not capture all the information about chances is something we should already have known from cases involving conditional probabilities. For instance P({O}|{O, α}) = 1 and P({α}|{O, α}) = 0, even though O and α are on par.

One standard solution to conditional probability case is infinitesimals. Perhaps P({α}) is an infinitesimal ι but P({O}) is exactly zero. In that case, we may indeed be able to make sense of (1) and (2). But infinitesimals are not a good model on other grounds. (See Section 3 here.)

Thinking about the difficulties with infinitesimals, I get this intuition: we want to get probabilistic information about the single-case event that has a higher resolution than is given by classical real-valued probabilities but lower resolution than is given by infinitesimals. Here is a possibility. Those subsets of the outcome space that have probability zero also get attached to them a monotone-increasing function from cardinalities to the set [0, 1]. If N is such a subset, and it gets attached to it the function fN, then fN(κ) tells us the probability that κ independent trials will yield at least one outcome in N.

We can then argue that fN(κ) is always 0 or 1 for infinite. Here is why. Suppose fN(κ)>0. Then, κ must be infinite, since if κ is finite then fN(κ)=1 − (1 − P(N))κ = 0 as P(N)=0. But fN(κ + κ)=(fN(κ))2, since probabilities of independent events multiply, and κ + κ = κ (assuming the Axiom of Choice), so that fN(κ)=(fN(κ))2, which implies that fN(κ) is zero or one. We can come up with other constraints on fN. For instance, if C is the union of A and B, then fC(κ) is the greater of fA(κ) and fB(κ).

Such an approach could help get a solution to a different problem, the problem of characterizing deterministic causation. To a first approximation, the solution would go as follows. Start with the inadequate story that deterministic causation is chancy causation with chance 1. (This is inadequate, because in the original dart-throwing case, the chance of missing the center is 1, but throwing the dart does not deterministically cause one to hit a point other than the center.) Then say that deterministic causation is chancy causation such that the failure event F is such that fF(κ)=0 for every cardinal κ.

But maybe instead of all this, one could just deny that there are meaningful chances to be assigned to events like the event of uncountably many trials missing or hitting the center of the target.

Sketch of proof of Proposition: The product space ΩI is the space of all functions ω from I to Ω, with the product measure PI generated by the product measures of cylinder sets. The cylinder sets are product sets of the form A = ∏iIAi such that there is a finite J ⊆ I such that Ai = Ω for i ∉ J, and the product measure of A is defined to be ∏iJP(Ai).

First I will show that there is an extension Q of PI such that Q(H)=0 (an extension of a measure is a measure on a larger σ-algebra that agrees with the original measure on the smaller σ-algebra). Any PI-measurable subset of H will then have Q measure zero, and hence will have PI-measure zero since Q extends PI.

Let Q1 be the restriction of P to Ω − N (this is still normalized to 1 as N is a null set). Let Q1I be the product measure on (Ω − N)I. Let Q be a measure on Ω defined by Q(A)=Q1I(A ∩ ΩN). Consider a cylinder set A = ∏iIAi where there is a finite J ⊆ I such that Ai = Ω whenever i ∉ J. Then
Q(A)=∏iJQ1(Ai − N)=∏iJP(Ai − N)=∏iJP(Ai)=PN(A).
Since PN and Q agree on cylinder sets, by the definition of the product measure, Q is an extension of PN.

To show that H is saturated nonmeasurable, we now only need to show that any PI-measurable set in the complement of H must have probability zero. Let A be any PI-measurable set in the complement of H. Then A is of the form {ω ∈ ΩI : F(ω)}, where F(ω) is a condition involving only coordinates of ω numbered by a fixed countable set of indices from I (i.e., there is a countable subset J of I and a subset B of ΩJ such that F(ω) if and only if ω|J is a member of B, where ω|J is the restriction of ω to J). But no such condition can exclude the possibility that a coordinate of Ω outside that countable set is in H, unless the condition is entirely unsatisfiable, and hence no such set A lies in the complement of H, unless the set is empty. And that’s all we need to show.

Friday, February 12, 2016

The Principle of Sufficient Reason and Probability

I just posted the paper here. Forthcoming in Oxford Studies in Metaphysics.

Abstract: I shall argue that considerations about frequency-to-chance inferences make very plausible some a localized version of the Principle of Sufficient Reason (PSR). But a localized version isn’t enough, and so we should accept a full PSR.

Friday, November 13, 2015

Bayesian divergence

Suppose I am considering two different hypotheses, and I am sure exactly one of them is true. On H, the coin I toss is chancy, with different tosses being independent, and has a chance 1/2 of landing heads and a chance 1/2 of landing tails. On N, the way the coin falls is completely brute and unexplained--it's "fundamental chaos", in the sense of my ACPA talk. So, now, you observe n instances of the coin being tossed, about half of which are heads and half of which are tails. Intuitively, that should support H. But if N is an option, if the prior probability of N is non-zero, we actually get Bayesian divergence as n increases: we get further and further from confirmation of H.

Here's why. Let E be my total evidence--the full sequence of n observed tosses. By Bayes' Theorem we should have:

P(H|E) = P(E|H)P(H)/[P(E|H)P(H) + P(E|N)P(N)].
But there is a problem: P(E|N) is undefined. What shall we do about this? Well, it is completely undefined. Thus, we should take it to be an interval of probabilities, the full interval [0,1] from 0 to 1. The posterior probability P(H|E), then, will also be an interval between:
P(E|H)P(H)/[P(E|H)P(H) + (0)·P(N)] = 1
and
P(E|H)P(H)/[P(E|H)P(H) + (1)·P(N)] ≤ P(E|H)/P(N) = 2n / P(N).
(Remember that E is a sequence of n fair and independent tosses if H is true.) Thus, as the number of observations increases, the posterior probability for the "sensible" hypothesis H gets to be an interval [a,1], where a is very small. But something whose probability is almost the whole interval [0,1] is not rationally confirmed. So the more data we have, the further we are from confirmation.

This means that no-explanation hypotheses like N are pernicious to Bayesians: if they are not ruled out as having zero or infinitesimal probability from the outset, they undercut science in a way that is worse and worse the more data we get.

Fortunately, we have the Principle of Sufficient Reason which can rule out hypotheses like N.

Friday, March 20, 2015

Are there more nonmeasurable sets than measurable ones?

Assume the Axiom of Choice.

Consider the subsets of the interval [0,1]. There 2c subsets of [0,1], where c is the cardinality of the continuum. Since the Cantor set C has measure zero, and every subset of a measure zero set is (Lebesgue) measurable, there are at least as many measurable subsets of [0,1] as subsets of the Cantor set. But the Cantor set has cardinality c, so there are 2c measurable subsets of [0,1].

On the other hand, however, if A is any nonmeasurable set, then adding and/or subtracting a set of measure zero to A won't change the nonmeasurability, so there are 2c nonmeasurable sets as well, since there are 2c ways of varying A by adding and/or substracting a subset of C.

But notice the curious role that the Cantor set played in my above arguments. The reason there are 2c measurable sets and 2c nonmeasurable sets is because there are 2c sets of measure zero, and we can vary any set by a measure zero set and preserve measurability / nonmeasurability.

This suggests another question. Say that two sets are equivalent provided that they differ by a set of measure zero. Are there more equivalence classes of nonmeasurable sets than of measurable sets?

Assuming the Axiom of Choice, the answer is yes. In fact:

  • Theorem: There are 2c equivalence classes of nonmeasurable sets while there are only c classes of measurable sets.

Why? Well, any measurable set differs from a Borel set by a set of measure zero, so there are no more equivalence classes of measurable sets than there are Borel sets, and there are c Borel sets. (And of course there are no fewer than c equivalence classes measurable sets: just consider the intervals [0,x].) Responding to a question I had asked on MathOverflow, Eric Wofsey yesterday showed that there are 2c equivalence classes of subsets of [0,1]. Since there are only c equivalence classes of measurable ones, there must be 2c equivalence classes of nonmeasurable ones.

In fact, we can say something stronger. Say that a subset A of [0,1] is saturated nonmeasurable (I used to call this "maximally nonmeasurable") provided that it is not only nonmeasurable but that every measurable subset of A has measure zero and every measurable superset of A has measure one. In other words, there is nothing measurable about A: we can't even give a non-trivial range of Lebesgue measures for A. (Technically: the inner measure is zero and the outer measure is one.) Then:

  • Theorem: There are 2c equivalence classes of saturated nonmeasurable sets.

Why? Well, this uses the same trick that Wofsey's proof did. Sierpiński showed in 1938 that you can find a disjoint family of perfect, and hence of cardinality c, subsets of the square [0,1]2 with the property that any set obtained by choosing one member from each element in the family is saturated nonmeasurable. Since as measure spaces [0,1]2 is isomorphic to [0,1], there will be a disjoint family F of subsets of [0,1], each of cardinality c, such that choosing one member of each element in F yields a saturated nonmeasurable set. For any AF, let φA be a one-to-one function from [0,1] to A. Define Ux={φA(x):AF} for x∈[0,1]. Then each Ux has inner measure zero and outer measure one, and the sets Ux are disjoint for different values of x. Any union of the sets Ux that (a) contains at least one of the sets (i.e., isn't empty) and (b) doesn't contain all the sets then also has inner measure zero and outer measure one. Why? Well, since it contains at least one of the sets, its outer measure is one. And since there is at least one that it doesn't contain, its complement has outer measure one and hence it itself has inner measure zero. But there are 2c such unions (since there are 2c possible unions given c nonempty disjoint sets, and throwing out the union that has no sets and the one that has all sets doesn't change that). Moreover, any two distinct such unions differ by at least one of the sets Ux, and hence differ by a set with inner measure one (and outer measure one for that matter), and so are not equivalent up to sets of measure zero.

Friday, May 23, 2014

Thomson's lamp and the Axiom of Choice

Consider Thomson's lamp: a lamp with a pushbutton switch that toggles it on and off. The lamp starts in the off position, and then in the next half minute the button is pressed, and in the next quarter it is pressed again, and then in the neight eighth again, and so on. Then at the end of the supertask, the lamp is either on or off.

Now keep the lamp but change the story. During each of the ever shorter intervals, a coin is flipped and the switch is pressed if it lands heads, and not pressed if it lands tails. Moreover, the final state of the lamp depends on the results of the coin flips in the following ways:

  1. The results of the coin flips determine the final state of the lamp.
  2. For any sequence of coin flip results, if any one (and only one) coin flip had a different, the lamp's final state would have been different, too.
Surprisingly, the existence of a lamp that would work in this way implies a version of the Axiom of Choice. To see this, notice that if the coin flips are independent and fair, then the subset of the probability space where the lamp's final state is on is nonmeasurable.[note 1]

But of course, on some technical assumptions, the existence of a nonmeasurable set requires a version of the Axiom of Choice. So if we read the Thomson's lamp story in such a way that the final outcome is determined by which presses are made and which aren't, in such a way that changing a single press changes the final outcome, that story seems to commit us to a version of the Axiom of Choice.

Conversely, it is easy to use the Axiom of Choice for pairs to prove the existence of a function such as would be implemented by the lamp.[note 2]

Thursday, December 12, 2013

Wednesday, February 6, 2013

An argument for a version of the Axiom of Choice

This is an argument for the Axiom of Choice where the sets we're choosing from are all subsets of the real numbers. The argument needs the notion of really independent random processes. Real independence is not just probabilistic independence (if you're not convinced, read this). I don't know how to characterize real independence, but here is a necessary condition for it. If S is a collection of really independent processes producing outcomes, and for each s in S, Us is a non-empty subset of the range of s (where the range of a process here is all the outcomes that it can generate) then it is metaphysically possible that each member s of S produces an outcomes in Us. (This need not hold for merely probabilistically independent processes.)

Now, let U be a set of disjoint non-empty subsets of the real numbers R. Let N be the cardinality of U. It is surely metaphysically possible to have N really independent random processes each of which has range R. For instance, one might have a multiverse with N universes, in each of which there is a random process that produces a particle at a normally distributed point from the emitter, and the outcome of the process can be taken to be the x-coordinate of where the particle is produced.

Now, there is a one-to-one correspondence between the members of U and the random processes. If r is one of the random processes, let Ur be the member of U that corresponds to it (after fixing one such correspondence). By real independence, it is metaphysically possible that for all r, the outcome of r is in Ur. Take a world w where this is the case. In that world, the set of outcomes of our processes will contain exactly one member from each member of U, and hence will be a choice set. But what sets of real numbers there are surely does not differ between worlds (I can imagine questioning this, though). So if in w there is a choice set, there actually is a choice set.

Granted, this only gives us the Axiom of Choice for subsets of the reals. But that's enough to generate the Banach-Tarski, Hausdorff and Vitali non-measurable sets. It's paradoxical enough.