Showing posts with label Newcomb Paradox. Show all posts
Showing posts with label Newcomb Paradox. Show all posts

Friday, April 24, 2015

Blackmail, promises and self-punishment

I was reading this interesting paper which comes up with "blackmail" stories against both evidential and causal decision theory (CDT). I'll focus on the causal case. The paper talks about an Artificial Intelligence context, but we can transpose the stories into something more interpersonal. John blackmails Patrick in such a way that it's guaranteed that if Patrick pays up there will be no more blackmail. As a good CDT agent, Patrick pays up, since it pays. However, Patrick would have been better off if he were the sort of person who refuses to pay off blackmailers. For John is a very good predictor of Patrick's behavior, and if John foresaw that Patrick would be unlikely to pay him off, then John wouldn't have taken the risk of blackmailing Patrick. So CDT agents are subject to blackmail.

One solution is to add to the agent's capabilities the ability to adopt a policy of behavior. Then it would have paid for Patrick to have adopted a policy of refusing to pay off blackmailers and he would have adopted that policy. One problem with this, though, is that the agent could drop the policy afterwards, and in the blackmail situation it would pay to drop the policy. And that makes one subject to blackmail once again. (This is basically the retro-blackmail story in the paper.)

Anyway, thinking about these sorts of cases, I've been playing with a simplistic decision-theoretic model of promises and weak promises—or, more generally, commitments. When one makes a commitment, then on this model one changes one's utility function. The scenarios where one fails to fulfill the commitment get a lower utility, while scenarios where one succeeds in fulfilling the commitment are unchanged in utility. You might think that you get a utility bonus for fulfilling a commitment. That's mistaken. For if we got a utility bonus for fulfilling commitments, then we would have reason to promise to do all sorts of everyday things that we would do anyway, like eat breakfast.

This made me think about agents who have a special normative power: the power to lower their utility function in any way that they like. But they lack the power to raise it. In other words, they have the power to replace their utility function by a lower one. This can be thought of in terms of commitments—lowering the utility value of a scenario by some amount is equivalent to making a commitment of corresponding strength to ensure that scenario isn't actualized—or in terms of mechanisms for self-punishment. Imagine an agent who can make robots that will zap him in various scenarios.

Now, it would be stupid for an agent simply to lower his utility function by a constant amount everywhere. That wouldn't change the agent's behavior at all, but would make sure that the agent is less well off no matter what happens. However, it wouldn't be stupid for the agent to lower his utility function for scenarios where he gives in to blackmail by agents who can make good predictions of his behavior and who wouldn't have blackmailed him if they thought he wouldn't give in. If he lowers that utility enough—say, by making a promise not to negotiate with blackmailers or by generating a robot that zaps him painfully if he gives in—then a blackmailer like John will know that he is unlikely to give in to blackmail, and hence won't risk blackmailing him.

The worry about the agent changing policies and thereby opening oneself to blackmail does not apply on this story. For the agent in my model has only been given the power to lower his utility function at will. He doesn't have the power to raise it. If the agent were blackmailed, he could lower his utility function for the scenarios where he doesn't give in, and thereby get himself to give in. But it doesn't pay to do that, as is easy to confirm. It would pay for him to raise his utility function for the scenarios where he gives in, but he can't do that.

An agent like this would likewise give himself a penalty for two-boxing in Newcomb cases.

So it's actually good for agents to be able to lower their utility function. Setting up self-punishments can make perfect rational sense, even in the case of a perfect rational agent, so as to avoid blackmail.

Wednesday, April 15, 2015

A paradox about prediction of belief

Sally is perfectly honest, knows for sure whether there has ever been life on Mars (she's just finished an enormous amount of NASA data analysis), and is a perfect predictor of my future beliefs. She then informs me that she knows what I will believe at midnight about whether there was once life on Mars, and she further informs me that:

  1. There was once life on Mars if and only at midnight tonight I will fail to believe that there was once life on Mars.
Moreover, I know that:
  1. I won't get any other evidence relevant to whether there was once life on Mars.
I'd love to know whether there was once life on Mars. I start off thinking:
Well, right now I have no belief either way, and I am unlikely to get any evidence before midnight. So by midnight I will also have no belief either way. And thus by Sally's information there was once life on Mars.
But of course as soon as I accept this argument, I start to believe that there was life on Mars. And I know that if I keep on believing this until midnight, then my belief is false. I quickly see the pattern, and I realize that I don't know what to think! But when I don't know what to think, I default to suspension of judgment. But this, too, leads me astray: For as soon as I think that the appropriate rational attitude for me is suspension of judgment, then I start thinking I will suspend judgment at midnight, and I then conclude that there was once life on Mars. And the circle starts again.

Now, I know I'm not perfectly rational. So I can get out of the circle by concluding that given how confusing this case is, I am probably not going to act rationally. So something non-rational will affect my beliefs by midnight, and I don't know what that will be, so I might as well not speculate until that happens. Sally knows what it will be, but I don't.

But suppose I am perfectly rational. I shall assume that a part of perfect rationality is knowing for sure you're perfectly rational, knowing for sure what you belief, and drawing all the right conclusions from one's evidence. What should I believe in the above case?

Monday, October 27, 2014

Yet another infinite population problem

There are infinitely many people in existence, unable to communicate with one another. An angel makes it known to all that if, and only if, infinitely many of them make some minor sacrifice, he will give them all a great benefit far outweighing the sacrifice. (Maybe the minor sacrifice is the payment of a dollar and the great benefit is eternal bliss for all of them.) You are one of the people.

It seems you can reason: We are making our decisions independently. Either infinitely many people other than me make the sacrifice or not. If they do, then there is no gain for anyone to my making it—we get the benefit anyway, and I unnecessarily make the sacrifice. If they don't, then there is no gain for anyone to my making it—we don't get the benefit even if I do, so why should I make the sacrifice?

If consequentialism is right, this reasoning seems exactly right. Yet one better hope that it's not the case that everyone reasons like this.

The case reminds me of both the Newcomb paradox—though without the need for prediction—and the Prisoner's Dilemma. Like in the case of the Prisoner's Dilemma, it sounds like the problem is with selfishness and freeriding. But perhaps unlike in the case of the Prisoner's Dilemma, the problem really isn't about selfishness.

For suppose that the infinitely many people each occupy a different room of Hilbert's Hotel (numbered 1,2,3,...). Instead of being asked to make a sacrifice oneself, however, one is asked to agree to the imposition of a small inconvenience on the person in the next room. It seems quite unselfish to reason: My decision doesn't affect anyone else's (I so suppose—so the inconveniences are only imposed after all the decisions have been made). Either infinitely many people other than me will agree or not. If so, then we get the benefit, and it is pointless to impose the inconvenience on my neighbor. If not, then we don't get the benefit, and it is pointless to add to this loss the inconvenience to my neighbor.

Perhaps, though, the right way to think is this: If I agree—either in the original or the modified case—then my action partly constitutes the a good collective (though not joint) action. If I don't agree, then my action runs a risk of partly constituting a bad collective (though not joint) action. And I have good reason to be on the side of the angels. But the paradoxicality doesn't evaporate.

I suspect this case, or one very close to it, is in the literature.

Friday, October 17, 2014

Too late!

Let's say that something very good will happen to you if and only if the universe is in state S at midnight today. You labor mightily up until midnight to make the universe be in S. But then, surely, you stop and relax. There is no point to anything you may do after midnight with respect to the universe being in S at midnight, except for prayer or research on time machines or some other method of affecting the past. It's too late for anything else!

This line of thought immediately implies two-boxing in the Newcomb's Paradox. For suppose that the predictor will decide on the contents of the boxes on the basis of her predictions tonight at midnight about your actions tomorrow at noon when you will be shown the two boxes. Her predictions are based on the state of the universe at midnight. Let S be the state of the universe being such as to make her predict that you will engage in one-boxing. Then until midnight you will labor mightily to make the universe be in S. You will read the works of epistemic decision theorists, and shut out from your mind the two-boxers' responses. But then midnight strikes. And then, surely, you stop and relax. There is no point to anything you may do after midnight with respect to whether the universe was in S at midnight or not, except for prayer or research on time machines or some other method of affecting the past, and in the Newcomb paradox one normally stipulates that such techniques are not relevant. In particular, with respect to the universe being in S at midnight tonight, it makes no sense to choose a single box tomorrow at noon. So you might as well choose two. Though, if you're lucky, by midnight tonight you will have got yourself into such a firm spirit of one-boxing that by noon tomorrow you will be blind to this thought and will choose only one box.

Tuesday, May 18, 2010

Is there a cost to two-boxing in Newcomb cases?

The following may be well-known to folks in the field, or it may be well-known to them to be mistaken.

It seems to be acknowledged by both sides that it is a significant cost of being causal decision theorist that one has to two-box in the Newcomb Paradox, and hence one loses out in Newcomb cases.

I suspect the causal decision theorist should not grant that this is a cost of the theory, or at least not a significant one. First, distinguish between Newcomb cases where there is weak counterfactual dependence between what one chooses and what the predictor predicted and ones where there is no weak counterfactual dependence. (I say that B weakly counterfactually depends on A provided that B might not have happened had A not happened.)

Case A: Weak counterfactual dependence. David Lewis thought that wherever there was counterfactual dependence between non-overlapping events, there was causal dependence. I think he was wrong. E.g., that God promised A counterfactually depends on A, since the non-occurrence of A entails the non-occurrence of the divine promise, at least given prior linguistic conventions. Nonetheless, I think that counterfactual dependence, and even weak counterfactual dependence, is generally a pretty good indicator of at least explanatory dependence. Perhaps something like this principle is true: If (a) B weakly counterfactually depends on A and (b) there is no event C such that (b1) C is not weakly counterfactually dependent on A and (b2) C together with B entails A, then B is explanatorily dependent on A. (Condition (b) rules out the divine promise case: let C be the linguistic conventions that God's act of promising depends on.) Maybe some additional conditions are needed. Nonetheless, I am optimistic that in the Newcomb case, from the existence of a weak counterfactual dependence between my choice and the prediction it follows that my choice is explanatorily prior to the prediction. (Condition (b) is satisfied, at least assuming one's choice is indeterministic. I don't know what to do about deterministic choices. I don't know if the notion makes sense. Nuel Belnap once said that to make a choice, you have to have choices.) But since the causal decision theorist should be willing to generalize causal dependence to explanatory dependence, the causal theorist in this case says to one-box.

Case B: No counterfactual dependence. In this case, the following counterfactual is true for Tamara the two-boxer: She would have got less had she chosen only one box. For had she chosen only one box, the predictor would still have made the same prediction that was in fact made, and Tamara would have done poorly.

But wouldn't the two-boxer have gotten a lot more had she been a one-boxer? Yes. Here, we need to distinguish two antecedents of counterfactuals:

  1. Tamara has a disposition to one-box in Newcomb cases like this one
  2. Tamara one-boxed in this Newcomb case.
And so:
  1. Had (1) been true, Tamara would have got a lot more than she got.
  2. Had (2) been true, Tamara would have got a bit less than she got.
This means that we need to distinguish between two different questions of rationality. Should Tamara have a disposition to one-box and should Tamara one-box in this case?

Perhaps the causal theorist could say: Tamara should have a disposition to one-box and should two-box in this case. There is nothing absurd about that sort of an answer, and according to SEP, it's a standard move here. There are, after all, standard cases where it is (narrowly self-interestedly) rational to be irrational, such as where you will be killed if you are thought to be a rational witness to a crime.

But there may be a better answer. Fact (3) does not entail that Tamara should have a disposition to one-box. All that fact (3) tells us is that if the only situations Tamara faces are Newcomb ones, then she would do better to have a disposition to one-box. But there are other possible situations where someone with a disposition to two-box would do better. For instance, cases where an evil two-boxing philosopher becomes a dictator and kills everyone who doesn't have a disposition to two-box. (Or cases where people are mistaken in thinking the predictor accurate.) So in Newcomb situations, Tamara would do better were she by disposition a one-boxer. But in crazy two-boxing philosopher situations, Tamara would do better were she by disposition a two-boxer. In other words, both of the competing theories have the consequence that in some worlds it is better to self-induce a disposition to abide by the opposite theory. And hence there seems to be little or no special cost here to being a causal decision theorist.