Showing posts with label moral risk. Show all posts
Showing posts with label moral risk. Show all posts

Friday, March 27, 2026

We should not make human-like AI

  1. Either AI with human-like behavior is a person or is not.

  2. We should not deliberately produce AI with human-like behavior when that AI is a person.

  3. We should not deliberately produce AI with human-like behavior when that AI is not a person.

  4. So, we should not deliberately produce AI with human-like behavior.

Obviously, the big question is whether (2) and (3) are true.

In favor of (3), a non-person with human-like behavior evokes emotional responses from us that are only apt as directed at persons. These emotional responses are of great moral importance to our life, and some of them are constitute a recognition of the object of the emotion as a being with dignity, a sacred being, and to have such emotional responses to something that lacks the relevant dignity blurs the central moral distinction between persons and non-persons.

In favor of (2), we have several arguments. Start with some Kantian ones. AI is an artifact. When we make artifacts, we make them to serve our purposes. To make a person to serve our purposes is to treat the person as a mere means to an end. And that’s wrong. This argument, I think, applies no matter what our purpose is: even if our purpose is to make the AI live its own free life. That’s still our purpose for it, and we have no right to impose a purpose on a person’s life.

Furthermore, by designing a digital person, we are designing, in its fundamentals what its basic purposes in life are and we are thus exercising a mode of control over another person that we have no right to have.

A final but least principled Kantian argument is that even if we “set free” the digital person—whatever exactly that means—it is pretty much impossible to protect digital persons from being enslaved by other humans.

There are also some non-Kantian arguments. If an AI is a person, that person will eventually be very cheap to keep in decent existence as compared to a flesh and blood human being. Since we have the duty to protect the life of a person when doing so is not an undue load on limited resources, we would have the duty to keep any person AI that we spin up running indefinitely. This is problematic as it ties the hands of future human generations, by imposing on them what one might call “an unnatural duty” to keep running all the AI persons that we make, with little benefit to the future human generations from this. The problem is the worse the greater the numbers of digital persons that we spin up. This problem is akin to the problem of frozen embryos in fertility clinics if these embryos are persons (which I think they are).

There is also a dilemma. We have the duty to protect the life of persons. But at the same time, there is something deeply unappealing about something of the level of sophistication of an ordinary human mental life extending an order of magnitude longer than the typical human life-span (rescaled as needed to take account of differences in processing speed). The most appealing religious accounts of afterlife involve a radical transformation, e.g., theosis or parinirvana. I think many of us rightly would feel that living a thousand years of the kind of life we now have isn’t appealing, though we wouldn’t mind an extra ten or twenty or maybe even hundred years. If we were ever able to indefinitely extend human life without an undue resource cost, we would find ourselves in an inextricable moral dilemma: on the one hand a duty to protect life when doing so does not carry an undue resource cost and on the other hand the monstrousness of living an order of magnitude longer than ordinary human life should be. But with digital persons, indefinite life extension would be easy, and so the dilemma would be unavoidable.

Finally, we would take a significant moral risk in designing digital persons. Training processes involve vasts amount negative feedback. We just do not know how unpleasant that might be.

Thursday, February 18, 2021

Moral risk

Say that an action is deontologically doubtful (DD) provided that the probability of the action being forbidden by the correct deontology is significant but less than 1/2.

There are cases where we clearly should not risk performing a DD action. A clear example is when you’re hunting and you see a shape that has a 40% chance of being human: you should not shoot. But notice that in this case, deontology need play no role: expected-utility reasoning tells you that you shouldn’t shoot.

There are, on the other hand, cases where you should take a significant risk of performing a DD action.

Beast Case: The shape in the distance has a 30% chance of being human and a 70% chance of being a beast that is going to devour a dozen people in your village if not shot by you right now. In that case, it seems it might well be permissible to shoot.

This suggests this principle:

  1. If a DD action has significantly higher expected utility than refraining from the action, it is permissible to perform it.

But this is false. I will assume here the standard deontological claim that it is wrong to shoot one innocent to save two.

Villain Case: You are hunting and you see a dark shape in the woods. The shape has a 40% chance of being an innocent human and a 60% chance of being a log. A villain who is with you has just instructed a minion to go and check in a minute on the identity of the shape. If the shape turns out to be a human, the minion is to murder two innocents. You can’t kill the villain or the minion, as they have bulletproof jackets.

The expected utility of shooting is significantly higher than of refraining from the action. If you shoot, the expected lives lost are (0.4)(1)=0.4, and if you don’t shoot the expected lives lost are (0.4)(2)=0.8. So shooting has an expected utility that’s 0.4 lives better than not shooting. But it is also clear, assuming the deontological claim that it is wrong to kill one to save two, that it is wrong to shoot in this case.

What is different from the villain case and the dangerous beast case is that in the Villain Case, the difference in expected utilities comes precisely from the scenario where the shape is human. Intuition suggests we should tweak (1) to evaluate expected utilities in a way that ignores the good effects of deontologically forbidden things. This tweak does not affect the Beast Case, but it does affect the Villain Case, where the difference in utilities came precisely from counting the life-saving benefits of killing the human.

I don’t know how to precisely formulate the tweaked version of (1), and I don’t know if it is sufficiently strong to covere all cases.

Wednesday, July 15, 2020

Catastrophic decisions

Kirk has come to a planet with two intelligent species in the universe, the Oligons and the Pollakons. There are a million Oligons and a trillion (i.e., million million) Pollakons. They are technologically unsophisticated, live equally happy lives on the same planet, but have no interaction with each other, and the universal translator is currently broken so Kirk can’t communicate with them either. A giant planetoid is about to graze the planet in a way that is certain to wipe out the Pollakons but leave the Oligons, given their different ecological niche, largely unaffected. Kirk can try to redirect the planetoid with his tractor beam. Spock’s accurate calculations give the following probabilities:

  • 1 in 1000 chance that the planetoid will now miss the planet and the Oligons and Pollakons will continue to live their happy lives;

  • 999 in 1000 chance that the planetoid will wipe out both the Oligons and the Pollakons.

If Kirk doesn’t redirect, expected utility is 106 happy lives (the Oligons). If Kirk does redirect, expected utility is (1/1000)(1012 + 106)=109 + 103 happy lives. So, expected utility clearly favors redirecting.

But redirecting just seems wrong. Kirk is nearly certain—99.9%—that redirecting will not help the Pollakons but will wipe out the Oligons.

Perhaps the reason intuition seems to favor not redirecting is that we have a moral bias in favor of non-interference. So let’s turn the story around. Kirk sees the planetoid coming towards the planet. Spock tells him that it has a 1/1000 chance that nothing bad will happen, and a 999/1000 chance that it will wipe out all life on the planet. But Spock also tells him that he can beam the Oligons—but not the Pollakons, who are made of a type of matter incapable of beaming—to the Enterprise. Spock, however, also tells Kirk that beaming the Oligons on board will require the Enterprise to come closer to the planet, which will gravitationally affect the planetoid’s path in such a way that the 1/1000 chance of nothing bad happening will disappear, and the Pollakons will now be certain, and not merely 999/1000 likely, to die.

Things are indeed a bit less clear to me now. I am inclined to think Kirk should rescue the Oligons (this may require Double Effect), but I am worried that I am irrationally neglecting small probabilities. Still, I am inclined to think Kirk should rescue. If that intuition is correct, then even in other-concerning decisions, and even when we have no relevant deontological worries, we should not go with expected utilities.

But now suppose that Kirk over his career will visit a million such planets. Then a policy of non-redirection in the original scenario or of rescue in the modified scenario would be disastrous by the Law of Large Numbers: those 1/1000 events would happen a number of times, and many, many lives will be lost. If we’re talking about long-term policies, then, it seems that Kirk should have a policy of going with expected utilities (barring deontological concerns). But for single-shot decisions, I think it’s different.

This line of thought suggests two things to me:

  • maximization of expected utilities in ordinary circumstances has something to do with limit laws like the Law of Large Numbers, and

  • we need a moral theory on which we can morally bind ourselves to a policy, in a way that lets the policy override genuine moral concerns that would be decisive absent the policy (cf. this post on promises).

Tuesday, March 5, 2019

More on moral risk

You are the captain of a small damaged spaceship two light years from Earth, with a crew of ten. Your hyperdrive is failing. You can activate it right now, in a last burst of energy, and then get home. If you delay activating the hyperdrive, it will become irreparable, and you will have to travel to earth at sublight speed, which will take 10 years, causing severe disruption to the personal lives of the crew.

The problem is this. When such a failing hyperdrive is activated, everything within a million kilometers of the spaceship’s position will be briefly bathed in lethal radiation, though the spaceship itself will be protected and the radiation will quickly dissipate. Your scanners, fortunately, show no planets or spaceships within a million kilometers, but they do show one large asteroid. You know there are two asteroids that pass through that area of space: one of them is inhabited, with a population of 10 million, while the other is barren. You turn your telescope to the asteroid. It looks like the uninhabited asteroid.

So, you come to believe there is no life within a million kilometers. Moreover, you believe that as the captain of the ship who has a resposibility to get the crew home in a reasonable amount of time, unless of course this causes undue harm. Thus, you believe:

  1. You are obligated to activate the hyperdrive.

You reflect, however, on the fact that ship’s captains have made mistakes in asteroid identification before. You pull up the training database, and find that at this distance, captains with your level of training make the relevant mistake only once in a million times. So you still believe that this is the lifeless asteroid. but now you get worried. You imagine a million starship captains making the same kind of decision as you. As a result, 10 million crew members get home on time to their friends and families, but in one case, 10 million people are wiped out in an asteroid. You conclude, reasonably, that this is an unacceptable level of risk. One in a million isn’t good enough. So, you conclude:

  1. You are obligated not to activate the hyperdrive.

This reflection on the possibility of perceptual error does not remove your belief in (1), indeed your knowledge of (1). After all, a one in a million chance of error is less than the chance of error in many cases of ordinary everyday perceptual knowledge—and, indeed, asteroid identification just is a case of everyday perceptual knowledge for a captain like yourself.

Maybe this is just a case of your knowing you are in a real moral dilemma: you have two conflicting duties, one to activate the hyperdrive and the other not to. But this fails to account for the asymmetry in the case, namely that caution should prevail, and there has to be an important sense of “right” in which the right decision is not to activate the hyperdrive.

I don’t know what to say about cases like this. Here is my best start. First, make a distinction between subjective and objective obligations. This disambiguates (1) and (2) as:

  1. You are objectively obligated to activate the hyperdrive.

  2. You are subjectively obligated not to activate the hyperdrive.

Second, deny the plausible bridge principle:

  1. If you believe you are objectively obligated to ϕ, then you are subjectively obligated to ϕ.

You need to deny (4), since you believe (3), and if (4) were true, then it would follow you are subjectively obligated to activate the hyperdrive, and we would once again have lost sight of the asymmetric “right” on which the right thing is not to activate.

This works as far as it goes, though we need some sort of a replacement for (4), some other principle bridging from the objective to the subjective. What that principle is is not clear to me. A first try is some sort of an analogue to expected utility calculations, where instead of utilities we have the moral weights of non-violated duties. But I doubt that these weights can be handled numerically.

And I still don’t know how to handle is the problem of ignorance of the bridge principles between the objective and the subjective.

It seems there is some complex function from one’s total mental state to one’s full-stop subjective obligation. This complex function is one which is not known to us at present. (Which is a bit weird, in that it is the function that governs subjective obligation.)

A way out of this mess would be to have some sort of infallibilism about subjective obligation. Perhaps there is some specially epistemically illuminated state that we are in when we are subjectively obligated, a state that is a deliverance of a conscience that is at least infallible with respect to subjective obligation. I see difficulties for this approach, but maybe there is some hope, too.

Objection: Because of pragmatic encroachment, the standards for knowledge go up heavily when ten million lives are at stake, and you don’t know that the asteroid is uninhabited when lives depend on this. Thus, you don’t know (1), whereas you do know (2), which restores the crucial action-guiding asymmetry.

Response: I don’t buy pragmatic encroachment. I think the only rational process by which you lose knowledge is getting counterevidence; the stakes going up does not make for counterevidence.

But this is a big discussion in epistemology. I think I can avoid it by supposing (as I expect is true) that you are no more than 99.9999% sure of the risk principles underlying the cautionary judgment in (2). Moreover, the stakes go up for that judgment just as much as they do for (1). Hence, I can suppose that you know neither (1) nor (2), but are merely very confident, and rationally so, of both. This restores the symmetry between (1) and (2).

Friday, March 1, 2019

Between subjective and objective obligation

I fear that a correct account of the moral life will require both objective and subjective obligations. That’s not too bad. But I’m also afraid that there may be a whole range of hybrid things that we will need to take into account.

Let’s start with clear examples of objective and subjective obligations. If Bob promised Alice to give her $10 but I misremember the promise and instead thinks he promised never to give her any more, then:

  1. Bob is objectively required to give Alice $10.

  2. Bob is subjectively required not to give Alice any money.

These cases come from a mistake about particular fact. There are also cases arising from mistakes about general facts. Helmut is a soldier in the Germany army in 1944 who knows the war is unjust but mistakenly believes that because he is a soldier, he is morally required to kill enemy combatants. Then:

  1. Helmut is objectively required to refrain from shooting Allied combatants.

  2. Helmut is subjectively required to kill Allied combatants.

But there are interesting cases of mistakes elsewhere in the reasoning that generate curious cases that aren’t neatly classified in the objective/subjective schema.

Consider moral principles about what one should subjectively do in cases of moral risk. For instance, suppose that Carl and his young daughter are stuck on a desert island for the next three months. The island is full of chickens. Carl believes it is 25% likely that chickens have the same rights as humans, and he needs to feed his daughter. His daughter has a mild allergy to the only other protein source on the island: her eyes will sting and her nose run for the next three months if she doesn’t live on chicken. Carl thus thinks that if chickens have the same rights as humans, he is forbidden from feeding chicken to his daughter; but if they don’t, then he is obligated to feed chicken to her.

Carl could now accept one of these two moral risk principles (obviously, these will be derivative from more general principles):

  1. An action that has a 75% probability of being required, and a 25% chance of being forbidden, should always be done.

  2. An action that has a 25% probability of being forbidden with a moral weight on par with the prohibition on multiple homicides and a 75% probability of being required with a moral weight on par with that of preventing one’s child’s mild allergic symptoms for three months should never be done.

Suppose that in fact chickens have very little in the way of rights. Then, probably:

  1. Carl is objectively required to feed chicken to his daughter.

Suppose further that Carl’s evidence leads him to be sure that (5) is true, and hence he concludes that he is required to feed chicken to his daughter. Then:

  1. Carl is subjectively required to feed chicken to his daughter.

This is a subjective requirement: it comes from what Carl thinks about the probabilities of rights, moral principles about what what to do in cases of risk, etc. It is independent of the objective obligation in (7), though in this example it agrees with it.

But suppose, as is very plausible, that (5) is false, and that (6) is the right moral principle here. (To see the point, suppose that he sees a large mammal in the woods that would suffice to feed his daughter for three months. If the chance that that mammal is a human being is 25%, that’s too high a risk to take.) Then Carl’s reasoning is mistaken. Instead, given his uncertainty:

  1. Carl is required to to refrain from killing chickens.

But what kind of an obligation is (9)? Both (8) and (9) are independent of the objective facts about the rights of chickens and depend on Carl’s beliefs, so it sounds like it’s subjective like (8). But (8) has some additional subjectivity in it: (8) is based on Carl’s mistaken belief about what his obligations are in cases of mortal risk, while (9) is based on what Carl’s obligations (but of what sort?) “really are” in those cases.

It seems that (9) is some sort of a hybrid objective-subjective obligation.

And the kinds of hybrid obligations can be multiplied. For we could ask about what we should do when we are not sure which principle of deciding in circumstances of moral risk we should adopt. And we could be right or we could be wrong about that.

We could try to deny (9), and say that all we have are (7) and (8). But consider this familiar line of reasoning: Both Bob and Helmut are mistaken about their obligations; they are not mistaken about their subjective obligations; so, there must be some other kinds of obligations they are mistaken about, namely objective ones. Similarly, Carl is mistaken about something. He isn’t mistaken about his subjective obligation to feed chicken. Moreover, his mistake does not rest in a deviation between subjective and objective obligation, as in Bob’s and Helmut’s case, because in fact objectively Carl should feed chicken to his daughter, as in fact (I assume for the sake of the argument) chickens have no rights. So just as we needed to suppose an objective obligation that Bob and Helmut got wrong, we need a hybrid objective-subjective one that Carl got wrong.

Here’s another way to see the problem. Bob thinks he is objectively obligated to give no money to Alice and Helmut thinks he is objectively obligated to kill enemy soldiers. But when Carl applies (5), what does he come to think? He doesn’t come to think that he is objectively required to feed chicken to his daughter. He already thought that this was 75% likely, and (5) does not affect that judgment at all. It seems that just as Bob and Helmut have a belief about something other than mere subjective obligation, Carl does as well, but in his case that’s not objective obligation. So it seems Carl has to be judging, and doing so incorrectly, about some sort of a hybrid obligation.

This makes me really, really want an account of obligation that doesn’t involve two different kinds. But I don’t know a really good one.