Rebecca Crootof has a current paper on what she calls the “kill box” problem for autonomous weapon systems.  It shows just how difficult it is to translate traditional just war values and laws into an AI context.  Basically, one of the most important rules for warfare is that you try to distinguish enemy combatants from non-combatants, and only target the former.  Of course any skim of the current news will show that this rule is frequently broken, but it seems reasonable to suppose that a world where militaries felt obliged to pay at least some attention to this rule is better than one where they do not.

Enter autonomous weapons systems (AWS, = where drones are rapidly going). Drones have typically been piloted by a human, who sees the target area from a camera on the drone and then engages it to destroy a particular target.  The military problem with this process is that it requires that the weapon be in more-or-less continuous communication with its operator, and it is relatively easy to disrupt communication, leaving the weapon useless.  In Ukraine, Russia has apparently taken to attaching drones to narrow fiber optic cables, which solves the radio-jamming problem but produces new problems of range, since the drone can now only go as far as its spool of cable lasts.  A couple of weeks ago, Russia seems to have taken the step that military ethicists have long-feared: a completely AI-powered drone.  Apparently aiming at a gas station, it crashed into a wall in Zaporizhzhia and killed three civilians.

AI powered drones are immune to both communications-jamming and can fly long distances without needing a fiber optic leash.  They can be assigned a target based on picture or description, and then sent on their way to destroy it. They thus present a series of problems around whether or not they can distinguish combatants from non-combatants, since this task is hard even for humans. As Crootof puts it, “to comply with the distinction requirement, an anti-personnel autonomous weapon system would need to be able to identify a complex range of unlawful human targets, including civilians, wounded and surrendering combatants, and other protected persons like medical and religious personnel” (5).

Crootof thinks that the difficulty of this task raises a different risk:

“Are technical standards relevant if autonomous weapon systems are deployed in environments where it would be reasonable for a commander to assume that there are no unlawful targets? Consider the SGR-A1, a stationary, armed robot used to patrol the demilitarized zone between North and South Korea. These robots identify every human being who enters the demilitarized zone as an enemy, on the grounds that the individual has entered a prohibited zone. Arguably, if every potential target in a defined zone is a lawful military objective, there’s no need for a distinction analysis. Taken to its logical extreme, this suggests that a berserker killer robot could be used in compliance with the distinction requirement—provided it is used only within a delineated space where there are no unlawful targets. If you’re concerned your system might make mistakes, only use it in mistake-free zones” (5)

That is, rather than fail to meet a distinction requirement, users of AWS will be tempted to sidestep the requirement altogether.  As she notes, there are precedents for such “kill boxes,” and the fear is that technical difficulties in designing systems that are capable of discriminating between combatants and non-combatants will drive states to simply designate entire areas as containing only combatants, no matter the reality on the ground.

Although Crootof doesn’t refer to this aspect of the problem, it seems to me that the temptation to establish kill boxes is going to there almost no matter how good the AWS are.  As Mark Swiatek pointed out as far back as 2012, an automated targeting system has to function like any other identification algorithm, which means that it will have an error rate and have to be calibrated for sensitivity and specificity.  As he puts the consequence:

“Programmed error rates translate into the killing of non-target individuals.  For example, if we were to employ an LAS with a tolerance for error of 10%, and 10 out of every 100 hundred individuals that the system kills shouldn’t have been targeted, then the system would be working properly, as programmed” (247).

In other words, the decision to use an AWS is a deliberate decision to kill a certain percentage of non-combatants.  As Swiatek notes, civilian casualties can be viewed as accidental in human systems and thus excusable by way of the doctrine of double effect, but using an autonomous system that identifies its targets requires a deliberate intention to kill a certain percentage of innocents.  In other words, the move to statistical determination takes a lot of the sense of chance out of targeting decisions.  In Biddle and Kukla’s terms, AWS targeting systems come with a very high degree of unavoidable epistemic risk.  Biddle and Kukla’s examples are largely medical, but their characterization seems to apply:

“Decisions such as how to operationalize concepts, what statistical models to use, and how to set an evidence bar, among others, have to be made, and any way of making them raises some epistemic risks and lowers others; one cannot make these decisions in abstraction from values and interests. These are paradigmatic phronetic risks” (221).

As they emphasize, these decisions emerge at an institutional, structural level, and they have to be actively managed.  Koray Karaca extends the point to binary classification models, and notes that terms like “high risk” have to be operationalized, and that these definitions are normative, not purely technical. The novel twist with AWS is that the act of managing epistemic risk may very well constitute a war crime just as much as not managing it.

Things are further complicated by base rate considerations.  If the prevalence of the target condition is low in the selected population, then even a system that looks accurate (has a high accuracy or AUC) will turn out to make a lot of mistakes.  Consider a prediction algorithm for a rare disease (1 in 1000 prevalence). A disease prediction algorithm that is 99% sensitive (so it picks up 99% of true positives) and only has a 1% false positive rate will nonetheless deliver 10 false positives for every true positive because of the low prevalence of the condition in the population.  So (ceteris paribus!) the chance that someone the test flags as having the disease only has a 9% chance of actually having it.

Given conditions of modern warfare and urban areas, your guess is as good as mine about the prevalence of combatants and non-combatants in a given space.  But assume for the sake of an example that 10% of the population is an enemy combatant and the presence of a 95% sensitive AWS with a 5% false positive rate. That seems pretty good, given the complexity of its task.  That means in a population of 1000 people, 100 of them are combatants.  The system will detect 95 of them.  Of the remaining 900 people, the system will flag 5% of them incorrectly.  So it will generate 45 false positives.  In short, to kill 95 enemy combatants, you are committing to also kill 45 innocent non-combatants.  That’s a lot of collateral damage to intentionally inflict (and it’s before you look at how many people shrapnel, falling debris, etc. are going to kill)!  In other words, the real-world performance of an AWS on anything other than a traditional battlefield will be a lot worse than its advertised accuracy.

Hence the temptation of kill boxes: the only way to get the false positive rate to zero is to have an area where the prevalence of combatants is 100%.

To return to Crootof, what is urgently needed, then, is the evolution of standards governing the creation and management of such areas.  Right now, “there is little extant clear law” (17) on the subject, and the states that do have policies have developed them for tactical reasons, such as coordinating attacks.  The need is urgent: “A kill box reverses the default presumption that an individual is a civilian until proven otherwise. To the extent that kill boxes are legitimized and proliferate, they threaten the distinction requirement and thus the entire edifice of international humanitarian law” (22).

The benefit is that “developing clear rules and standards for kill boxes will facilitate appropriate use, both because parties to a conflict will understand what is required for ex ante compliance and because it will be easier for outside entities to identify violations” (18).  Simultaneously, it will “formally reaffirm[]the import and applicability of distinction, clarify[] what compliance entails in the kill box context, and establish[]enforcement mechanisms for noncompliance” (19).  Such clarification would have a side benefit as well, enabling allies to coordinate around an area and avoid shooting at each other.

As Crootof notes, there are there are also numerous risks.  One obvious one is that humans-in-the-loop systems often fail spectacularly, mainly serving to provide a fall for when a system (inevitably) fails – a “moral crumple zone.” There are a variety of reasons why human supervision of AI is prone to failure (Crootof’s own coauthored work on this is very good), but an AI system that could setup kill boxes might very well setup human supervisors to be both functionally uninvolved in the process and then responsible for when the AI draws it around a school or an apartment building, rather than having the human be involved enough that they could actually take meaningful responsibility.

Another is bias or overfitting based on bad training data.  For example, “a training data set of ‘combatants’ that included only images of young men with beards, for example, would be simultaneously over- and under-inclusive, as young male civilians and older unbearded combatants would be more frequently misclassified.”  If AI surveillance determined that all those present on a city block were combatants, that would protect against bad training data, or at least the visibility of the problems it causes.  Crootof also notes that AI systems tend to be escalatory. 

A final risk is that kill-boxes will then legitimate the idea that civilians are to blame for their own deaths:

“They may encourage a more general acceptance of the proportion, amounts, or kinds of harms civilians are expected to absorb, or they may foster a more general devaluing of civilian life. If they are allowed to be established on a more permanent basis … kill boxes may lend support to narratives that blame the victims themselves for being in the zone, regardless of whether the kill box is invasive, unmarked, or illegitimate. In short, while the State establishing a kill box should bear the burden of ensuring that there are no civilians or other unlawful targets within it, there is a risk that the proliferation of kill boxes may promote an assumption that it is the civilian’s obligation to learn of and avoid kill boxes” (23).

I have no ability to do better than point to the depth of the problems around AWS targeting systems.  We are told repeatedly the high-tech weapons are precisely targeted, which is no doubt true.  But targeting systems have their limits, and those matter in the context of AWS. At least, how we talk about them is going to have to change.  On the other hand, perhaps there is less new here than one might think: this entire discussion reminds me of William Spanos’ Heideggerian critique of the US bombing campaigns in Vietnam. There, in an exercise of blatant technological Enframing, when confronted with an inability to find and correctly identify enemy combatants, the US took to napalming entire forests. Kill boxes make that decision explicit and up-front.

Posted in

Leave a comment