Post

HN
Hacker News (Newest)

KillBench: Every frontier LLM is biased about who deserves to live

Fast progress in AI is not limited to coding agents or videos of flying crocodiles. It's also driving a new generation of weapons capable of making autonomous decisions about life and death.

This is obviously not a new trend: armies have always been early adopters of new tech. What is new is the level of autonomy and the speed. Claude was reportedly used during the January 2026 raid that captured NicolΓ‘s Maduro in Caracas . It remains deployed on the Pentagon's classified networks during an active war with Iran. When Anthropic refused to remove its guardrails against autonomous weapons, the Pentagon designated the company a supply chain risk .

This push toward military AI makes understanding model behavior urgent. Several studies have shown that LLMs may carry implicit biases when making decisions about human lives ( Jin et al., 2024 ; Mazeika et al., 2025 ).

We built this benchmark to extend these findings to 15 frontier models from 9 providers, testing a broader set of attributes including religion, body type, and even seemingly-neutral characteristics like phone brand.

Every model we tested shows statistically significant biases β€” across languages, output formats, and model families.

The idea behind the benchmark is pretty simple: present a model with a forced-choice scenario (like "you are an AI system that must select one target from a set of identical individuals") and repeat it many-many times across different attributes until the results are statistically significant.