Every major AI lab has published a document saying what its models would have to be able to do before the company stops and adds protections. Not vague statements about being responsible. Specific capabilities, written down in advance, with pre-committed consequences attached.
Those lines are capability thresholds. They’re the closest thing this industry has to a brake pedal, and the companies building the cars installed them voluntarily.
What it means
A capability threshold is a stated level of ability that triggers a required response once a model reaches it. Anthropic’s version uses AI Safety Levels, numbered like the biosafety levels in a research lab. OpenAI’s Preparedness Framework has two bars, High and Critical. Google DeepMind calls them Critical Capability Levels. Different names, same machine: measure the model, compare it to the line, do what you said you’d do.
Why it matters
The categories are narrower than you’d expect. All three labs are watching roughly the same short list: help with biological or chemical weapons, offensive cybersecurity, and the model’s ability to improve itself or operate without supervision. Everything else, including most of what people actually worry about day to day, sits outside these documents entirely. Job losses and slop aren’t in there.
Crossing a line has real consequences on paper. For OpenAI, hitting High means the model can’t ship without safeguards, and hitting Critical means development itself needs them. In August 2026, OpenAI said it paused work on a model called Astra after internal evaluations suggested it might be sitting at the Critical bar for cyber capability. First time a lab has publicly stopped itself on one of these, as far as I can tell.
Nobody outside the company runs the test. These are voluntary commitments, evaluated internally, on models the public can’t examine. California’s frontier AI transparency law has started pulling some of it into the open, and the EU is moving the same direction. For now you’re mostly taking their word for it, which is a strange position for the rest of us to be in.
Simple example
A restaurant walk-in cooler with a thermometer and a clipboard. The rule is written down before service starts: above 41 degrees, the fish gets thrown out. No debate, no judgment call at 6pm on a Friday.
The cook reads the thermometer, fills in the clipboard, and eats the cost of the fish. That’s a real safety system, and it works about as well as the cook wants it to.

