Glossary

Model welfare

Definition

The open question of whether sufficiently capable AI systems have morally relevant internal states and, if so, what obligations follow for the labs that build and run them.

Why it matters

Andrei's adversarial work touches model welfare directly: the same probes that detect deception under pressure also describe what 'pressure' looks like to a system. Treating welfare seriously, even at low credence, sharpens evaluation design and constrains otherwise corner-cutting prompts.

Example

An evaluation that systematically threatens shutdown to extract a behaviour is making a welfare claim — That the threat is fine — Whether or not the team running it admits to one.

Read further