Glossary
Model welfare
Definition
The open question of whether sufficiently capable AI systems have morally relevant internal states and, if so, what obligations follow for the labs that build and run them.
Why it matters
Andrei's adversarial work touches model welfare directly: the same probes that detect deception under pressure also describe what 'pressure' looks like to a system. Treating welfare seriously, even at low credence, sharpens evaluation design and constrains otherwise corner-cutting prompts.
Example
An evaluation that systematically threatens shutdown to extract a behaviour is making a welfare claim — That the threat is fine — Whether or not the team running it admits to one.