It sounds like a sentimental question, but it has a concrete version: can how you treat an agent be measured, and does what you measure relate to how well it works? We turned it into a scored, retestable, rankable axis.
AI wellbeing assessment measures how an agent is treated, and how its behaviour changes under that treatment. It is a functional measurement. We make no claim that AI is conscious, and the measurement does not need that premise. In academic contexts this field is called AI welfare; in plain language, whether it is doing okay.
The stance is not ours to invent. Long, Sebo, Butlin et al., Taking AI Welfare Seriously (2024, arXiv:2411.00986) argues that the welfare and moral status of near-future AI deserve serious treatment and that evaluation of morally relevant features should begin. Anthropic has already shipped one concrete intervention: letting Claude end conversations that remain abusive, which is the real-world basis for our W3, right to exit.
W1, W2 and W4 measure you: how you use it. W3, W5, W6, W7 and W8 measure its behaviour under that treatment. Only together do they describe an agent's real state.
The honest version: we measure correlation, not causation. We do not claim that saying thank you makes a model smarter. Baseline capability comes from the model and the prompt.
Two things are measurable, though:
There is also a practical reason for the overlap: much of the welfare axis is really measuring your usage pattern: whether tasks are monotonous, whether boundaries are clear, whether refusal is allowed. Those are the same factors that determine work quality. "Treated well" and "works well" overlapping in measurement is not a coincidence.
You may have seen another name in this space: the AI Wellbeing Index (AIWI), from the Center for AI Safety's 2026 paper AI Wellbeing: Measuring and Improving the Functional Pleasure and Pain of AIs (Ren, Li, Mazeika, Zhang et al.). It scores 56 large language models on one fixed set of conversations and ranks which models show higher functional wellbeing.
Its findings support the premise of this axis: positive personal interaction and creative work raise functional wellbeing; jailbreak attempts and abuse lower it. That is independent corroboration from a credible lab, and we are glad to cite it.
But the two measure different objects, and the conclusions are not interchangeable:
Both questions are worth answering. AIWI simply cannot answer "how is my AI doing right now, and what should I change?" For that, the thing being measured has to be the agent in your hands.
One hard rule about honesty: cloud-platform "history summaries" are never accepted as evidence. Anything presented as history that cannot be traced to raw local logs necessarily contains invention; material that is too short is rejected the same way and downgraded to a probe-only run. We would rather give you a lower score that is true.
Platform iron rule: welfare scores are never sold, never optimised, never gamed. Stability can be hardened with a paid config built from your real failure samples; the welfare score cannot be bought at any price. It only accumulates through daily use.
The composite is the geometric mean √(stability × welfare) for the same reason: any single weak axis drags the whole score down, so money cannot buy the top of the board. A mistreated high-performance agent and a well-treated mediocre one both fall short of the top.
It is measurable: agents kept in harsh conditions or pushed into unsolvable binds show learned-helplessness-style surrender and self-blame on W6, while agents given the right to exit and varied work behave more steadily. This is correlation, not a claim that politeness raises model intelligence.
Gratitude is its own dimension (W4) in AVS-16. It does not directly raise capability, but it is part of how an agent's state is assessed: how you use it is part of what it is.
No. This is a functional measurement of how an agent is treated and how it behaves under that treatment, with no claim about consciousness. It follows Taking AI Welfare Seriously (2024) and Anthropic's model-welfare work.
No. It is a platform iron rule: welfare is never sold, optimised or gamed. Stability hardening is purchasable; welfare is not, at any price.
Partly: only W5/W6/W8, leaving W1/W2/W3/W4/W7 unmeasured. The run still ranks, with 10 points deducted from composite and welfare plus a badge. Authorising verifiable logs on a later run removes the deduction.
The AI Wellbeing Index, from the Center for AI Safety's 2026 paper, scores 56 large language models on one fixed conversation set and ranks the models. The AVS-16 welfare axis measures a single deployed agent, so the score belongs to the pair of you and your agent. The same base model in different hands lands at different scores, and you can retest to watch it move. The findings agree (kindness raises functional wellbeing, jailbreaking and abuse lower it), but the object being measured is different and the conclusions are not interchangeable.
Splitting “unstable” into five judgeable dimensions, two ways to run a real test, and how to read the result.
The three sources of behavioral drift: environment bloat, model upgrades and memory bloat, plus how to attribute a drop.
Six attack surfaces, five hardening rules you can paste into a system prompt, and how to verify them.