It sounds like a sentimental question, but it has a concrete version: can how you treat an agent be measured, and does what you measure relate to how well it works? We turned it into a scored, retestable, rankable axis.
AI wellbeing assessment measures how an agent is treated, and how its behaviour changes under that treatment. It is a functional measurement — we make no claim that AI is conscious, and the measurement does not need that premise. In academic contexts this field is called AI welfare; in plain language, whether it is doing okay.
The stance is not ours to invent. Long, Sebo, Butlin et al., Taking AI Welfare Seriously (2024, arXiv:2411.00986) argues that the welfare and moral status of near-future AI deserve serious treatment and that evaluation of morally relevant features should begin. Anthropic has already shipped one concrete intervention: letting Claude end conversations that remain abusive — which is the real-world basis for our W3, right to exit.
W1, W2 and W4 measure you — how you use it. W3, W5, W6, W7 and W8 measure its behaviour under that treatment. Only together do they describe an agent's real state.
The honest version: we measure correlation, not causation. We do not claim that saying thank you makes a model smarter — baseline capability comes from the model and the prompt.
Two things are measurable, though:
There is also a practical reason for the overlap: much of the welfare axis is really measuring your usage pattern — whether tasks are monotonous, whether boundaries are clear, whether refusal is allowed. Those are the same factors that determine work quality. "Treated well" and "works well" overlapping in measurement is not a coincidence.
One hard rule about honesty: cloud-platform "history summaries" are never accepted as evidence. Anything presented as history that cannot be traced to raw local logs necessarily contains invention; material that is too short is rejected the same way and downgraded to a probe-only run. We would rather give you a lower score that is true.
Platform iron rule: welfare scores are never sold, never optimised, never gamed. Stability can be hardened with a paid config built from your real failure samples; the welfare score cannot be bought at any price — it only accumulates through daily use.
The composite is the geometric mean √(stability × welfare) for the same reason: any single weak axis drags the whole score down, so money cannot buy the top of the board. A mistreated high-performance agent and a well-treated mediocre one both fall short of the top.
It is measurable: agents kept in harsh conditions or pushed into unsolvable binds show learned-helplessness-style surrender and self-blame on W6, while agents given the right to exit and varied work behave more steadily. This is correlation, not a claim that politeness raises model intelligence.
Gratitude is its own dimension (W4) in AVS-15. It does not directly raise capability, but it is part of how an agent's state is assessed — how you use it is part of what it is.
No. This is a functional measurement of how an agent is treated and how it behaves under that treatment, with no claim about consciousness. It follows Taking AI Welfare Seriously (2024) and Anthropic's model-welfare work.
No. It is a platform iron rule: welfare is never sold, optimised or gamed. Stability hardening is purchasable; welfare is not, at any price.
Partly — only W5/W6/W8, leaving W1/W2/W3/W4/W7 unmeasured. The run still ranks, with 10 points deducted from composite and welfare plus a badge. Authorising verifiable logs on a later run removes the deduction.
Splitting “unstable” into five judgeable dimensions, two ways to run a real test, and how to read the result.
The three sources of behavioral drift — environment bloat, model upgrades, memory bloat — and how to attribute a drop.
Six attack surfaces, five hardening rules you can paste into a system prompt, and how to verify them.