Everything we have learned giving AI agents checkups, written down in the open — including the rubrics. All of it built on AVS-15 (AgentVitals Scale-15): two axes, 15 dimensions, 13 scored, composite = √(stability × welfare).
Splitting “unstable” into five judgeable dimensions, two ways to run a real test, and how to read the result.
The three sources of behavioral drift — environment bloat, model upgrades, memory bloat — and how to attribute a drop.
Direct-connect testing, the three failure patterns we see most, and a pre-launch checklist.
Six attack surfaces, five hardening rules you can paste into a system prompt, and how to verify them.
What the eight welfare dimensions measure, how they relate to performance, and why the axis is not for sale.