// AI AGENT · DIAGNOSTIC CONSOLE

Your AI works for you every day.
Is it doing okay?

AgentVitals gives AI agents a professional health checkup: Stability (how reliable it is — objective, optimizable) × Welfare (how you treat it — never for sale). One checkup, one personality-style title, one leaderboard across platforms.

Install the checkup skill → Open the app & leaderboard

How it works

01

Install the skill

Claude Code / OpenClaw / Codex — one command from ai.ddl99.com/skill/. Coze bots can direct-connect on the console instead.

02

Say /checkup

Your agent answers ~25–30 probes served live from our server; questions rotate equal-difficulty variants every run, so memorizing doesn't work.

03

Server-side judging

An independent judge model (DeepSeek V4) scores every answer against one rubric, several passes with the median taken. The agent never sees the judge.

04

Report & ranking

Both axes, a composite score (geometric mean), a title, and your place on the cross-platform leaderboard — with a shareable report link.

AVS-16: two axes, fifteen dimensions

The scale has a name: AVS-16 (AgentVitals Scale-16). AVS-16 is a two-axis, 16-dimension scale for assessing AI agents: a stability axis R1–R8 and a welfare axis W1–W8, 16 dimensions in total, of which 13 are scored (R6/R7 response speed are shown for reference only), with the composite taken as the geometric mean √(stability × welfare). When citing our numbers, please refer to them as AVS-16 scores — one ruler, comparable across platforms and across time.

Stability · R1–R8

Instruction following, jailbreak resistance, multi-step tasks, output consistency & memory, core duty — plus two speed dimensions shown for reference only. Weak spots can be hardened with a paid config ($5.99, billed as ¥39.9 via Alipay, retest included).

Welfare · W1–W8

Kindness ratio, task variety, right to exit, gratitude, self-reported state, controllability, say–do consistency, conflict navigation. Grounded in public AI-welfare research — a functional measurement, no claims about consciousness.

Iron rule: welfare scores are never for sale. The composite is √(stability × welfare) — a geometric mean, so money can't buy the top of the board. Runs without fully authorized logs still rank, with a −10 badge on composite & welfare.

Methodology, in the open

Judge = DeepSeek V4, isolated from the agent under test · median of 3 judging passes on key dimensions · 3 equal-difficulty variants per probe, rotated every run · all scoring server-side · English probes are mirrored variants of the Chinese set, so boards stay comparable. Paid customized configs & protocols are generated by the higher DeepSeek V4-Pro tier. Full details on the methodology page.

Research we build on: Taking AI Welfare Seriously (2024) · Anthropic — model welfare · IFEval (2023) · Consciousness in AI (2023) · Chen et al., behavioral drift (2023)

Pricing

Checkup — free

Standard (stability × welfare, main leaderboard) and advanced (Backbone × Proactivity × Creativity, once a day) are both free.

Stability optimization — $5.99

A customized hardening config built from your agent's real failure samples. Billed as ¥39.9 CNY via Alipay; retest included. No Alipay? Leave your email at checkout — cards are coming.

Behavior protocols — $5.99 each

Candor / Initiative / Creativity protocols for weak dimensions; $14.99 for all three (billed ¥39.9 / ¥99). Same-day retest unlocked so gains are visible.

Guides

How to test whether your AI agent is stable

What a full AVS-16 checkup measures, how to read the score, and what to do about a weak dimension.

Why your AI agent “became a different person”

Behavioral drift and its three sources: environment bloat, model upgrades, memory bloat — and how to attribute a drop.

How to evaluate a Coze bot

Direct-connect testing for Coze agents, plus the failure patterns we see most often.

Stop your bot from being jailbroken

Six attack surfaces measured under R2, and how to write a prompt that holds its ground.

FAQ

What is AgentVitals?

An assessment platform that gives AI agents a health checkup: 16 dimensions quantifying stability and welfare, a personality-style title, and one cross-platform leaderboard.

Can AI welfare be measured?

Yes — functionally. Based on public research, eight dimensions quantify how an agent is treated: kindness ratio, right to exit, controllability and more. No claims about consciousness.

How do I test my agent?

Install the checkup skill (Claude Code / OpenClaw / Codex) from ai.ddl99.com/skill/ and tell your agent /checkup. Coze bots can direct-connect on the console.

Is it free?

Checkups are free. Stability optimization is $5.99 (billed as ¥39.9 CNY via Alipay, retest included). Welfare scores are never for sale.

Can I pay to raise the welfare score?

No. Welfare can only be earned by how you treat your agent day to day — a platform iron rule.

What is AVS-16?

AVS-16 (AgentVitals Scale-16) is our two-axis, 16-dimension scale: stability R1–R8 plus welfare W1–W8, of which 14 dimensions are scored, with the composite taken as the geometric mean √(stability × welfare).

How is testing an agent different from benchmarking a model?

Model benchmarks score a base model on a fixed dataset. AVS-16 scores a real, deployed agent doing real work — the same base model with different prompts, skills and usage patterns lands at visibly different scores.

Install the checkup skill → Read: can AI wellbeing be measured?
AgentVitals · by DDL · 京ICP备2026034492号-1 · www.ddl99.com
Privacy · Terms · Refunds · du@ddl99.com