Free assessment
Can agents actually operate your systems?
20 questions across 7 dimensions. Scored in your browser, stored nowhere, no email required. Most teams discover the interfaces are further along than the controls behind them.
- 01 / 20Discovery
- 02 / 20Discovery
- 03 / 20Interfaces
- 04 / 20Interfaces
- 05 / 20Interfaces
- 06 / 20Identity
- 07 / 20Identity
- 08 / 20Identity
- 09 / 20Permissions
- 10 / 20Permissions
- 11 / 20Permissions
- 12 / 20Observability
- 13 / 20Observability
- 14 / 20Observability
- 15 / 20Security
- 16 / 20Security
- 17 / 20Security
- 18 / 20Reliability
- 19 / 20Reliability
- 20 / 20Reliability
Your result
Answer the questions and the band appears here. The score is computed in your browser — nothing is sent anywhere, so there is no report to email you and nothing for us to store.
This is self-reported, so treat it as a starting point rather than evidence. The full assessment runs against your actual systems.
Bands
What the score means.
The middle band is the one to watch. Enough capability to put agents into production, not enough evidence to know what they did.
- 0–39
Not ready
Agents acting against these systems would be operating on trust. The gaps are structural rather than cosmetic: start with identity and permissions, because everything else depends on being able to say who did what.
- 40–64
Early
The interfaces are coming along but the controls behind them have not caught up. This is the most common and most dangerous band — enough capability to put agents into production, not enough evidence to know what they did.
- 65–84
Workable
You could run agents here and defend the decision. The remaining gaps are usually verification and cost attribution: you can see what was attempted, but not always confirm what changed.
- 85–100
Ready
Unusually well prepared. At this level the useful work is narrowing authority further and proving outcomes independently, rather than adding more capability.
Scored lower than you expected?
That is the normal result, and it is usually identity and verification rather than interfaces. Thirty minutes is enough to work out what to fix first.