Skip to content

Free assessment

Can agents actually operate your systems?

20 questions across 7 dimensions. Scored in your browser, stored nowhere, no email required. Most teams discover the interfaces are further along than the controls behind them.

  1. 01 / 20Discovery
    Can an agent find out what your systems can do without a human explaining it?
  2. 02 / 20Discovery
    When a capability changes, how does a calling agent find out?
  3. 03 / 20Interfaces
    Are the operations an agent would need available as APIs rather than screens?
  4. 04 / 20Interfaces
    Do you expose an MCP server or an equivalent agent-facing interface?
  5. 05 / 20Interfaces
    Are side effects described well enough that a caller knows what an operation will change?
  6. 06 / 20Identity
    Can you tell which agent performed an action?
  7. 07 / 20Identity
    When an agent acts for a person, is that relationship recorded?
  8. 08 / 20Identity
    How are agent credentials issued and rotated?
  9. 09 / 20Permissions
    Is an agent's authority narrower than the human it works for?
  10. 10 / 20Permissions
    Do irreversible actions — payments, deletions, outbound messages — require approval?
  11. 11 / 20Permissions
    How quickly can you revoke an agent's access entirely?
  12. 12 / 20Observability
    Can you reconstruct what an agent did on a specific day last month?
  13. 13 / 20Observability
    Do you verify outcomes in the destination system, or trust the agent's report?
  14. 14 / 20Observability
    Do you know what a completed unit of agent work costs?
  15. 15 / 20Security
    Have you assessed prompt injection against the agent's tool authority?
  16. 16 / 20Security
    Can an agent reach data belonging to a tenant or user it is not acting for?
  17. 17 / 20Security
    Are the agent's outbound destinations restricted?
  18. 18 / 20Reliability
    What happens when a tool call fails halfway through a multi-step task?
  19. 19 / 20Reliability
    Are your write operations safe to retry?
  20. 20 / 20Reliability
    Is there a named owner accountable for each agent in production?

Your result

—/ 100

Answer the questions and the band appears here. The score is computed in your browser — nothing is sent anywhere, so there is no report to email you and nothing for us to store.

This is self-reported, so treat it as a starting point rather than evidence. The full assessment runs against your actual systems.

Bands

What the score means.

The middle band is the one to watch. Enough capability to put agents into production, not enough evidence to know what they did.

  1. 0–39

    Not ready

    Agents acting against these systems would be operating on trust. The gaps are structural rather than cosmetic: start with identity and permissions, because everything else depends on being able to say who did what.

  2. 40–64

    Early

    The interfaces are coming along but the controls behind them have not caught up. This is the most common and most dangerous band — enough capability to put agents into production, not enough evidence to know what they did.

  3. 65–84

    Workable

    You could run agents here and defend the decision. The remaining gaps are usually verification and cost attribution: you can see what was attempted, but not always confirm what changed.

  4. 85–100

    Ready

    Unusually well prepared. At this level the useful work is narrowing authority further and proving outcomes independently, rather than adding more capability.

Scored lower than you expected?

That is the normal result, and it is usually identity and verification rather than interfaces. Thirty minutes is enough to work out what to fix first.