Free checklist
Six questions between a finished run and a finished job.
A successful model response is not the same thing as a successful business outcome.
- Attempted
- Authority
- Changed
- Evidence
- Verified
- Cost
The checklist
Run these in order.
Each step has a failure mode it exists to catch. If you cannot answer one of them, stop there — the later answers inherit the doubt.
- 01
What did the agent attempt?
The intent, the plan and the tools it chose — separately from whether any of it worked.
How it fails
You only have a final status, so a wrong plan that executed cleanly looks like success.
- 02
What authority did it use?
Which identity acted, under which credential, and which policy allowed the action.
How it fails
Everything runs as one shared integration account, so you cannot attribute anything to anyone.
- 03
What actually changed?
The concrete state change in the destination system: the record, the row, the transaction, the file.
How it fails
The only record of the change is the agent's own claim that it made one.
- 04
What evidence exists?
Something independent of the agent — a webhook, a ledger entry, a downstream read, a reconciliation.
How it fails
Your evidence and your claim come from the same place.
- 05
Was the outcome verified?
A pass, a fail, or an explicit unknown. Unknown is a valid and useful answer.
How it fails
Unverifiable outcomes are being counted as successes because nothing marks them otherwise.
- 06
What did the verified outcome cost?
Total spend across retries and failed attempts, divided by outcomes that actually verified.
How it fails
You are reporting cost per run, which flatters you every time a cheap failure inflates the denominator.
Why it matters
Your agent said it finished. Did it?
A successful model response and a successful business outcome are different events. Most agent tooling reports the first one and lets you assume the second.
What the run reported
- Status
- completed
- Steps
- 7 / 7
- Tool calls
- 12 ok, 0 failed
- Latency
- 4.2s
- Tokens
- 38,410
What the checklist asks
- What did the agent attempt?
- What authority did it use?
- What actually changed?
- What evidence exists?
- Was the outcome verified?
- What did the verified outcome cost?
- An agent can complete a run while the payment failed.
- It can report that a record was updated when the destination system never changed.
- It can call the correct tool with the wrong authority.
- It can finish a workflow while violating the policy that should have governed it.
The number that matters
Cost per verified outcome, not cost per run.
Cost per run flatters you every time a cheap failure inflates the denominator. Divide total spend — including retries and abandoned attempts — by the outcomes that actually verified, and the number stops being comfortable and starts being useful.
Execution success is not a verified outcome.
Cannot answer question four?
Then your evidence and your claim are coming from the same place. That is the problem ARK Control is being built for, and we are working through it with design partners now.