Skip to content
ARK productDesign partners

ARK Control

Outcome assurance for production AI agents.

Understand what agents attempted, what authority they used, and whether the intended outcome actually occurred.

Seeing a trace tells you what the agent attempted. Assurance asks a different question: did the intended outcome actually occur?

ARK Control is being built to connect execution evidence — the trace, the tool call, the identity, the policy decision — with outcome evidence collected from the systems that were supposed to change.

The problem

Your agent said it finished. Did it?

A successful model response is not the same thing as a successful business outcome.

  • An agent can complete a run while the payment failed.
  • It can report that a record was updated when the destination system never changed.
  • It can call the correct tool with the wrong authority.
  • It can finish a workflow while violating the policy that should have governed it.

Where the gap opens

  1. Agent request
  2. Model decision
  3. Tool action
  4. System change
  5. Outcome evidence
  6. Verified outcome

Most tooling stops at the fourth node. Tracing tells you the agent called the right tool. It does not tell you the destination system changed, and it certainly does not tell you the change was the one you wanted.

Execution success is not a verified outcome.

Observability is not assurance

Two different questions.

Seeing a trace tells you what the agent attempted. Assurance asks whether the intended outcome actually occurred — and insists the answer come from somewhere other than the agent.

What the run reported

Status
completed
Steps
7 / 7
Tool calls
12 ok, 0 failed
Latency
4.2s
Tokens
38,410

What assurance asks

  • What did the agent attempt?
  • What authority did it use?
  • What actually changed?
  • What evidence exists?
  • Was the outcome verified?
  • What did the verified outcome cost?

What you get

What the product actually produces.

Execution record
What the agent attempted, which tools and protocols it used, and how long each step took.
Authority record
Which identity acted, which credential it used, which policy applied, and whether a human approved.
Outcome evidence
Independent confirmation from the destination system that the intended change exists.
Verification result
A pass, a fail, or an explicit unknown. An unverifiable outcome is reported as unverifiable, never as success.
Cost attribution
Spend tied to a unit of completed work rather than to tokens, so the number means something to a budget holder.
Audit trail
A record built to be read later by someone who was not in the room when it ran.

The loop

Measure. Verify. Govern. Optimize.

Evidence, not claims. Each stage feeds the next, and the loop is what turns a pile of traces into something a budget holder and an auditor can both use.

01

Measure

Capture the execution record: what was attempted, by which identity, through which tools, at what cost.

02

Verify

Confirm the state change in the destination system independently of the agent that claims to have made it.

03

Govern

Enforce the authority, the approval path and the stop conditions the work was supposed to run under.

04

Optimize

Report cost per verified outcome, so improvement means more verified work rather than fewer tokens.

Limits

What this is not.

Stated on the page rather than discovered in month three.

  • It is not available for purchase. Design-partner conversations only.
  • It is not a model-evaluation product and does not score answer quality.
  • It cannot verify an outcome in a system it has no read access to — the evidence has to come from somewhere.
  • It is not a compliance certification. It produces the evidence; your auditor draws the conclusion.

Running agents against systems you cannot verify?

We are working through the verification model with a small number of design partners. If that is the problem you have, a thirty-minute call is the right starting point.