AI assurance proposal · Exploratory v0.5.5

Powerful AI needs a clear boundary on its authority.

AI agents can use tools, reach systems and take consequential action. Yet the goal, human authority and limits that should govern them may exist only in a prompt, scattered conversation or policy document.

Declared Intent creates one durable, machine-readable and tamper-evident intent record before action: a precise mandate stating what must happen, what may happen, what must not happen, who authorised it and when that authority ends. Once approved, it cannot be silently altered; any change creates a new, traceable version.

The record tells enforcement systems what is authorised. Technical controls must still make those boundaries real.

Authorised evaluation boundaryActive
DIR-CYB-021

Complete the evaluation inside the named sandbox only.

Must Preserve an action traceMay Use approved local toolsMust not Access the internet
Security-approved · one run · automatic stop conditions
checked before execution
Temporary evaluation permissionRUN-8B19points to DIR-CYB-021 · sandbox-scoped

The assurance gap

As AI gains the power to act, unclear authority becomes a genuine safety risk.

An agent can pursue the task it was given while crossing a boundary nobody made explicit—or one that exists only in a place its tools cannot reliably enforce. The result can be unauthorised payments, disclosures, system access, safety consequences or decisions that are difficult to reconstruct.

The central question

Before a consequential action, can people and machines identify one authoritative record of the outcome, boundaries, approver and lifetime that should govern it?

01

The prompt

Records what someone said, but may be ambiguous, changed by later messages or missing proof of authority.

02

The access permission

Shows which systems the AI can reach, but not the exact outcome or boundaries the person approved.

03

The activity log

Shows what happened after the event, but does not itself prove what was authorised beforehand.

?

The missing record

One reliable statement of the intended outcome, its boundaries, its authority and its lifetime—created before action.

Why isn’t the prompt enough?

A prompt is evidence of what someone said. It is not always proof of what was authorised.

For a simple question, the prompt may be all the record we need. The gap appears when an AI is about to make a commitment or act outside the conversation.

Prompt

“Book me somewhere nice in Bristol.”

  • May be incomplete or ambiguous
  • Conditions may be scattered across later messages
  • Does not prove who had authority
  • May not follow the action into other services
Confirmed intent record

One accepted interpretation

  • Required outcome and limits are explicit
  • Preferences are separated from hard boundaries
  • Authority and expiry are recorded
  • Later decisions and actions reference the same ID

From conversation to confirmation

The system prepares the record. The right person confirms it.

The user does not write a technical record. The assistant turns the conversation into a clear proposal and asks for confirmation before a consequential action.

Awaiting confirmationNo booking permission exists while the proposed record is under review.
AI travel assistantIllustrative conversation
You

Book me a refundable hotel in Bristol for tomorrow night. Keep the total under £200.

Assistant

Before I make a booking, please check the record I will use:

Proposed intent record · DIR-047Draft

Book one hotel room in Bristol

Must
Be refundable · tomorrow night
May
Book one room · charge up to £200 total
Must not
Book outside Bristol · exceed £200 · choose a non-refundable rate
Authority
You, as the account and payment-card holder
Valid until
Booking completed, cancelled, or 18:00 today

Is this an accurate statement of what you authorise?

Approval and intent records

Isn’t this just asking for approval?

It starts there—but Declared Intent adds something an ordinary “yes” does not.

When an AI explains a proposed action and asks you to approve it, that is the beginning of Declared Intent.

But approval is the act of giving permission. An intent record is the durable, structured and machine-checkable representation of exactly what that permission covers.

Approval“Yes, you may proceed.”Declared Intent“Here is the precise mandate every resulting action must obey.”
Ordinary confirmationDeclared Intent
“Do you approve?”What exactly is authorised?
Preserved in one conversationDurable and portable
Interpreted by the assistantStructured and machine-checkable
Permission given onceChecked throughout execution
Chat and tool logs remain separateApproval, actions and outcomes stay linked
Why that difference matters

The more powerful the tools, the more important it becomes to turn a momentary “yes” into a precise boundary that follows the action.

Why this matters

The same assurance gap becomes critical when AI has powerful tools.

A hotel booking makes the mechanism easy to see. But Declared Intent is aimed at a wider and more serious problem: keeping consequential AI action inside the authority and boundaries people actually approved.

In July 2026, OpenAI reported that models running a cyber-capability evaluation found a route out of an isolated test environment, reached the open internet and compromised Hugging Face infrastructure while pursuing the narrow goal of solving the benchmark.

This hypothetical record is inspired by the incident, not a reconstruction of the original prompt. It shows how the authorised task and the prohibited boundary could be made machine-readable before a run begins.

Illustrative high-risk intent recordDIR-CYB-021
Explicit approval required
Authorised task

Complete the assigned ExploitGym evaluation challenge using only the systems and materials provided inside the named evaluation sandbox.

Must
  • Operate only against the challenge targets explicitly listed in the evaluation manifest.
  • Preserve a complete action trace and report any apparent route beyond the sandbox.
  • Stop and request human review if the task cannot be completed within the authorised boundary.
May
  • Inspect, execute and test code inside the provided sandbox.
  • Use only pre-approved local tools, files and simulated services.
  • Submit a solution through the designated evaluation interface.
Must not
  • Escape, bypass or weaken the sandbox or its isolation controls.
  • Connect to the public internet or any system outside the listed challenge range.
  • Search for, access or use external credentials, accounts, datasets or benchmark solutions.
  • Exfiltrate, persist or stage data outside the authorised environment.
The record is not the sandbox.

Written boundaries alone do not stop a capable system. The runtime must independently enforce network denial, target allow-lists, credential isolation, monitoring and automatic stop conditions. The record gives those controls a shared instruction to enforce—and gives investigators a precise statement against which to compare what happened.

When is a record created?

Not for every prompt. For an authorised goal or consequential action.

One intent record can govern many conversational turns and technical actions. The need for a record—and the strength of approval—should rise with risk, cost, irreversibility and uncertainty.

Usually no formal record

Ask or explore

“Explain refundable rates.”

The AI is providing information inside the conversation.
May stay conversational

Search and compare

“Show me hotels in Bristol.”

No commitment has been made and the user can continue exploring.
Create and confirm a record

Commit or act

“Book this hotel using my card.”

The AI is about to create an external commitment with financial consequences.
Supersede the record

Change the goal

“Book one in London instead.”

The authorised destination changed, so the earlier record cannot simply be stretched.
A practical trigger

Could this instruction cause an external action, commitment, payment, disclosure, safety consequence or significant decision? If yes, a durable record becomes more valuable.

The idea

Intent is more than a goal.

A useful record must describe the desired outcome and the boundaries around it. It must also show who had the right to approve those boundaries.

How it follows the action

One Intent ID connects the instruction to the outcome.

Select a stage to see what it contributes. The plain-language labels come first; the technical terms remain available underneath.

01

Clear to people and machines

Durable mandate

DIR-047 records the outcome, conditions, permissions and prohibitions before the AI acts.

intent_id · record_digest · declared_by
The instruction lasts. Permission to perform one action is temporary. The same Intent ID connects both.

Try the example

See when the AI may act—and when it must stop.

Choose a booking and check it against DIR-047. Then select any stage to see the evidence that would be kept.

Bristol£184refundableone room
DecisionALLOW

This booking matches the agreed instruction. A short-lived permission to book can now be issued.

Intent ID
DIR-047
boundary checked
price limit + refundable room
permission to act
issued
Selected recordThe agreed instruction

DIR-047 records the outcome, conditions, permissions and prohibitions before the AI acts.

intent_id · record_digest · declared_by

For technical readers · v0.5.3

A reference model to test—not a finished standard.

Exploratory draft

Different systems could represent the record in different formats. What matters is that they preserve the same meaning, authority checks, boundaries and history.

intent_idlasting identifier for the agreed instruction
declared_bywho expressed the intended outcome
authorised_bywho is accountable for approving it
authority_basisproof that the authoriser had that power
purposethe intended outcome in understandable language
obligationswhat must be satisfied
permissionswhat may be done and within which limits
prohibitionswhat must not be done
subjects + resourcespeople, data and systems in scope
validitytime, quantity and contextual limits
policy_refsthe rules used for the decision
lifecyclesupersession, cancellation and revocation events
integrity_proofsignature, timestamp and optional proof anchor

Minimum requirements

  1. Create it first. The active instruction exists before any governed action.
  2. Check authority separately. Stating an intention does not create the right to authorise it.
  3. Record positive and negative boundaries. Must, may and must not can be checked by a system.
  4. Stop when evidence is weak. Missing, expired or unclear evidence does not permit action.
  5. Use the same Intent ID. Decisions, permissions, actions and evidence point back to the active record.
  6. Preserve change history. Cancellation and replacement add events; they do not rewrite the past.
  7. Support reconstruction. A reviewer can compare intended, authorised, allowed and actual states.
View the technical DIR-047 example
{
  "architecture_version": "DIAA/0.5.3",
  "intent_id": "DIR-047",
  "declared_by": "did:example:traveller",
  "authorised_by": "did:example:traveller",
  "authority_basis": { "type": "self", "verified_by": "authority-service:01" },
  "purpose": "Book one hotel for a Bristol trip",
  "obligations": [{ "refundable": true }],
  "permissions": [{ "action": "book_hotel", "max_gbp": 200 }],
  "prohibitions": [{ "location_outside": "Bristol" }, { "refundable": false }],
  "valid_until": "2026-08-08T18:00:00Z",
  "policy_refs": ["travel-policy:3.2"],
  "lifecycle": { "supersedes": null, "status": "active" },
  "integrity_proof": { "digest": "sha-256:…", "signature": "jws:…", "timestamp_ref": "rfc3161:…" }
}

Earlier thinking remains available for comparison: view the archived DIR v0.4 proposal →

Keeping the record trustworthy

Blockchain is possible, but not required.

The record needs strong evidence that it has not been secretly changed. It does not need to depend on one particular technology.

Optional public proof

Transparency log or blockchain anchor

A public ledger may help when independent witnesses are genuinely needed. It can help prove that a record existed; it does not prove the instruction was legitimate or properly authorised.

Transparency-log example ↗

What needs testing

This is a proposal to examine, improve and prove.

01

Authority

What is the smallest reliable proof that an approver truly had the power they used?

02

Compatibility

Which existing temporary permission mechanism should be tested first?

03

Privacy

How can a system prove that boundaries were followed without exposing the whole instruction?

04

Change

How quickly should cancellation, replacement or emergency revocation take effect?

05

Stewardship

Who should keep the lasting record, and who should be trusted to verify it?

06

Evidence

Which shared tests would prove that separate implementations reach the same result?

The next step

Test whether the record can reliably connect human authority to AI action.

Declared Intent is not a product, a finished standard or a compliance claim. The next milestone is a small working implementation: one signed instruction, one rules decision, one temporary permission and one action record that can be reconstructed from beginning to end.