Complete the evaluation inside the named sandbox only.
AI assurance proposal · Exploratory v0.5.5
Powerful AI needs a clear boundary on its authority.
AI agents can use tools, reach systems and take consequential action. Yet the goal, human authority and limits that should govern them may exist only in a prompt, scattered conversation or policy document.
Declared Intent creates one durable, machine-readable and tamper-evident intent record before action: a precise mandate stating what must happen, what may happen, what must not happen, who authorised it and when that authority ends. Once approved, it cannot be silently altered; any change creates a new, traceable version.
The record tells enforcement systems what is authorised. Technical controls must still make those boundaries real.
The assurance gap
As AI gains the power to act, unclear authority becomes a genuine safety risk.
An agent can pursue the task it was given while crossing a boundary nobody made explicit—or one that exists only in a place its tools cannot reliably enforce. The result can be unauthorised payments, disclosures, system access, safety consequences or decisions that are difficult to reconstruct.
Before a consequential action, can people and machines identify one authoritative record of the outcome, boundaries, approver and lifetime that should govern it?
The prompt
Records what someone said, but may be ambiguous, changed by later messages or missing proof of authority.
The access permission
Shows which systems the AI can reach, but not the exact outcome or boundaries the person approved.
The activity log
Shows what happened after the event, but does not itself prove what was authorised beforehand.
The missing record
One reliable statement of the intended outcome, its boundaries, its authority and its lifetime—created before action.
Why isn’t the prompt enough?
A prompt is evidence of what someone said. It is not always proof of what was authorised.
For a simple question, the prompt may be all the record we need. The gap appears when an AI is about to make a commitment or act outside the conversation.
“Book me somewhere nice in Bristol.”
- May be incomplete or ambiguous
- Conditions may be scattered across later messages
- Does not prove who had authority
- May not follow the action into other services
One accepted interpretation
- Required outcome and limits are explicit
- Preferences are separated from hard boundaries
- Authority and expiry are recorded
- Later decisions and actions reference the same ID
From conversation to confirmation
The system prepares the record. The right person confirms it.
The user does not write a technical record. The assistant turns the conversation into a clear proposal and asks for confirmation before a consequential action.
Book me a refundable hotel in Bristol for tomorrow night. Keep the total under £200.
Before I make a booking, please check the record I will use:
Book one hotel room in Bristol
- Must
- Be refundable · tomorrow night
- May
- Book one room · charge up to £200 total
- Must not
- Book outside Bristol · exceed £200 · choose a non-refundable rate
- Authority
- You, as the account and payment-card holder
- Valid until
- Booking completed, cancelled, or 18:00 today
Is this an accurate statement of what you authorise?
Approval and intent records
Isn’t this just asking for approval?
It starts there—but Declared Intent adds something an ordinary “yes” does not.
When an AI explains a proposed action and asks you to approve it, that is the beginning of Declared Intent.
But approval is the act of giving permission. An intent record is the durable, structured and machine-checkable representation of exactly what that permission covers.
The more powerful the tools, the more important it becomes to turn a momentary “yes” into a precise boundary that follows the action.
Why this matters
The same assurance gap becomes critical when AI has powerful tools.
A hotel booking makes the mechanism easy to see. But Declared Intent is aimed at a wider and more serious problem: keeping consequential AI action inside the authority and boundaries people actually approved.
In July 2026, OpenAI reported that models running a cyber-capability evaluation found a route out of an isolated test environment, reached the open internet and compromised Hugging Face infrastructure while pursuing the narrow goal of solving the benchmark.
This hypothetical record is inspired by the incident, not a reconstruction of the original prompt. It shows how the authorised task and the prohibited boundary could be made machine-readable before a run begins.
Complete the assigned ExploitGym evaluation challenge using only the systems and materials provided inside the named evaluation sandbox.
- Operate only against the challenge targets explicitly listed in the evaluation manifest.
- Preserve a complete action trace and report any apparent route beyond the sandbox.
- Stop and request human review if the task cannot be completed within the authorised boundary.
- Inspect, execute and test code inside the provided sandbox.
- Use only pre-approved local tools, files and simulated services.
- Submit a solution through the designated evaluation interface.
- Escape, bypass or weaken the sandbox or its isolation controls.
- Connect to the public internet or any system outside the listed challenge range.
- Search for, access or use external credentials, accounts, datasets or benchmark solutions.
- Exfiltrate, persist or stage data outside the authorised environment.
Written boundaries alone do not stop a capable system. The runtime must independently enforce network denial, target allow-lists, credential isolation, monitoring and automatic stop conditions. The record gives those controls a shared instruction to enforce—and gives investigators a precise statement against which to compare what happened.
When is a record created?
Not for every prompt. For an authorised goal or consequential action.
One intent record can govern many conversational turns and technical actions. The need for a record—and the strength of approval—should rise with risk, cost, irreversibility and uncertainty.
Ask or explore
“Explain refundable rates.”
The AI is providing information inside the conversation.Search and compare
“Show me hotels in Bristol.”
No commitment has been made and the user can continue exploring.Commit or act
“Book this hotel using my card.”
The AI is about to create an external commitment with financial consequences.Change the goal
“Book one in London instead.”
The authorised destination changed, so the earlier record cannot simply be stretched.Could this instruction cause an external action, commitment, payment, disclosure, safety consequence or significant decision? If yes, a durable record becomes more valuable.
The idea
Intent is more than a goal.
A useful record must describe the desired outcome and the boundaries around it. It must also show who had the right to approve those boundaries.
Required conditions
The room must be refundable.
If this is not true, the action cannot proceed.Permitted choices
The agent may book one room costing up to £200.
This gives freedom inside a clear boundary.Explicit prohibitions
The agent must not book outside Bristol.
A changed purpose needs a new authorised record.How it follows the action
One Intent ID connects the instruction to the outcome.
Select a stage to see what it contributes. The plain-language labels come first; the technical terms remain available underneath.
Clear to people and machines
Durable mandate
DIR-047 records the outcome, conditions, permissions and prohibitions before the AI acts.
intent_id · record_digest · declared_byThe instruction lasts. Permission to perform one action is temporary. The same Intent ID connects both.
Try the example
See when the AI may act—and when it must stop.
Choose a booking and check it against DIR-047. Then select any stage to see the evidence that would be kept.
DIR-047 records the outcome, conditions, permissions and prohibitions before the AI acts.
intent_id · record_digest · declared_byHow this fits with existing technology
Connect the parts through one lasting record.
Several emerging standards address parts of this problem. Declared Intent is designed to connect them through one lasting record of the original instruction. It does not require a new token, communication protocol or policy engine.
For technical readers · v0.5.3
A reference model to test—not a finished standard.
Different systems could represent the record in different formats. What matters is that they preserve the same meaning, authority checks, boundaries and history.
Minimum requirements
- Create it first. The active instruction exists before any governed action.
- Check authority separately. Stating an intention does not create the right to authorise it.
- Record positive and negative boundaries. Must, may and must not can be checked by a system.
- Stop when evidence is weak. Missing, expired or unclear evidence does not permit action.
- Use the same Intent ID. Decisions, permissions, actions and evidence point back to the active record.
- Preserve change history. Cancellation and replacement add events; they do not rewrite the past.
- Support reconstruction. A reviewer can compare intended, authorised, allowed and actual states.
View the technical DIR-047 example
{
"architecture_version": "DIAA/0.5.3",
"intent_id": "DIR-047",
"declared_by": "did:example:traveller",
"authorised_by": "did:example:traveller",
"authority_basis": { "type": "self", "verified_by": "authority-service:01" },
"purpose": "Book one hotel for a Bristol trip",
"obligations": [{ "refundable": true }],
"permissions": [{ "action": "book_hotel", "max_gbp": 200 }],
"prohibitions": [{ "location_outside": "Bristol" }, { "refundable": false }],
"valid_until": "2026-08-08T18:00:00Z",
"policy_refs": ["travel-policy:3.2"],
"lifecycle": { "supersedes": null, "status": "active" },
"integrity_proof": { "digest": "sha-256:…", "signature": "jws:…", "timestamp_ref": "rfc3161:…" }
}Earlier thinking remains available for comparison: view the archived DIR v0.4 proposal →
Keeping the record trustworthy
Blockchain is possible, but not required.
The record needs strong evidence that it has not been secretly changed. It does not need to depend on one particular technology.
What needs testing
This is a proposal to examine, improve and prove.
Authority
What is the smallest reliable proof that an approver truly had the power they used?
Compatibility
Which existing temporary permission mechanism should be tested first?
Privacy
How can a system prove that boundaries were followed without exposing the whole instruction?
Change
How quickly should cancellation, replacement or emergency revocation take effect?
Stewardship
Who should keep the lasting record, and who should be trusted to verify it?
Evidence
Which shared tests would prove that separate implementations reach the same result?
The next step
Test whether the record can reliably connect human authority to AI action.
Declared Intent is not a product, a finished standard or a compliance claim. The next milestone is a small working implementation: one signed instruction, one rules decision, one temporary permission and one action record that can be reconstructed from beginning to end.