DARINORCold Lab

// Writing

How to Build an Agent Audit Trail Before You Need It

By
9 min read
Incident response#agentic-ai#incident-response#detection-engineering#agent-architecture

Two clocks decide whether an agent incident can be investigated at all. The provider's retention clock — 30 days on OpenAI's published API defaults, zero under zero-data-retention terms — is theirs. The clock on storage you control is the one you haven't picked yet, because there's nothing on it.

"Add logging" is the reflex here, and it fails in a specific way: everything logs fine, and the review still opens on three questions with no answers. What did the model see. Why did it pick that tool. Whose authority did it act under. None of those is a log line. They're artifacts — and artifacts are built, not enabled.

The case that they're missing is a separate post. This one assumes you're convinced and asks the harder question: where does each capture live, and what does the build actually look like? The good news is that you already run the component that can do all of it. The harness sees every assembled context before the model does and every tool call before it executes. It's the only layer in the stack that touches everything worth capturing — which makes it the only place a trail can come from without new infrastructure between your agent and its model.

Logging is a pipeline you switch on. An audit trail is a set of answers you guarantee. Build for the guarantee and the pipeline comes free.

Capture where the harness already touches

Four artifacts carry an agent investigation: the assembled context, the reasoning state, the per-action identity, and the environment's configuration. Each has a natural capture point in a loop you already run.

Snapshot the payload before the call. The harness assembles context — system prompt, retrieved documents, memory entries, the user turn — and the moment before the provider call is the only moment that payload exists in one place. Hash it with something fast; SHA-256 takes seconds. Write the hash into the run record and park the full snapshot in storage you control. The hash proves what ran. The snapshot explains it. If storage is tight, hash everything and snapshot what's expensive to reconstruct — retrieval results and memory entries top that list, because they're the parts that differ between runs and vanish after them.

Extend the tool gate. If you're validating tool calls before they execute, the validation point already sees the tool, the arguments, and the decision to allow. Extend the record it writes: append the context hash and whatever reasoning state preceded the call. Where the model surfaces a reasoning trace, store it. Where it doesn't — several providers never expose one, and some won't by policy — store its inputs: the context hash plus the tool schemas is what lets a reviewer reconstruct the decision instead of guessing at it. Not as good as the trace. Infinitely better than the shrug.

Record the delegation. Every scoped credential the harness mints — a workload identity, a per-task token, a delegated grant — is provenance waiting to be written down. An agent acting as you puts your name on every log line, and the countermeasure is per-action attribution: identity, grant, scope, stamped onto the same record as the action they authorized. Workload identity federation already scopes the machine hop; this is the field on the form.

Pin the environment state. Config, tool versions, allowlists, MCP server versions — snapshotted per run and hashed. "It worked last week" resolves to a drift diff or it resolves to nothing.

On standards: OpenTelemetry's GenAI conventions have settled the model-call attributes enough to build on, but session, agent identity, and multi-agent handoffs are still experimental, with names that changed between releases (OpenTelemetry GenAI semantic conventions, 2026). Emit what's settled. Don't hang the forensically load-bearing fields on attribute names marked experimental — the parts of the trail that are yours shouldn't depend on the parts that aren't.

And if you run a gateway between your application and the provider, it's a second capture point for the payload — useful redundancy for the context layer. But gateways see payloads, not reasoning and not identity. They're a supplement to the harness capture, never the trail itself.

Make the sink un-rewritable

Here's the uncomfortable inheritance from the diagnosis: the trail is written by the thing under investigation. A compromised agent — or one running on poisoned context — writes its own alibi with perfect confidence. The fix isn't a format. It's plumbing, and it's the same argument the containment work made: what the runtime never holds, no prompt can reach.

Concretely:

  • Append-only sink, separate credentials. The agent's runtime gets a write-only identity for the sink. It can append. It cannot read, update, or delete. The read path belongs to your review tooling — different credential, preferably different infrastructure.
  • The write path exits the agent's reach. If the trail lands in a store the agent can also touch through a tool — the same database it writes rows to, the same bucket its artifacts go to — assume it gets touched. One thing should be able to write the sink: the harness.
  • Version the writer. Record schema version, harness version, model version on every entry. A trail you can't interpret six months later was never evidence; it was noise with a retention bill.

The agent doesn't need write access to its own history. Taking it away isn't an optimization — it's the line between evidence and autobiography.

The record itself

One append per step. Six fields:

FieldExampleWhy it exists
run_id, step7f3c…, step 4 of 9Correlates every artifact to one run, in order
context_hashsha256:a6f1…Proves the exact payload without storing it inline
context_refyour storage URIPoints at the full snapshot
tool + argsdeploy.rollback {env: prod}What it did, verbatim, post-validation
decision basistrace ref or context hashWhy the call happened, or the inputs that reconstruct it
identity + grantwid/deploy-bot, scope prod:readWhose authority, exactly
env hashconfig sha256:9d2e…The environment the run actually ran in

That's seven. The seventh is the one most teams skip and every review needs. The journal your harness should already keep — one record per logical action, append-only — is this record's skeleton; the audit fields bolt provenance and context onto what you already write. You're not building a new pipeline. You're thickening one.

Retention, sampling, drift

Retention is a decision, and the diagnosis post's arithmetic still applies: every default layer evaporates, so either you own a copy or there is no trail. Storage you control, a window you chose, aligned with however long you keep authentication logs — that's the evidence class an agent incident actually belongs to. Decided before the incident, because afterwards the evidence is already aging out on somebody else's clock.

Two pressures will attack the decision, and both have answers:

  • Cost. Full-fidelity snapshots per step are enormous. Sample the reads if you must. Never sample the writes — anything that mutates state, spends money, or sends a message gets the full record. And write the sampling policy down next to the retention window, because "that step was sampled away" is a bad sentence to say in a postmortem.
  • Drift. "It worked last week" resolves to a diff or it resolves to nothing. The per-run env hash makes drift visible after the fact; the cheap ritual makes it visible before — diff this week's agent config against last week's on a schedule, the way you'd review a firewall rulebase. Same discipline, smaller blast radius.

And classify the trail like the data it mirrors, not like a log file. A snapshot of everything the model saw is a snapshot of everything your org fed it — customer records, credentials that leaked into prompts, documents nobody should have been able to retrieve. The audit trail is a new data store with breach surface of its own. A trail that becomes the incident is a known genre; don't author a new entry.

What this build cannot do

It cannot capture what a provider won't expose. Where reasoning isn't surfaced, you hold inputs, and reconstruction is inference. Say so in the review: "here's what it saw; here's the decision it plausibly made." A trail that overstates its own certainty is just a better-formatted guess.

It cannot cross the trust boundary. When your agent hands work to another agent — another team's, another org's — your record ends where their loop begins. The graph has no boundaries of its own, and the other side's trail is theirs to fail to keep. Record the handoff itself: what was sent, under which grant, to whom. That much is yours.

It cannot make replay the truth. Models are nondeterministic; re-running the prompt reproduces a plausible story, not the one that happened. The snapshot is evidence. Replay is a hypothesis test. Keep the distinction — it's the whole game.

And it rots quietly if nobody reads it. Schemas drift. Retention jobs truncate on schedule. The sink fills. A trail that has never been pulled is a backup that has never been restored — technically present, functionally a rumor. Drill it: once a quarter, take one production run and answer the three questions from your own copy, cold.

An audit trail nobody has exercised isn't a control. It's a liability with a retention bill. Pull one this quarter — the incident is a terrible time to learn your own schema.

If you run one agent in production, do this week: add two fields — the context hash and the identity stamp — to whatever journal it already writes, pointed at storage it can't rewrite. Then pull one run from last Tuesday and answer, from your own copy: what did it see, why each tool call, whose authority. Any answer that starts with "we'd have to ask the provider" means the build isn't done.

Packet captures and disk images exist because those systems were built to leave evidence. Agents weren't. Build yours like the review is already scheduled — one day, it will be.