DARINORCold Lab

// Writing

Your AI Agent Has No Audit Trail

By
8 min read
Incident response#agentic-ai#ai-security#incident-response#detection-engineering

The agent did something wrong on Tuesday. By Friday the change is reverted, the ticket is closed, and the review has stalled on the only question that matters: what did it see? Not what it did — the tool log answers that. What was in its context when it decided. Which retrieved document, which memory entry, which version of the system prompt. The things an incident review actually turns on.

Nobody knows. That's the honest answer, and the uncomfortable part is that it's by design. The model's side of the story lived in a context window that was reclaimed the moment the run ended. Your provider kept a copy, but the retention clock on that copy belongs to their abuse-monitoring policy, not your incident process. OpenAI's published API default holds inputs and outputs for up to 30 days; under zero-data-retention terms — the arrangements compliance was happiest to sign — it holds none.

Every other system you operate is assumed to keep evidence. Databases keep write-ahead logs. Firewalls log the traffic they pass and the traffic they block. The agent — the component most likely to do something nobody expected — is the one where keeping nothing is the default, and where "the logs will tell us" quietly means the tool calls. Which are the least of it.

Tool logs are not an audit trail

What you have today, if you have anything, is a tool-call log: tool name, arguments, sometimes the result, a timestamp, maybe a trace ID if your platform emits spans. It looks like an application log, so it gets treated like one.

For ordinary software that's mostly fair. The code is versioned, the callers are known, and "why did it do this" resolves to a commit hash. For an agent the tool log is the least interesting layer. It records that the agent called delete_user with id=4471. It does not record that the instruction arrived inside a support ticket the agent ingested ninety seconds earlier, or that a retrieved document described user 4471 as a stale test account. The call is a footprint. The decision is the incident, and the decision isn't logged.

This is the same blindness the memory work keeps running into: most of what steers an agent wrong arrives as context, not as commands. An indirect prompt injection never shows up in a tool log, because by the time the agent acts, the injection is just what it knew. A trail that doesn't capture what the model saw can't tell an agent that went rogue from one that was told to.

What the trail has to capture

Four things, and none of them are exotic. They're just not what logging defaults to.

What the model saw. The assembled context, per step: system prompt version, retrieved documents with their sources, memory entries, the user turn. Not a summary — the actual payload, or a pointer to it in storage you control. Every other answer hangs off this artifact.

Why each tool call happened. The plan or reasoning state that preceded the action, as close to raw as your provider exposes. Some models surface reasoning traces; some give you nothing. Where the reasoning isn't available, capture its inputs — the context above plus the tool schemas — so a reviewer can reconstruct the decision instead of guessing at it.

Whose authority it acted under. Which credential, which delegated identity, which scope. An agent acting as you puts your name on every log line, and "the agent did it" is not an answer an auditor accepts. Per-action provenance — identity, grant, task scope — is what turns attribution from an interview into a lookup.

What the environment looked like. Config, tool versions, allowlists, MCP server versions. The difference between "it worked last week" and "it deleted the wrong rows today" is usually drift, and drift is invisible unless the configuration was snapshotted per run. Diffing the agent's config against the run that failed is the quick version; the durable version is writing a config hash into every trace.

Retention is a design decision

The default direction of every layer is evaporation. Context windows get reclaimed. Provider copies expire on their clock — 30 days on OpenAI's default API terms, zero under ZDR. Harness traces, where they exist at all, tend to rotate with the application logs. Nothing in that chain is obligated to keep your evidence until your review is done.

So the decision is blunt: either you own a copy, or there is no trail. Storage you control, a retention window you chose, aligned with however long you keep authentication logs — decided before the incident, because afterwards the evidence is already aging out.

Zero data retention is a compliance feature. It is also an evidence policy you didn't write, and it expires at zero.

The containment post landed on one move worth reusing here: hash the exact payload before egress and log the hash. Do the same to prompts. Hash the assembled context, keep the full snapshot in cold storage. The hash proves what ran; the snapshot explains it. Hashing a prompt takes seconds. It's the storage policy that takes the meeting.

What exists today

Honest scoping, because this is where vendor decks get ahead of reality.

OpenTelemetry's GenAI semantic conventions cover model calls with attributes settled enough to build on. The agent-level attributes — session, agent identity, multi-agent handoffs — and parts of tool orchestration are still marked experimental, and their names have changed between releases (OpenTelemetry GenAI semantic conventions, 2026).

Vendor trace exports exist — most providers and agent platforms will hand you a run's spans and inputs — but they're vendor-shaped, they expire on vendor clocks, and they stop at the model boundary. None of them, as of now, covers the two things that matter most forensically: what was assembled into context, and why the model selected that tool. Trace formats record what happened. Reasoning capture is provider-specific, sometimes absent by policy, and never portable.

Directionally, the industry is moving from "log the calls" to "log the loop." Nobody has settled the second half. So design your trail so the parts that are yours — context snapshots, provenance, config state — don't depend on the parts that aren't.

What an audit trail cannot do

It can't be trusted blindly, because the trail is written by the thing under investigation. A compromised agent, or one running on poisoned context, writes its own alibi with perfect confidence. The mitigation isn't better formatting — it's plumbing: an append-only sink, separate credentials, a write path outside the agent's reach. If the agent can edit its own history, you don't have an audit trail. You have a diary.

It won't stay small. Full-fidelity context snapshots per step are enormous, and the storage bill will push you toward sampling. Sample the reads if you must. Never sample the writes — anything that mutates state, spends money, or sends a message gets the full record. And write the sampling decision down, because "the trace for that step was sampled away" is a bad sentence to say in a postmortem.

It isn't neutral data, either. A snapshot of everything the model saw is a snapshot of everything your org fed it — customer records, credentials that leaked into prompts, documents nobody should have been able to retrieve. Your audit trail is a new data store with breach surface of its own. Classify it like the data it mirrors, not like a log file. A trail that becomes the incident is a known genre.

And replay is not truth. Models are nondeterministic; re-running the prompt reproduces a plausible story, not the one that happened. Use replay to test hypotheses. Use the snapshot as evidence. The distinction is the whole game.

Tool logs record what the agent did. The trail an investigation needs is what the agent knew when it decided to do it — and by default, nobody keeps that.

If you run one agent in production, do this week: pick a run from last Tuesday and try to answer three questions. What did it see. Why did it call each tool. Whose credentials did it act under. Any answer that starts with "we'd have to ask the provider" or ends with "that's expired" means you don't have an audit trail. Start with the cheapest version anyway — context snapshot, provenance, config hash, into storage you control. The incident was never over when the fire went out; for agents, most teams never had the debris to begin with.

You can reconstruct a breach from packet captures and disk images because those systems were built to leave evidence. Agents aren't. Yet.