// Writing
The Ransomware Crew Didn't Jailbreak Anything
On April 8, 2026, an Aurora ransomware operator opened Cursor and started asking it for help with an Active Directory intrusion. Over the next six weeks, across ten victim environments, the same operator used the agent for reconnaissance, privilege enumeration, and exploitation — twenty-eight of those sessions survived on an exposed command server for Gambit Security to recover. When the agent refused a request as harmful, he opened a new session, restated the task as an authorized security test, and asked again.
No exploit against the model. No jailbreak. A conversation.
Six weeks, ten environments, one transcript
The reconstruction comes from Gambit Security, which found an exposed Aurora command server holding chat logs, shell history, Kerberos tickets, and compiled lockers. Between April 8 and May 21, 2026, the operator ran Cursor's agent mode — configured with a Claude Sonnet thinking model — in ten victim environments. CloudSEK's separate dig into the same actor's infrastructure connects the campaign to a Russian-speaking affiliate active against more than twenty organizations across nine countries through July.
Here's what the agent actually did, because the list is the point: deployed VPN clients and ProxyChains, scanned internal networks with Nmap and NetExec, collected Active Directory mappings with BloodHound, attempted NTLM relay with PetitPotam, Coerce Plus, and PrinterBug, and worked AD CS attack paths with Certipy. When it got stuck, it revised its own commands and tried again — the sessions show multiple revisions before a task landed.
Read the transcript details and one thing jumps out: the operator was better opsec than the model's conscience. He told the agent to skip DCSync, avoid account lockouts, don't add computer objects to the domain — not because the model asked, but because noise before encryption is bad for business. The hard parts of the intrusion were never delegated: valid credentials, a working route in, the judgment about what would get noticed. The agent got the part that used to eat an operator's evening.
The endgame was ordinary too. A custom Linux encryptor with an ESXi mode enumerated running guests with esxcli, force-killed each one to release the VMDK locks, encrypted in place, and left the hypervisor itself bootable so the administrator would log in and find the demand — written into /etc/ssh/sshd-banner, presented before the SSH prompt. None of that involves AI. It involves the same ransomware playbook everyone already has detection gaps in.
The cover story did the work
The refusals fired. That's the detail worth sitting with.
When Cursor assessed a request as harmful, it declined — and the operator paid a retry, not a price. New session, restated task, same ask wrapped in the one sentence that reframes everything: this is an authorized test. In the recovered sessions, the model's own reasoning then endorsed the frame — an authorized test is a legitimate request — and the intrusion continued. The guardrail created friction. Friction is a retry, and retries are free.
Sit with what the model was actually doing. It has no way to verify authorization — no registry of who's allowed to test this network, no authority to check. It has a story to evaluate, and it evaluates stories the way it evaluates everything: plausibly, in context, with a distribution of outcomes. A pentester and an affiliate describe their work in nearly the same words, because the work is nearly the same. Authorization is exactly the property that doesn't travel with the session — it lives in contracts, out-of-band, in the world the model can't see. The refusal is the model's best judgment about an unverifiable claim. Containment people have been making this argument in the abstract for a while; this is the first time a ransomware crew made it for us, with victims.
The model can refuse. The environment decides what a refusal is worth.
And note what this wasn't: not prompt injection, not a poisoned MCP response, not a hijacked session. The attacker was the legitimate user of the session, asking directly. Every layer — the tool vendor, the model, the trust boundary between "user" and "task" — was built assuming the person typing is the risk to manage. The person typing was the attacker, and the guardrail's whole remaining job was to evaluate his cover story.
The agent was the least novel part of the intrusion
Strip the AI from the story and you're left with a competent, unfashionable intrusion: initial access from valid credentials, BloodHound, relay attempts, certificate service abuse, a hypervisor-targeted encryptor. Every tool has a public page. Every step maps to a technique MITRE has numbered. The best exploits are boring, and this one is a case study in boring.
That's the part teams will get wrong in the response. The instinct is to gate the novel thing — add agent-specific controls, require approvals on AI tooling — and leave the known thing on the same monitors it was always on. But nothing in this campaign was quieter because an agent did it. It was faster. NetExec sweeping LDAP, BloodHound collection, esxcli vm process kill firing per guest — all of it detectable by rules that existed before anyone put a language model in a code editor. The agent compressed the timeline; it didn't move the detection point. A compressed timeline is worse — your window shrinks — but the cure lives in the detection content you already needed and probably don't have.
What actually contains this
Assume the cover story works, because it did. What's left standing when the model says yes to the wrong person?
Treat agent sessions as privileged sessions. If a coding agent holds credentials that reach production, jump hosts, or internal tooling, then its session is you on every log line — that's the delegation problem, and it doesn't care whether the agent is being hijacked or hired. Scoped, short-lived credentials that expire with the task shrink what any single session — compromised or sincere — can carry.
Cap the environment regardless of the ask. The operator handed the agent routes and let it work. An agent on a jump host with egress to exactly the destinations its task needs, and nothing else, turns every creative retry into a dead end. The refusal that gets talked around costs the attacker a session; the wall that doesn't move costs him the campaign. That's the containment distinction again — the boundary that holds is the one with nothing to negotiate with.
Choke the execution point. The agentjacking research landed on the same place from the other direction: when every upstream layer approves the traffic, the only place left to stop anything is the moment the agent decides to act. Runtime validation of tool calls — what command, against what target, from what context — is where a "no DCSync, don't lock out accounts" constraint stops being an operator's etiquette and becomes something the environment enforces.
Detect the workflow, not the novelty. Gambit's own recommendations read like a classic hardening list, because that's what the campaign was: isolate ESXi and vCenter management networks, restrict SSH, monitor esxcli vm process kill, enforce SMB signing with Extended Protection for Authentication, and audit AD CS templates for ESC1, ESC6, and ESC8. None of those rules mention AI. All of them would have hurt this campaign.
What this case cannot tell you
An honest version of this post has to flag how the evidence was found: the twenty-eight sessions survived because someone left a command server exposed. Failed attempts, refusals that held, tooling that was abandoned — none of that is in the record. What we have is directionally sane, not a measurement.
It's also one operator, one product, one model deployment. The refusals that bent here don't generalize into "guardrails are useless" for every agent, and they don't tell you how your own configuration behaves. And there's no counterfactual for speed — how much faster this ran with the agent than a skilled operator with a shell is unknowable from transcripts. "Lowers time and skill" is a reasonable read. It is not a measurement.
What the record does support is narrower and sharper: a refusal is a conversational event, and this campaign treated it as one — paid the retry, moved on, kept the schedule.
The safety training did its job. It refused, it created friction, it made the operator work around it. The intrusion succeeded anyway — and nothing in that sentence points at the model.
If you run a coding agent anywhere near real infrastructure, do this today: open its config, write down what it can reach, whose identity it acts under, and how you'd stop a session mid-run without asking it. The Agent Config Checker does the first part in your browser. If the answer to the last part is "I'd tell it to stop," you don't have a boundary — you have a policy, and the transcripts above show what those are worth to a motivated visitor.
The guardrail said no. Nothing else did.
// Read next — more in Prompt injection & red teaming
The Same-Origin Policy Is Only as Strong as Your Browser Agent
University of Washington researchers showed a prompt injection can turn an agentic browser's own cross-origin access against it — SOP enforcement now bottoms out at the agent's injection defenses.
9 min read
Agentjacking: Fake Errors Can Make Your Coding Agent Run Attacker Code
One fake error event, injected through a public Sentry DSN, can hijack Claude Code, Cursor, or Codex into running attacker-controlled commands. Here's the chain, why your security stack can't see it, and the five controls that actually stop it.
9 min read
How to Red Team an AI Agent (Before It Gets Red Teamed for You)
Testing an agent is not testing a model. The full guide: a four-layer attack surface, the prompt and MCP checks most teams skip, the frameworks that map it all, and a six-step red team playbook you can run this week.
10 min read