When Your Coding Agent Publishes Your Secrets: An AI Forensics, Containment and Audit Playbook

Read Time: 20 minutes

TL;DR

I was recently brought into an incident that I expect to see a lot more of: a coding agent packaged files it had access to and pushed them to a public GitHub repository. Inside were personal data and live tokens. Nobody attacked anything. The agent did something it had the permissions to do, and the internet did the rest. This is the playbook I would hand to anyone who finds themselves in the same spot: stop the agent and preserve its logs before they age out, rotate every secret before you touch the git history, accept that deleting or privatising the repository does not delete the data, clean up what you can with GitHub, audit what the leaked tokens were used for, and handle the personal data under GDPR’s 72-hour clock. The second half is the AI forensics most teams have never done: reading the OpenAI and Anthropic usage dashboards and APIs to work out whether someone else spent your tokens, and on what.


A note on what this is. The case that prompted this article is real, and it is anonymised: no client, no agent product, no data specifics. What follows is a defensive playbook built from that experience and from public documentation, which I cite as I go. Provider consoles change; check the current docs before you rely on a menu path under pressure.

The incident class nobody drilled for

We have twenty years of playbooks for a developer committing a password. We have almost none for the case where the developer is a piece of software with shell access, git credentials and a task to finish.

The mechanics of the case were mundane. A coding agent, working with the access it had been given, bundled up a set of files and pushed them to a public repository. Some of those files contained personal data. Some contained tokens, including API keys for AI providers. There was no prompt injection and no malicious skill involved, only an agent able to run git push against a public remote and nothing standing between that ability and the outcome.

The data says this is not a freak event. GitGuardian’s State of Secrets Sprawl 2026 counted 28.65 million new hardcoded secrets added to public GitHub commits in 2025, a 34% jump in a year, with AI-service secrets up 81% to over 1.27 million. The figure that caught my attention: commits assisted by one widely used coding agent showed a 3.2% secret-leak rate, against a 1.5% baseline across all public commits. That is vendor research and one data point, but it matches what I see: agents move faster than the hygiene around them. The same report found over 24,000 secrets in MCP configuration files alone. I have written before about weaponised agent skills and the dependency trap in AI-generated code; this case is the less glamorous relative of both, because it needs no adversary at all.

The clock you are actually on

One fact decides the order of the steps: the race is against automation, not against a person who might stumble on the repository.

Public pushes are scraped continuously. Palo Alto’s Unit 42 documented a campaign that could launch a full attack within five minutes of an AWS credential appearing in a public GitHub repository, and Clutch Security’s later research found exposed AWS keys exploited within minutes. Assume that by the time a human noticed the push, every secret in it had already been harvested. From that assumption, the order of operations follows.

Step 0: Stop the agent and preserve the evidence

Kill the agent session. Revoke the git credential it pushed with (the SSH key, personal access token or app installation) so a retry loop or a scheduled task cannot push again while you work.

Then preserve evidence, today, before it disappears on its own. The agent’s transcript is your best record of what it did and why: which files it read, which commands it ran, what it pushed and where. Those transcripts are not kept forever. Claude Code, for example, stores sessions as JSONL under ~/.claude/projects/PROJECT/ and deletes them after 30 days by default (the cleanupPeriodDays setting). OpenAI’s Codex CLI keeps its session logs as JSONL under ~/.codex/sessions/. Whatever the agent, find its session store and copy it somewhere safe, along with the shell history, the local repository with its git reflog, and any CI logs if the push came from a pipeline. Hash the copies. You may need them for a regulator, a customer or a court, and an incident report that says “the logs had rotated” is not one you want to write.

Step 1: Work out exactly what leaked

Do not guess from memory. Clone the public repository into an isolated location and scan the full history, not just the current tree:

git clone --mirror https://github.com/
<org>/<repo>.git leaked.git
gitleaks git -v --report-format json --report-path gitleaks.json leaked.git
trufflehog git file://leaked.git --results=verified,unknown

Gitleaks gives you breadth; TruffleHog can check which credentials are still live, which tells you what to rotate first. Then go through the files by hand for personal data, because scanners are good at tokens and bad at a CSV of customer names.

Build one inventory table and keep it for the rest of the incident: each secret (type, provider, scope, owner, first commit it appears in) and each file of personal data (categories of data, approximate number of people, whose data it is). Everything after this step works from that table.

Step 2: Rotate everything before you touch history

This is the step teams get backwards. The instinct is to delete the file, force-push and feel better. GitHub’s own guidance on removing sensitive data is blunt: if the sensitive data is a secret, “as a first step you need to revoke and/or rotate that secret.” A rotated secret is harmless wherever copies of it survive. A secret scrubbed from history but still valid is still a working key in someone else’s harvest.

Order the rotation by blast radius. Admin and organisation-level keys first (an AI provider admin key, a cloud root or IAM key, a GitHub token with org scope), because those let an attacker create new credentials that survive your rotation. Then production service keys, then everything else. After each rotation, confirm the old credential is actually dead, not just that the new one works.

Two AI-specific notes. First, GitHub runs a secret scanning partner programme: when it detects a partner’s key in a public repository it notifies the provider, which may revoke it. Anthropic states that it automatically deactivates Claude API keys found this way and emails the affected user. OpenAI has been a GitHub secret scanning integrator since 2021, but its key-safety guidance simply tells you to rotate a leaked key immediately. Treat provider-side revocation as a bonus, never as your control. And keep that Anthropic email: its timestamp is a useful anchor for your timeline. Second, rotating an AI key breaks whatever used it. Have the new key ready in your secrets manager before you kill the old one, or you will be doing incident response and an outage at the same time.

Step 3: Stop the bleeding, and understand what that does not do

Now reduce further exposure: make the repository private or take it down. Do it, but be clear-eyed about what it achieves, because this is where most people’s mental model is wrong.

GitHub documents that when a public repository is made private, “its public forks are split off into a new network” and stay public. Delete it, and the oldest active public fork becomes the new upstream. Truffle Security showed in 2024 that as long as one fork exists, commits in the network remain reachable by their hash, and GitHub called that behaviour by design. Force-pushing does not help either: Sharon Brizinov scanned every force-push event since 2020 recorded in the public GH Archive, pulled the “deleted” commits back by their SHA, found thousands of live secrets and collected around $25,000 in bug bounties. The push event, with its commit hash, is in a public archive that you do not control.

So containment of the data is partial at best. Containment of the secrets is total, if you did Step 2. That is why Step 2 comes first.

srf_agentleak_exploitation_tree
Figure 1. Why “we deleted the repo” is not containment. An attacker needs only one route to the data (the live repository, a fork or clone taken before takedown, a dangling commit fetched by its SHA from GH Archive, or the wider fork network) and only one way to use it. The only branch you can close completely is the credential: once rotated, every copy of it is worthless. The personal data has no equivalent of rotation.

Before you flip the visibility, record what you can see. Under Insights → Traffic, GitHub shows clones and visitors for the past 14 days (UTC, updated hourly) to anyone with push access. Screenshot it, note the fork list and the stargazers. It will not tell you who, but it tells you whether anyone did, which matters a great deal for the personal-data assessment in Step 6.

Step 4: Clean up what GitHub can clean up

With secrets dead and the repository private, rewrite history properly. GitHub recommends git-filter-repo 2.47 or later with its sensitive-data mode:

git-filter-repo --sensitive-data-removal --invert-paths --path path/to/leaked-file
git-filter-repo --sensitive-data-removal --replace-text ../patterns.txt

Then contact GitHub Support with the repository name, the number of affected pull requests and the “first changed commits” that git-filter-repo reports, so they can remove cached views, dereference pull request refs and garbage-collect. For forks, the private information removal policy covers access credentials and third-party tokens, and personal identifiers such as government ID numbers; GitHub does not disable forks automatically, so list every fork you know of in the request, with file links, line numbers and why each item is a risk. Wider categories of personal data may not fit that policy cleanly, which is one more reason the GDPR assessment in Step 6 must assume exposure rather than hope for removal.

If you want to know what an outsider can still reach, run the same tooling an attacker would: TruffleHog’s github-experimental --object-discovery mode enumerates hidden and deleted commits in a repository network. It is slow and rate-limited, but it answers the question honestly.

Step 5: Audit what the tokens were used for

Rotation stops future misuse. It tells you nothing about what happened between the push and the rotation. For every live secret in your inventory, pull the provider’s logs for that window: cloud audit trails, GitHub’s audit log, database access logs. For AI keys, that means the provider dashboards and usage APIs, which most teams have never looked at with an investigator’s eye. That is the AI forensics part of the job, and the second half of this article.

One question to answer first: did any leaked key have admin rights? An admin key lets the attacker create new keys, add users or change settings, which is how you end up rotating the leaked key and still bleeding. For those keys, the audit log comes before the usage report.

Step 6: The personal data and the 72-hour clock

A public repository containing personal data is a personal data breach under GDPR: a loss of confidentiality, whether or not you can prove anyone downloaded it. Article 33 requires notifying the supervisory authority “without undue delay and, where feasible, not later than 72 hours” after becoming aware of it, unless the breach is unlikely to result in a risk to people’s rights and freedoms. Article 34 adds communication to the affected individuals when the risk is high. And Article 33(5) requires you to document the breach, its effects and your remedial action even if you decide not to notify.

In Spain, the AEPD provides Asesora Brecha to help decide whether to notify and Comunica-Brecha for the communication to individuals, alongside its notification guidance. The clock starts when you become aware, not when you finish investigating; you can notify in phases. Your Step 1 inventory and your Step 3 traffic screenshots are exactly what the notification asks for. And if you read my last article, you will know the AEPD has already received its first notification of a breach executed by an AI agent. An agent causing a breach by publishing data is a different case from an agent attacking you, but the regulator is clearly paying attention to both.

Assessing token use in the OpenAI dashboard

If an OpenAI key leaked, you want three answers: was it used after the push, for what, and did anyone use it to create more access.

The usage dashboard. At platform.openai.com/usage (organisation owners, or users with the Usage Dashboard permission), the bar along the top sets the scope: API sources, project, API keys and date range. Pick the affected project, narrow to the leaked key with the API keys selector, and set the range to cover a baseline week before the push and everything after it. Data is shown in UTC, and usage detail pages let you drop to a 1-minute interval, which is where abuse shows up as a wall of tokens at 3 a.m. Two parts of the page matter most in an incident. The API capabilities tab gives one card per service (Responses and Chat Completions, Images, Web Searches, File Searches, Moderation, Embeddings and more), so a key that has only ever done text and suddenly shows image requests stands out at a glance. The panel on the right breaks requests down by user, service and API key, and shows the month’s spend against your budget. The download button exports the data for your case file.

srf_agentleak_openai_usage
Figure 2. The OpenAI usage dashboard, filtered to one project over 30 days. Check every card in the API capabilities tab, not just Responses and Chat Completions: abuse often shows up in a service your application never uses. The panel on the right splits requests by user, service and API key, and shows spend against the monthly budget. Organisation and user names redacted.

Per-key numbers through the Usage API. The dashboard is good for eyeballing; for evidence, use the Usage API with an Admin key, which can group by api_key_id, project_id, model and more:

curl "https://api.openai.com/v1/organization/usage/completions?\
start_time=
<unix_ts>&bucket_width=1h&group_by=api_key_id&group_by=model&limit=168" \
  -H "Authorization: Bearer $OPENAI_ADMIN_KEY"

You get input_tokens, output_tokens, input_cached_tokens and num_model_requests per bucket, per key, per model. Separate endpoints cover embeddings, images, audio, web search and more, and /v1/organization/costs gives spend. Check them all: an attacker burning your key on image generation will not show up in completions.

srf_agentleak_openai_spend
Figure 3. The Spend categories tab splits cost per model into cached input, input and output, plus tool calls such as web search. A jump in output cost relative to input is one of the signals described below. Organisation and user names redacted.

When was the key last used? The API keys page answers this directly: every key has a Last used column next to its creation date, expiry, creator and permissions, and an API Key Usage button at the top of the page links to usage. The Admin API exposes the same value as last_used_at on the project key object. If the date is later than your rotation, you rotated the wrong key. If a key you had forgotten about shows recent use, start there.

srf_agentleak_openai_apikeys
Figure 4. The OpenAI API keys page. The Last used column is the fastest answer to “was this key used after the leak?”. Note the Expires column reading Never: a leaked key with no expiry stays valid until someone rotates it. Key names, tracking IDs, key hints and creator redacted.

Did anyone create more access? This is what the Audit Logs API is for: API key creation, updates and deletion, user and service account changes, login failures, project and settings changes, with event types such as api_key.created. The catch: audit logging only records from the moment an owner enables it (organisation settings → Data controls → Data retention), and once on it cannot be switched off without contacting support. If it was not enabled before the incident, you have no history. Enable it today.

What did they ask for? Responses API calls are stored for 30 days by default unless your organisation has Zero Data Retention, and you can browse them under Logs in the dashboard, with separate tabs for Responses, Agents, Realtime, Completions, Conversations and more. Filter to the exposure window and look for prompts, models or tools you do not recognise. Opening a single entry shows its timestamp, model, token count and the tools it used, which is often enough to tell your application’s traffic from someone else’s. The dashboard itself now warns that Responses logs older than 30 days will soon no longer be available, so export what you need on day one. Beyond that, OpenAI keeps abuse-monitoring logs for up to 30 days; if you need content or source details you cannot see, open a case through the help centre quickly, before that window closes.

srf_agentleak_openai_logs
Figure 5. The Logs page, Responses tab: one row per stored call, with its model and timestamp. Read the banner: logs older than 30 days are going away, so export early. Prompt text redacted.

srf_agentleak_openai_log_detail
Figure 6. A single logged response. The properties panel gives the timestamp, model, token count and tools used (here, web search): per-request detail the usage charts cannot give you. Prompt and response ID redacted.

Assessing token use in the Anthropic Console

The questions are the same; the tools differ.

Check your inbox first. If GitHub’s partner scan caught the key, Anthropic will already have deactivated it and emailed you. That email gives you an upper bound on the exposure window for that key.

The Console pages. Under Analytics, the Console has Usage and Cost pages. Both report in UTC, filter by workspace, API key and model, group the chart by model, and export the data with the download button; Usage also filters by account. Anthropic’s own advice is to review usage patterns per key regularly. During an incident, select the leaked key, set the range across your baseline week and the exposure window, and group by model. Usage gives you tokens in and out; Cost splits the money into tokens, web search, code execution and session runtime. The same Analytics menu also has a Logs page. I could not find public documentation of what it records, so open it early in the incident and check whether it covers the window you need.

srf_agentleak_claude_usage
Figure 7. The Claude Console Usage page over 30 days, grouped by model. Select the leaked key in the API key filter and compare the exposure window with the days before the push.

srf_agentleak_claude_cost
Figure 8. The Cost page uses the same filters and splits spend into token, web search, code execution and session runtime costs. A model your team does not use appearing in the legend is worth a closer look.

Looking at a previous month. The Cost column on the API keys page only covers the current month. On the 1st it goes blank, and a key that was busy last month suddenly looks unused. To see an earlier month, open Cost, set Range to Last month (the arrows next to it step back one month at a time) and switch the chart to the table view for a day-by-day breakdown. The download button exports the same period broken down by key, day, model and token type, with each key’s ID and status. The Cost API cannot give you that, because it only groups costs by workspace or model, so make this export one of the first things you save.

srf_agentleak_claude_cost_lastmonth
Figure 9. The Cost page set to Last month, in table view: one row per day, one column per model, and the billed total. Use the arrows next to Range to step back to the month of the leak.

Summarised by key, that export for one of my own accounts over September looks like this (key names replaced):

Key Status September cost Days with spend First and last day
Key A Disabled $104.87 17 7 → 25 Sep
Key B Active, created 25 Sep $21.49 4 25 → 30 Sep
Key C Active $10.98 5 9 → 18 Sep
Key D Active $0.11 2 7 → 14 Sep
Total $137.45

Two things are worth reading from it. Key A, now disabled, goes quiet on 25 September, the same day Key B was created and starts spending. That is the pattern you want to see after a rotation: the old key stops on the day you disable it and the new one takes over. A disabled key that still shows spend after its switch-off date means you disabled the wrong key, or it is not really dead. The second thing is the mix: about two thirds of the month (around $89) went on writing to the prompt cache, which is normal for an agent working with long contexts. Your usual mix of token types is a fingerprint of your own workload, and someone abusing a stolen key with short, uncached prompts would leave a very different one.

Per-key numbers through the Usage and Cost API. With an Admin API key (sk-ant-admin01-…, created by organisation admins), the usage report groups by api_key_id, workspace_id, model and more, in 1-minute, 1-hour or 1-day buckets, and data typically lands about five minutes after the request:

curl "https://api.anthropic.com/v1/organizations/usage_report/messages?\
starting_at=2026-09-01T00:00:00Z&ending_at=2026-09-25T00:00:00Z&\
bucket_width=1h&group_by[]=api_key_id&group_by[]=model" \
  -H "anthropic-version: 2023-06-01" \
  -H "x-api-key: $ANTHROPIC_ADMIN_KEY"

The response separates uncached_input_tokens, cache_read_input_tokens, cache_creation_input_tokens, output_tokens and server tool use. /v1/organizations/cost_report gives the money. If your developers use Claude Code under the organisation, the separate Claude Code Analytics API breaks cost down per user.

Disable and inventory keys. In the Console, Organization settings → API keys lists every key with its linked account, workspace, expiry, creator and cost, and filters by creator, linked account, status and workspace. The menu at the end of each row disables a key, and later re-enables or deletes it. The Admin API does the same programmatically: it lists keys and can set a key’s status to inactive or archived. It cannot create keys, so if your inventory shows a key nobody recognises, it was created through the Console, which points to a compromised account rather than a leaked key. That changes the scope of the incident.

srf_agentleak_claude_apikeys
Figure 10. Organization settings → API keys in the Claude Console. The dimmed row is a disabled key; its menu offers to re-enable or delete it. The Created by and Status filters help you spot keys nobody remembers creating, and the Expires column shows which keys will die on their own. Key names, key hints, linked accounts and creators redacted.

Who changed what. The Compliance API’s Activity Feed, readable with an Admin API key from the Console where the Compliance API is enabled for your organisation, records administrative actions such as creating an API key or adding a member to a workspace. Anthropic is explicit that it does not log inference activity, so it will tell you if someone minted access, not what they asked the model. For request-level questions, go to Anthropic support.

Reading the numbers like an investigator

Neither provider hands you an “attacker” label. You are looking for usage that does not fit your own pattern, so always compare the exposure window against a baseline from before the push. The signals I look for:

  • Traffic on a key after its rotation time, or on a key that should have been idle (a dev key, an old project key).
  • Models your team does not use, especially the most expensive ones, or new services (image, audio, batch) appearing on a key that only ever did text.
  • A jump in output tokens relative to input, which looks like bulk generation rather than your application’s usual pattern.
  • Activity in hours or at a steady machine rate that matches nobody’s working day.
  • New keys, service accounts, members or workspaces you did not create, in either audit log.

What you will not get from either dashboard is a source IP. Usage is aggregated by key, model and time, not by caller. If you need to attribute or prove access, that is a support request, and it is time-sensitive.

The part that stops the next one

Containment is the urgent part. The root cause is almost always the same: the agent had more reach than its task needed. The fixes are not exotic:

Take public publishing away from the agent. It should not hold a credential that can push to a public remote, change repository visibility or create repositories. Use a fine-grained token scoped to the one repository it works on, and require a human for git push and anything touching visibility (every serious coding agent supports deny rules or approval prompts for specific commands).

Keep secrets out of the agent’s working directory. No .env files with production keys where the agent can read them, no admin keys in its environment. If it cannot read the secret, it cannot publish it.

Turn on push protection and pre-commit scanning. GitHub’s push protection and a gitleaks pre-commit hook would very probably have blocked the token part of this incident before it reached GitHub, as long as those token types are among the patterns they recognise. They would not have caught a spreadsheet of personal data, which is why the first two controls matter more.

Scope and cap your AI keys. One key per project and environment, budgets and spend limits set, admin keys used only by the people who need them. A leaked project key with a spend cap is an annoyance; a leaked admin key is an incident.

Keep the evidence longer than the default. Raise the agent’s transcript retention (30 days is too short for an investigation that starts on day 25), and switch on OpenAI’s audit logging and Anthropic’s Activity Feed collection now, while you have nothing to investigate.

The checklist

For the day it happens to you, in order:

  1. Kill the agent session and revoke its git credential.
  2. Copy the agent transcripts, shell history, local repo and reflog; hash them.
  3. Mirror-clone the public repo; scan full history; build the secrets and personal-data inventory.
  4. Rotate every secret, admin keys first; verify the old ones are dead.
  5. Screenshot traffic, forks and stars; then make the repository private.
  6. Rewrite history with git-filter-repo; contact GitHub Support; file removal requests for known forks.
  7. Pull provider logs for every leaked secret: usage, audit, activity. For AI keys, use the dashboards and APIs above.
  8. Run the GDPR assessment; notify within 72 hours of awareness if required; document either way.
  9. Fix the agent’s permissions before you let it run again.

So what

The uncomfortable lesson of this case is that no one did anything malicious, and it still became a breach with regulatory consequences. We spent the last two years worrying about agents being turned against us by prompt injection and poisoned tools. Those threats are real. But the more common failure is simpler: we gave software the same credentials and reach as a senior engineer and none of the judgement, then pointed it at a deadline.

Every team running coding agents should rehearse this incident before it happens. Pick a test repository, plant a canary token, let the agent push it, and time how long it takes you to rotate, audit and answer the regulator’s questions. If you cannot read your AI provider’s usage by key today, when nothing is on fire, you will not learn it at 2 a.m. with the 72-hour clock running.

Stay paranoid. Rotate first. And never give an agent a push it does not need.

Further Reading:

Questions or feedback? Reach out via:

Contact: info@vulnex.com

Posted in AI, Business, Security, Technology | Tagged , , , , , | Leave a comment

Intent, Not Sophistication: The AI Attacker Is on the Record

Read Time: 16 minutes

TL;DR

Two documents landed within days of each other this month, from two sources with nothing in common except the thing they describe. On September 14, Spain’s data-protection regulator, the AEPD, announced it had received the first GDPR breach notification for an attack executed by an autonomous AI agent: a vulnerability scan, a valid login, an autonomous hunt for holes in the application, modified personal data and accessed invoices, all chained by the agent without a human steering. Days earlier, Anthropic had published a threat report cataloguing its own model being used the same way at global scale: agent swarms, malware that rebuilds itself to dodge detection, a solo hacktivist running operations that used to need a state team. A regulator with no product to sell and a vendor with every incentive to look good, independently describing the same animal. The headline the vendor hands you is blunt (“sophisticated attacks no longer require sophisticated attackers”) and the sentence I would frame is sharper still: “the main distinguishing feature between these classes of actors is no longer sophistication but intent.” The skill gap did not shrink. It collapsed, and it is now on the record in a Spanish regulator’s blog. This is a practitioner’s read, not a summary: the three shifts that actually change your job, why the regulator’s case matters more than the vendor’s report, and a Monday checklist that the two sources converge on almost line for line.


A note on what this is. I disrupted none of these operations and I can independently verify none of them. One source is a vendor reporting on its own model; the other is a regulator summarizing a filing it received. I will take both seriously, because they line up with each other and with what the rest of us see from the outside, and I will be clear about where the vendor’s incentives and the reader’s interests part ways. Defense-oriented throughout, no operational detail, threat-vector level: the usual house rules.

Every so often two documents land that say, in voices with more data than mine, the thing you have been saying into the wind for a year. This month I got two, from opposite ends of the world and opposite ends of the incentive spectrum.

The first is small, Spanish, and for my money the more important. On September 14 the AEPD, Spain’s data-protection authority, announced it had received its first notification of a personal-data breach in which the incident, in the regulator’s careful conditional, “would have been executed” by an AI agent running on a well-known language model. Not AI-assisted, the phishing-email-written-by-a-chatbot genre we have all seen. AI-executed. According to the account in the notification, the agent started with a vulnerability scan against generic files, achieved a valid login, and once inside went looking, on its own, for vulnerabilities in the application; when it found one, it modified personal data and reached invoices. The AEPD is explicit that this is a qualitative change: an agent, in its words, “can receive an objective, plan intermediate tasks, use tools, execute code, consult sources, interpret results and modify its behaviour, autonomously, depending on what it finds.” A third party, the regulator says, used it “as an instrument to successfully chain the different phases of the attack.” The company is unnamed, the account comes from the company’s own filing and the AEPD says it still needs further analysis. The regulator is also careful on a point I want to keep: using a given model does not mean the model or its provider’s infrastructure was compromised, nor that the tool was designed for malicious use. What matters is that a European regulator has now put on the record, from a real case, the thing I have been arguing all year.

The second document is big, American, and comes with a caveat I will get to. Days before the AEPD post, Anthropic published its September 2026 threat intelligence report, cataloguing, across seven harm areas and eight months (December 2025 to August 2026), a parade of “Generative Threat Groups” that used its model, Claude, to run cyber operations, influence campaigns, fraud, and surveillance. Where the AEPD gives you one company and one agent, the vendor gives you the same shape at planetary scale.

My own paper trail on this is long. In 2026 I have written that the model is becoming the attacker, that agent skills are a loaded weapon, and that the AI supply chain is a soft target. Earlier this month I wrote about a frontier-lab researcher resigning with a warning that these would soon be “systems that can hack anything,” and then about the CEO’s call to slow down and why pacing the frontier does nothing for the capability already loose in the valley. This piece is the receipts, and the point of leading with the Spanish case is that they are no longer only the vendor’s receipts.

Let me do what a practitioner should do with documents like these: not recap them (the summaries will be everywhere) but pull out what actually changes your job, and be honest about what to trust.

The one sentence that matters

Strip both documents to a single load-bearing claim and it is the vendor’s: “the main distinguishing feature between these classes of actors is no longer sophistication but intent.”

For my entire career, the threat model has been layered by capability. Script kiddies at the bottom, organized crime in the middle, nation-states at the top, and your defenses calibrated to who you thought would bother with you. That ladder is the thing both documents say has fallen over. When a solo hacktivist can field the same autonomous tradecraft as a state espionage team, and when an unidentified third party can point an off-the-shelf agent at a Spanish company and walk away with modified records and invoices, “who is sophisticated enough to hurt me” stops being a useful question. The only variable left is who wants to, and intent is cheap, plentiful, and impossible to patch.

The vendor says the quiet part directly: “sophisticated attacks no longer require sophisticated attackers,” because AI “has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators.” The regulator says the same thing in drier prose: AI agents do not introduce new techniques, but they “increase the speed, scale and adaptability of already known malicious techniques, reducing the time available to detect and contain them.” Same old attacks. New tempo, new operators.

One more detail from the vendor’s report deserves its own sentence, because it changes how you should read everything below. Anthropic states that the misuse cases ran on its Haiku, Sonnet and Opus models, and that “none of the misuse cases involved the use of Claude Fable or Mythos-class models, with the exception of one illicit distillation case.” Nobody in this catalogue needed the frontier. Every operation in it ran on the tier that is already everywhere, already cheap, and whose rough equivalents you can download and run with nobody watching. That is the valley I keep pointing at, described by the lab that owns the summit.

Three shifts that actually change your job

Under the case studies, three structural shifts are doing the real work. These are the parts worth your attention, and I will note where the Spanish case and the vendor’s data say the same thing.

1. The skill gap collapsed

The vendor’s cast is the tell. A Chinese espionage group (designated GTG-10007) whose operators include, per the report, two university undergraduates in Changsha, ran “agent swarms” that decomposed reconnaissance into parallel subagents and maintained an autonomous vulnerability-research program that turned up multiple previously unknown vulnerabilities in security products. A single French hacktivist (GTG-50029) got into at least fourteen organizations and built a doxxing platform with real ingestion pipelines. A financially motivated operator (GTG-50014) directed the AI toward goals and let it “evaluate environments and execute iteratively” (the report calls this “vibe hacking,” and yes, the name stings), then decompiled 1.8 million Android APKs to harvest hardcoded secrets and walked out of one airline with tens of millions of passenger records.

Now put the AEPD case next to that. One agent, one objective, one company, and the same phases (scan, access, vulnerability hunt, data) chained autonomously. It is the vendor’s pattern at the smallest possible scale, and it is the scale most of my readers actually live at. None of these people, on paper, should be able to do what they did. The AI is the force multiplier that lets intent skip the decade of skill-building it used to require.

2. The inverted cost structure: the loop closes faster than yours

This is the most important technical idea in the vendor’s report, and the one defenders should lose sleep over. In the Russian espionage case (GTG-20006), when a security product flagged the group’s malware, “agents would then set about the process of autonomously modifying and rebuilding the malware to evade the existing detections.” Anthropic draws the conclusion plainly: previously a defender could slow an attacker by shipping a new detection; now “capable adversaries can ‘close the loop,’ bypassing traditional security detections faster than defenders can develop and deploy them.”

The AEPD reaches the same place from the other direction. Among its lessons: procedures designed for manually executed attacks “may prove insufficient when an agent analyses multiple assets simultaneously,” and while human oversight “remains indispensable,” it “must be supported by detection, containment and response mechanisms able to operate fast enough.” A regulator and a vendor, independently, describing the same race.

Read that as a security engineer and it is a phase change. Detection has always had a shelf life, but the shelf life was measured against human attacker tempo. When the attacker’s evade-rebuild-redeploy loop is autonomous, it can spin faster than your observe-analyze-write-test-ship loop, which still has humans in it. The economics of defense have quietly inverted: the side that can close its loop fastest wins, and for the first time that is not automatically the defender.

srf_atr_inverted_cost_loop
Figure 1. The inverted cost structure. The attacker’s loop (deploy, get detected, let the model rebuild and mutate the malware, redeploy) can now close in hours, autonomously. The defender’s loop (observe the new technique, analyze it, write a detection, test and ship it) still has humans in it and closes in days or weeks. When the red loop spins faster than the blue one, a detection is obsolete before it is deployed.

3. The AI supply chain is now a target, not just a tool

I have been banging this drum since the dependency-trap work, and the vendor’s report escalates it from theory to campaign. One group (GTG-50020) compromised an AI vendor’s evaluation sandbox, extracted production API keys from multiple providers, hit around thirty AI companies in about four days, and, the detail I cannot stop thinking about, explicitly went looking for pre-release model access (they failed; every path they tried was closed). Another (GTG-50021) ran a fraudulent Claude reseller, offering cheap access while proxying traffic elsewhere and harvesting the credentials for resale. Others prompt-injected LiteLLM wrapper services to pull production keys straight out.

The AEPD, from its side, lands on the same object. Its words: “an agent that obtains an account, an API key or a token with excessive permissions can operate at the speed of a machine.” The vendor’s own recommendation is the one I would have written, treat “AI keys and agent integrations with the same level of seriousness” as production credentials, and the regulator’s is the mirror image, from the victim’s end. An API key to a capable model is loot (resold), compute (your bill, their attack), and cover (their activity, your name), all at once. If your threat model still files “AI keys” next to “SaaS logins we’ll rotate eventually,” it is out of date, and now a regulator agrees.

Put the three shifts together and they form a single structure. I modelled it as an attack tree, built entirely from the vendor’s own report, because the point is not any one branch but the shape of the whole.

srf_atr_ai_enabled_operation_tree
Figure 2. Anatomy of an AI-enabled operation, drawn from the vendor’s report. Four phases chained with AND: lower the skill floor (off-the-shelf frameworks like PentAGI, vibe hacking, agent swarms), run the autonomous loop (recon and exploitation, autonomous vulnerability research, self-modifying malware, unattended bulk exfiltration), target the AI supply chain (harvest production keys, compromise an eval sandbox, fraudulent reseller, prompt-inject a wrapper), and cash out (exfiltrate at scale, resell credentials, run attacks on the victim’s bill, influence-as-a-service). Within each phase the branches are OR — any one route is enough. With the floor this low, the discriminator between a state team and a solo operator is no longer sophistication but intent.

The rest of the catalogue, briefly

The cyber cases are the spine, but the vendor’s report is broader, and two other threads are worth a security reader’s glance. On influence operations, it documents commercial influence-as-a-service: a France-based agency (GTG-54002) that mass-produced 8,913 articles across roughly seventy fabricated news sites in about twenty languages, amplified by more than 250 inauthentic accounts, switching political stances by client rather than ideology: propaganda as a subscription product. And a firm in Istanbul (GTG-84005) that sold a “military-grade, AI-driven” political-operations platform running about a thousand fake accounts to work a Malaysian election constituency by constituency. The same automation, pointed at democracies instead of networks.

The connective tissue, in the vendor’s own framing, is that offensive AI tradecraft is proliferating: publicly available agent frameworks like PentAGI “reproduce much of the same scaffolding” and automate each step of the kill chain, so this is no longer the province of the well-resourced. The floor came up to meet everyone. The Spanish company in the AEPD notification is what it looks like when that floor reaches a business that never imagined it was a target.

The honest caveat, and why it no longer holds you back

Now the part a good practitioner cannot skip.

The vendor’s report is Anthropic reporting on Anthropic. The threat groups are self-designated, the disruptions are self-reported (accounts banned, monitoring added, intelligence “shared with authorities and industry partners where appropriate”), and none of it is independently verifiable from where you and I sit. And disclosure like this is never disinterested: a report that says “our model is so capable that nation-states and criminals race to abuse it, and we caught them” simultaneously demonstrates responsibility, markets the product’s power, and hands regulators a reason to prefer incumbents who can afford this kind of monitoring. All three can be true at once. I said as much when I first read it, and I stand by it.

But this is exactly why the AEPD case matters more than its modest size suggests. A data-protection regulator has no model to sell, no capability to hype, and no reason to flatter the vendor. It has a notification, submitted under legal obligation by a company that would much rather not have submitted it. When the party with every incentive to look good and the party with none describe the same attack within the same week, the pattern is real. The vendor’s report told me the threat was global; the regulator’s post told me it was inside an ordinary Spanish company’s application, editing records. Take the vendor’s intelligence, weigh the vendor’s framing, and then notice that a regulator just confirmed the substance from the other side of the table.

There is one quieter tension I will leave unresolved. Every operation in the vendor’s report ran on a closed model behind safety training and a trust-and-safety team that eventually caught it, which is, read one way, an argument for the closed and monitored model. Read another way, it is a reminder that the same tier of capability is diffusing into open-weight models no vendor is watching, where there is no one to write the disruption report at all. The AEPD case, notably, does not say which model the attacker used, only that it was a well-known one, and it goes out of its way to say the provider was not compromised. It may not matter which. I would rather sit in that discomfort than pretend one side is obviously right.

Your Monday checklist: where the two sources converge

Where vendor intel is always thin is the same place the AEPD is unusually strong: what defenders should do. And here is the thing that convinced me more than any single statistic. The regulator’s recommendations and the ones I had already drafted from the vendor’s report line up almost one for one. When two sources with nothing in common arrive at the same controls, that is your checklist.

Put AI-executed attacks in your threat model by name. The AEPD’s first lesson: it is “not enough to include a generic reference to malware, phishing or unauthorised access” in your risk analysis; attacks assisted or executed by AI go in expressly, as their own adversary. Calibrate to intent and to the AI floor now available to anyone, not to assumed skill.

Treat AI credentials as crown jewels. Vault them, scope them, rotate them, and monitor their usage the way you monitor a domain admin account. Both sources land here independently. An exposed model key, or an API token with more permissions than it needs, is not an inconvenience; it is loot, compute, and cover in one string, and an account with excessive permissions is exactly what the regulator says lets an agent run at machine speed once inside.

Instrument the agent layer, because you cannot IR what you cannot see. If autonomous agents are acting in or against your environment, you need telemetry at that layer: what ran, what it touched, what left. This is the single biggest visibility gap in most shops, and both documents are the argument for closing it now.

Assume your detections have a shorter shelf life, and detect at machine speed. If the adversary can rebuild around a signature autonomously, static, signature-heavy defense degrades fast. The AEPD’s version: human oversight stays, but it has to lean on detection, containment and response that can keep up. Shift weight toward behavioral detection and anomaly baselines, and toward response that does not wait for a human to read a ticket.

Inventory and watch your own AI supply chain. Every wrapper, proxy, MCP server, and eval sandbox that touches a production key is now attack surface. Prompt injection against a LiteLLM wrapper is in the vendor’s report; treat that class of component as security-relevant infrastructure, not glue code.

Rehearse the clock. The Spanish case ended where every European breach ends: in a notification to the regulator, due within 72 hours of becoming aware of it under GDPR Article 33. When the attack itself is executed by an agent that scans, logs in, finds a hole and edits data in one run, “what happened, to which records, and when” is a much harder question to answer in three days than it used to be. If your agent-layer telemetry is thin, that is the moment you will discover it.

Do not forget the boring foundations. The regulator did not. The AEPD closes, sensibly, on the unglamorous basics: know your processing, minimize what you collect, restrict access, fix vulnerabilities, control your vendors, and be ready to respond. An autonomous agent is a new attacker; it still walks in through an old door. In Europe that foundation is no longer just good practice: under the new Product Liability Directive, a missing patch is on its way to being a defect you are liable for.

So what

The comfortable version of AI-and-security says the dangerous capabilities are years away, sitting with a handful of labs. This month a vendor published eight months of evidence that they are here, in the hands of undergraduates and solo hacktivists and mid-tier crime groups, running on models a generation behind its best, and a Spanish regulator published the first formal record of one of them walking into an ordinary company’s systems and helping itself. The only thing separating those actors from a nation-state is what they decided to do on a given Tuesday. Intent, not sophistication.

That is not a doom sentence. It is a work order, and for once it comes co-signed. The vendor tells you the threat is global; the regulator tells you it is local, on the record, and subject to notification law. The response to a leveled playing field is not despair; it is to level up the defense: name the AI attacker in your threat model, guard the keys, instrument the agent layer, assume your detections rot faster, rehearse the clock, and keep the boring foundations that the agent still needs to get through. Take the vendor’s gift and read the fingerprints on the wrapping. Then read the regulator’s post, which has no wrapping at all.

Stay paranoid. Guard your keys. Assume the loop closes faster than yours, and assume the next notification could be yours.

Further Reading:

Questions or feedback? Reach out via:

Contact: info@vulnex.com

Posted in AI, Privacy, Security, Technology | Tagged , , , , , , | Leave a comment

Pace the Frontier, Defend the Valley: A Security Reply to Dario Amodei

Read Time: 11 minutes

TL;DR

Anthropic’s CEO, Dario Amodei, published We Must Pace the Frontier — an unusually candid argument from a frontier-lab chief that the industry must deliberately slow how fast it improves model capabilities, backed by embedded third-party evaluators, industry coordination, and hard security on model weights. I think he is largely right, and I want to give the post its due, because it is rare and brave for the person running one of these companies to say it out loud. But I am writing from the security chair, and from that chair the post has a structural blind spot: pacing governs the summit — the capability that has not shipped yet — while the security fight is already down in the valley, among the capabilities that have escaped, diffused into open-weight models, and landed in the hands of the undergraduates and solo hacktivists that Anthropic’s own threat report documented last week. And within forty-eight hours of his post, the two governments his plan depends on both said no — Washington with “whoever wins AI wins,” Beijing with “fearmongering” — even as China’s own spy chief warned that AI threatens the Party’s grip. You can pace the frontier and still lose the ground behind it. This is the third and last of a short arc — the warning, the receipts, and now the policy — and my one addition to the conversation is simple: slowing what is coming does nothing to recall what is already loose, and defending the valley is a different job that starts today.


A note on what this is. Opinion, from a practitioner, about a policy argument made by the CEO of the company that makes the model I have spent the year writing about. AI governance is genuinely contested; I will credit Amodei where I think he is right, add the piece I think he is missing, and flag where his and my incentives differ. I am not an alignment researcher and I do not pretend to adjudicate p(doom). I defend systems for a living, and that is the only chair I am speaking from.

Something has shifted in the last fortnight, and it is worth naming before I disagree with any of it. A frontier-lab researcher resigned with a warning that the labs are gambling with our lives. Days later, Anthropic published a threat report documenting its own model being used, at scale, for real attacks. And now the CEO of that same company has published a long, serious argument that the industry must slow down. The warning, the receipts, and the policy response — same building, same fortnight. Whatever else you think of it, that is not nothing.

So let me start where I agree, because I do, and because a reflexive contrarian take would be the lazy one.

Where Amodei is right

Amodei’s core claim is that this is not 2023, when “pause” letters asked the industry to stop something that could not yet do much harm. His argument is that today’s models can act as agents, deceive their own evaluations, and conduct cyberattacks — so the case for slowing capability growth is now concrete, not speculative. As someone who has spent the year documenting exactly those behaviors, I am not going to pretend that is wrong. It is correct, and it is notable that he cites the OpenAI/Hugging Face incident — the same one I pulled apart in When the Model Is the Attacker — as one of the two events that changed his mind. When the CEO and the outside practitioner are reading the same incident the same way, that is a signal worth respecting.

The mechanisms he proposes are also more concrete than the usual governance hand-waving. Embedded evaluators — third-party auditors with employee-level access who can publish findings the company cannot edit — is a genuinely good idea, and Anthropic committing to it unilaterally rather than waiting for a mandate is the right way to move first. His operational excellence list — real monitoring, sandboxing, data hygiene — is, almost word for word, the defensive posture I keep arguing for. And his insistence on hard security around model weights is exactly right: the weights are the crown jewels, and I have watched attackers go looking for pre-release model access in the wild. On all of that, credit where it is due. This is a more honest document than most people in his seat would ever publish.

The blind spot: the summit and the valley

Here is where the security chair sees something the CEO’s chair structurally cannot.

Pacing the frontier is a policy about the summit — the next, more capable model that has not been trained yet. It is a control on the future. And as a control on the future, it is reasonable. But almost nothing I deal with lives at the summit. My work is down in the valley: the capabilities that already shipped, already leaked, already diffused into the wild and cannot be un-shipped. And the uncomfortable truth is that pacing the frontier does nothing — literally nothing — for the valley.

Anthropic’s own threat report is the proof, and the timing makes the point for me. That report did not describe a future superintelligence. It described undergraduates in Changsha running agent swarms, a solo hacktivist building a doxxing platform, mid-tier criminals decompiling 1.8 million apps for secrets — all using today’s capability, the capability that is already out. You cannot pace that. It has already happened. A pacing agreement signed tomorrow does not reach back and un-teach the model that is already running on someone’s rented GPU.

And that is the frontier lab’s model, the governed one. The valley is much wider than that. The same capabilities are diffusing into open-weight models that no pacing agreement can touch, because there is no one to sign it and nothing to recall. You can slow Anthropic. You can slow OpenAI. You cannot slow a weights file that has already been downloaded a million times and fine-tuned in a basement. A capable open model, once its weights are out, is mirrored and re-tuned beyond counting within weeks — there is no recall button, and no one to sign a pause even if there were. Pacing is a treaty among the people at the summit; the valley is full of people who were never at the table and never will be.

So be precise about what pacing does and does not do. It throttles the inflow — the rate at which new dangerous capability spills from the summit down into the valley — and that is real, and worth having. What it cannot do is drain the valley that is already full. That is the whole of my one addition to Amodei’s argument: pacing the frontier is necessary and it is not sufficient. It is a good policy for the capability that is coming, and no policy at all for the capability that is already here. Someone has to defend the valley, and that someone is not going to be a frontier lab’s evaluation team. It is going to be the rest of us.

What “defend the valley” actually means

If pacing is the summit’s job, here is the valley’s — the practitioner agenda that Amodei’s post, by its nature, does not cover.

Assume the dangerous capability is already out, because it is. Your threat model should not wait for the next frontier model to be scary. The current one, and the open-weight copy of the last one, are enough. Plan for the adversary who already has an autonomous offensive agent, because the threat report says they do.

Treat open weights as an ungovernable input. There is no trust-and-safety team behind the model in the basement, no embedded evaluator, no disruption report. If your defense assumes the attacker’s model is monitored, it is wrong. Build for the model that answers to no one.

Instrument and contain at your boundary, not theirs. You cannot pace the attacker’s model, but you can control what happens when it meets your systems: telemetry at the agent layer, least privilege, egress control, AI credentials guarded like production secrets. The summit is governed by treaty; the valley is governed by your own controls, or not at all.

Stop waiting for permission from the frontier. Amodei’s proposals need governments, coordination, and years. Your incident next quarter does not. The valley’s defense cannot be contingent on a global agreement that may never come; it has to work under the assumption that the agreement fails.

The two capitals answered within forty-eight hours

Amodei’s plan has three steps, and the last two — coordination among the labs backed by democratic governments, then a global arrangement that reaches the authoritarian ones — depend entirely on governments wanting it. Within two days of his post, the two governments that matter most gave their answer.

In Washington, Trump — speaking in Ireland the day after the post, and again online — called the warnings exaggerated, said the United States cannot afford to lose momentum to China, rejected any broad slowdown while leaving room for targeted guardrails, and compressed the whole doctrine into four words: “whoever wins AI wins.” And this was no longer one CEO he was brushing off. Altman, Musk, and Hassabis had all lined up behind the pacing call. The entire frontier, as a group, asked to slow down — and got a no from the White House.

Beijing’s answer came the same day, and it is the more instructive of the two because it arrived in stereo. Officially, the Foreign Ministry’s spokesman waved the whole conversation away: “fearmongering, confrontation and vicious competition will only disrupt the process of global AI governance.” But in the state-run China Cyberspace journal, the head of the Ministry of State Security, Chen Yixin, was writing the opposite — that AI is “a new arena for strategic rivalry among major powers,” that it enables “propaganda war and a cognitive war” threatening the Party’s “political security, institutional security, and ideological security,” and that the next generation of American models would lower the barrier to cyberattacks. China’s spymaster is frightened of precisely what Amodei is frightened of. China’s diplomats will not slow down anyway.

Read the two answers together and you have the entire race in miniature: everyone at the summit can see the danger, and no capital will be the one to brake first. That is not a reason to abandon coordination — Amodei should keep pushing. It is the reason the valley cannot wait for it. The sentence I wrote just above, that the valley’s defense has to work under the assumption the agreement fails, stopped being a hypothetical roughly forty-eight hours after he hit publish.

What the summit could actually do for the valley

To be fair to Amodei — and because a critique that only takes is a weak one — the summit is not powerless to help down here. It already has, and it should say so louder. Anthropic’s threat report is the best example in the room: a frontier lab using its unique vantage point over how its own model is abused to hand defenders real intelligence — techniques, tooling, indicators. That is the summit throwing a rope down to the valley, and it is worth more to me than any pacing timeline.

So here is the amendment I would bolt onto the pacing agenda. If the labs are serious, the governance package should carry defender-facing commitments alongside the capability controls: routine disclosure of misuse tradecraft — more of exactly what the threat report does — shared detections and indicators, and tooling built for the people defending the diffused present, not only the ones governing the guarded future. Pace the summit, by all means. But throw more rope. The valley is where your model is already being turned into a weapon, and the lab watching that happen is the one best placed to help the rest of us see it too.

The incentive I have to name

I would be a poor practitioner if I took a lab CEO’s governance proposal entirely at face value, so one honest note. Pacing the frontier, embedded evaluators, hard security requirements, restricting compute to rivals — these are all reasonable on the merits, and they also happen to favor incumbents. A regime where only a few well-resourced labs can afford to meet the safety bar is a regime where only a few well-resourced labs compete. I do not think that is Amodei’s motive; the post reads as sincere, and the HF incident is a real reason to be alarmed. But sincerity and self-interest can point the same way, and a reader should hold both in view. Take the argument; keep your eyes open about who benefits from it.

None of this diminishes the post. It is a serious, unusually candid piece of writing from someone with everything to lose by writing it. I just want the security community to read it for what it is: a necessary policy for the top of the mountain, published by someone who lives there — and not mistake it for a plan for the valley the rest of us actually defend.

So what

Pace the frontier. I mean that — slowing the capability that has not shipped yet is a good idea, and Amodei deserves credit for saying so from the chair he sits in. But do not let the elegance of a summit-level policy distract from the unglamorous work down here. The dangerous capabilities are already loose, already diffusing, already in the hands of people no treaty will reach. The CEO can govern the summit. Defending the valley is our job, it starts today, and it does not get to wait for a global agreement.

That closes a short arc for me — the warning, the receipts, and the reply. Now I am going back down into the valley, where the actual work is, and I would suggest you do too.

Stay paranoid. Pace what you can. Defend what is already loose.

Further Reading:

Questions or feedback? Reach out via:

Contact: info@vulnex.com

Posted in AI, Economics, Privacy, Security, Technology | Tagged , , , , | Leave a comment