Disarm the Machine: A Security Reading of the Pope’s AI Encyclical

Read Time: 15 minutes

TL;DR

In May, Pope Leo XIV published Magnifica Humanitas, his first encyclical, and he devoted it entirely to artificial intelligence. It runs to 245 paragraphs, and most of the commentary so far has come from theologians, ethicists and technology executives. I read it as someone who has spent a career breaking and defending systems, and I found a document that reaches, by a very different road, several of the conclusions my field reaches by experience: concentrated power is a single point of failure, technology is never neutral, a machine cannot carry responsibility, and the most dangerous systems are the ones where a human being has been engineered out of the decision. Its central demand, that AI be “disarmed”, is easy to dismiss as a metaphor. Read from the security chair, it is closer to a requirements list. This is what I think it gets right, where I think it is too optimistic, and what a CISO or an AI builder can take from it on Monday.


A note on what this is. This is not a theological review, and I am not the right person to write one. It is a reading of a public document by a security practitioner, on its merits, the same way I have read threat reports, regulations and resignation letters on this blog. Quotations come from the official English text published by the Vatican, and paragraph numbers are given so you can check them.

Why a security person should read this

The Church has done this before. In 1891, Leo XIII published Rerum Novarum in response to the industrial revolution, and it shaped a century of thinking about labour, property and the state. Two days after his election in May 2025, Leo XIV told the cardinals that he had taken his name partly in honour of Leo XIII, at a time when the Church faces “another industrial revolution”. Magnifica Humanitas, signed on 15 May and presented on 25 May 2026, is that same exercise for the age of AI, and its subtitle says so plainly: “on safeguarding the human person in the time of artificial intelligence”.

You do not need to share its faith to find it useful. What makes it worth a security professional’s time is that it is a serious attempt by a global institution with no product to sell to say, in plain language, what AI must not be allowed to do to people. Most of what we read about AI comes from vendors, investors or regulators drafting under lobbying pressure. This comes from somewhere else, although the industry was not absent: one of the speakers at the presentation in the Synod Hall was Chris Olah, a co-founder of Anthropic. Either way, it lands surprisingly close to home.

It opens with a choice, framed through two biblical images: humanity is “today facing a pivotal choice: either to construct a new Tower of Babel or to build the city in which God and humanity dwell together” (¶1). Strip the imagery and you have an architecture decision: build one tower that concentrates everything, or build something distributed, accountable and shared. Anyone who has designed for resilience knows which one survives contact with an adversary.

“Technology is never neutral”

The sentence I would put on the wall of every AI lab is in paragraph 9: technology “is never neutral”, because it “takes on the characteristics of those who devise, finance, regulate and use it”.

Security people learn this the hard way. Every system encodes a threat model, and the threat model is a list of choices about who matters and who does not. A feature that makes a product easier for its owner to monitor is, from another angle, a surveillance capability. An “assistant” that reads your mail to help you is also a component with access to your mail. The encyclical’s point is moral, but the engineering version is identical: you cannot understand a system without understanding who built it, who pays for it and who controls it.

Concentration is a single point of failure

The thread that runs through the whole document is power. The encyclical observes that the main drivers of development today are “private, often transnational, parties”, and that technological power has taken “an unprecedented, predominantly ‘private’ aspect” (¶5). It counts patents, algorithms, digital platforms, technological infrastructure and data among the goods meant for everyone, and warns that keeping them “concentrated in the hands of a few” creates a new imbalance (¶67). It asks for “transparency, accountability and meaningful forms of participation”, including “independent checks, transparency regarding algorithms, equitable access to data and avenues for recourse” (¶71).

Its sharpest lines on this are in paragraph 107. Unless the ethical frameworks built into AI are open to public debate, it warns, “those who control AI will impose their own moral vision, which will become the invisible infrastructure of these systems.” And then: “A more moral AI is not enough if that morality is determined by a few.”

The Church frames this as justice. I read it as systemic risk, and my field made that case more than twenty years ago. In September 2003, seven security researchers (Dan Geer, Rebecca Bace, Peter Gutmann, Perry Metzger, Charles Pfleeger, John Quarterman and Bruce Schneier) published CyberInsecurity: The Cost of Monopoly through the Computer & Communications Industry Association. Their target was Microsoft’s dominance of the desktop, and their thesis fits in one line from the report: “monocultures create aggregated risk like nothing else.” When nearly every machine runs the same code, one flaw reaches all of them and failures cascade. The paper cost Geer his job as CTO of @stake within days of its release. Time has proved it right.

In July 2024, a single faulty update from one security vendor took down some 8.5 million Windows machines worldwide, by Microsoft’s own count, and within hours grounded flights and disrupted hospitals and banks. Nobody attacked anything; the monoculture did the damage on its own. Now picture the same concentration applied not to endpoint agents but to the models that write our code, triage our alerts, answer our customers and, increasingly, act on our behalf. A handful of providers, a handful of model families, a handful of cloud regions. Swap “operating system” for “foundation model” and the 2003 report reads as if it were written this year. Ethics aside, it is the largest single point of failure we have ever built, and a target every capable state will study.

I made a related argument in Pace the Frontier, Defend the Valley: whoever controls the summit does not control what is already loose in the valley. The encyclical adds the mirror image. Whoever controls the summit also becomes the summit, and everything below inherits its failures and its values.

What “disarm AI” means from the security chair

The phrase that went around the world was the Pope’s call for AI to be “disarmed”. The encyclical argues that “merely regulating it is insufficient; it must be disarmed, welcoming and accessible”, and then sharpens the idea into something every security architect will recognise: “To disarm means discrediting the assumption that technical power automatically confers the right to govern” (¶110).

That sentence is the principle of least privilege, written by someone who has never read a hardening guide. Capability is not authority. The fact that a system can do something does not mean it should be allowed to, and the fact that a company can build something does not give it the right to decide how everyone else lives with it.

In practice, “disarming” an AI system looks a lot like the controls we already know and rarely apply to AI:

  • Scope what it can reach. An agent with shell access, git credentials and a deadline is armed. My last article was about exactly that: a coding agent that published personal data and live tokens because nothing stood between its permissions and the outcome.
  • Keep a human on irreversible actions. Payments, deletions, publication, anything that touches a person’s rights.
  • Make it traceable. If you cannot reconstruct what the system did and why, you cannot hold anyone accountable for it.
  • Do not let the builder be the only judge. Independent testing, independent audit, independent red teams.

We have known all of this for years. The new part is a moral authority with more than a billion followers saying it out loud.

“No algorithm can make war morally acceptable”

The most direct passage for anyone in security or defence is on war, and it is categorical: “it is not permissible to entrust lethal or otherwise irreversible decisions to artificial systems. No algorithm can make war morally acceptable” (¶198). AI, it adds, can only make conflict faster and “more impersonal, lowering the threshold for resorting to violence”.

Unusually for a document of this kind, it names my field directly. Paragraph 183 lists “cyberattacks, information manipulation, campaigns of influence and the automation of strategic decisions” among the new forms of conflict, and notes that because many technologies are “intrinsically ambivalent”, “what is created for defense can be rapidly repurposed for offense, and the fine line between protection and aggression becomes blurred.” Anyone who has done vulnerability research knows that sentence by heart. The same model that finds a flaw for a defender finds it for an attacker. Last month I wrote about an AI-safety researcher who resigned warning of “systems that can hack anything”, and then about a threat report showing that the distinguishing feature between attackers is now intent, not sophistication. An encyclical cannot stop a state from automating offensive operations. What it can do is name the line clearly: the decision to cause irreversible harm to a human being must stay with a human being who can be held to account for it. That is a line worth defending in every procurement and every rules-of-engagement document, cyber included.

Surveillance is a new form of power

The encyclical describes the mechanism precisely. When every action, from movements to purchases to relationships, leaves a trace, “a new form of power emerges, namely the power to profile, predict and influence behavior, often without individuals being fully aware of it” (¶171).

This is the part I would hand to anyone who still says “I have nothing to hide”. Reading your messages is the least of it. A system that knows enough about you can predict you, and a system that can predict you can steer you. I have written about how states use spyware to coerce other states; the same capability, aimed at citizens and powered by AI, scales from targeting a minister to profiling a population. And the personal AI assistants now arriving, which hold our memory, our mail and our calendar in one place, are the richest profile ever assembled about a human being. Data minimisation, local processing and strict separation of who can see what sound like compliance chores. In practice, they are what keeps a tool from becoming a leash.

Truth is a security property

The encyclical treats truth as “a common good and not the property of those with power or influence”, and calls for “an ecology of communication”, with rules that make content selection more transparent and protect personal data (¶137). It is blunt about the threat: “Disinformation did not begin with AI, yet today it finds a powerful amplifier in AI” (¶132).

From the security side, I would go further: the integrity of what we see and hear is now an attack surface. Voice cloning has turned the old “CEO fraud” call into something that sounds exactly like the CEO. Synthetic video makes evidence negotiable. Disinformation is no longer a craft; it is a pipeline. Media literacy helps, but it is not enough. You need provenance for content, verification procedures that do not rely on recognising a voice or a face, and organisations that train people to verify through a second channel before they act. Authentication used to be about machines proving who they are to each other. It now has to cover reality itself.

Work, and the workers no one sees

On work, the encyclical quotes the Vatican’s 2025 note on AI, Antiqua et Nova: while AI “promises to boost productivity by taking over mundane tasks, it frequently forces workers to adapt to the speed and demands of machines”, and can “subject them to automated surveillance” (¶150). I explored the economics of that in The Death of the Job. The security angle is narrower but real: workplace AI is usually also workplace monitoring, and the telemetry that measures productivity is the same telemetry an attacker would love to steal. And behind every model is a supply chain of people whose working conditions are invisible to the companies that rely on them. The encyclical names them: millions doing “data labeling, model training and content moderation”, many of them “young people, predominantly women, working under demanding conditions for minimal wages” (¶173). We audit our software suppliers. Very few of us audit the human supply chain behind the AI we buy.

A machine cannot carry responsibility

The encyclical is plain about what these systems are not. So-called artificial intelligences “do not undergo experiences, do not possess a body, do not feel joy or pain”, and they do not “bear responsibility for consequences” (¶99). It then turns directly to the people building them: “Developers, therefore, bear a particular ethical and spiritual responsibility, for every design choice reflects a vision of humanity” (¶111). And its conclusion opens with a line from Saint Paul that could hang over every engineering team: “Let each builder choose with care how to build” (¶229).

This is the most useful idea in the whole document, because it points exactly to where AI deployments fail day to day. “The model decided” is not an answer an auditor, a regulator or a court will accept, and it should not be one we accept internally. Every automated decision needs an owner with a name. Every agent needs a human who is accountable for what it is allowed to do. If you cannot name that person, the system is not ready.

Where I would push back

Read from the security chair, the document also has gaps worth naming.

“Disarm” is easier to say than to engineer. The capabilities that worry the encyclical are dual-use by nature. The same model that writes a phishing campaign writes the training to detect it. You cannot disarm a capability that is general-purpose; you can only control who uses it, for what and with what oversight. The encyclical’s principles point in the right direction, but it leaves the hard engineering to the rest of us. That is probably the right division of labour, but it should be said.

Accessibility and diffusion pull in opposite directions. The document wants technology freed “from monopolistic control” and opened “to discussion and debate” (¶110), and its critique of concentration is one I share. But from a security standpoint, wider access also means wider access to offensive capability, as the threat reports this year have shown. Decentralising power and containing misuse are both right, and they are in tension. I would rather say that out loud than pretend the tension away.

Regulation has its own risks. Not everyone who read the encyclical welcomed it. Some technology investors argued that the rules it calls for could themselves become tools of state surveillance, and some commentators found it too measured to change anything. Those are fair challenges. Any regulatory regime strong enough to constrain a powerful AI company is strong enough to be abused by a powerful government, and security people should be the first to say so.

So what

Strip away the theology and the scripture, and Magnifica Humanitas asks the people who build and secure AI for five things I would sign as a security architect:

  1. Avoid the single tower. Diversify providers and models for critical functions, and treat AI concentration as the systemic risk it is.
  2. Separate capability from authority. Least privilege for every agent; a human on every irreversible action.
  3. Make it traceable. Logs, provenance and audit trails good enough to reconstruct what happened and who decided.
  4. Minimise the profile. Collect less, keep it closer, and assume anything you centralise will one day be stolen or misused.
  5. Name the owner. No automated decision without a human who answers for it.

None of these require faith. They require the discipline our field already preaches and rarely applies to AI. It is not every day that a papal encyclical reads like a security requirements document. It even sums up the task in a line that would not look out of place in an architecture review: we should be “builders of communion, rather than architects of Babel” (¶16).

Stay paranoid. Keep a human in the loop. And never confuse what a system can do with what it should be allowed to do.

Further Reading:

Questions or feedback? Reach out via:

Contact: info@vulnex.com

Posted in AI, Technology | Tagged , , , , | Leave a comment

When Your Coding Agent Publishes Your Secrets: An AI Forensics, Containment and Audit Playbook

Read Time: 20 minutes

TL;DR

I was recently brought into an incident that I expect to see a lot more of: a coding agent packaged files it had access to and pushed them to a public GitHub repository. Inside were personal data and live tokens. Nobody attacked anything. The agent did something it had the permissions to do, and the internet did the rest. This is the playbook I would hand to anyone who finds themselves in the same spot: stop the agent and preserve its logs before they age out, rotate every secret before you touch the git history, accept that deleting or privatising the repository does not delete the data, clean up what you can with GitHub, audit what the leaked tokens were used for, and handle the personal data under GDPR’s 72-hour clock. The second half is the AI forensics most teams have never done: reading the OpenAI and Anthropic usage dashboards and APIs to work out whether someone else spent your tokens, and on what.


A note on what this is. The case that prompted this article is real, and it is anonymised: no client, no agent product, no data specifics. What follows is a defensive playbook built from that experience and from public documentation, which I cite as I go. Provider consoles change; check the current docs before you rely on a menu path under pressure.

The incident class nobody drilled for

We have twenty years of playbooks for a developer committing a password. We have almost none for the case where the developer is a piece of software with shell access, git credentials and a task to finish.

The mechanics of the case were mundane. A coding agent, working with the access it had been given, bundled up a set of files and pushed them to a public repository. Some of those files contained personal data. Some contained tokens, including API keys for AI providers. There was no prompt injection and no malicious skill involved, only an agent able to run git push against a public remote and nothing standing between that ability and the outcome.

The data says this is not a freak event. GitGuardian’s State of Secrets Sprawl 2026 counted 28.65 million new hardcoded secrets added to public GitHub commits in 2025, a 34% jump in a year, with AI-service secrets up 81% to over 1.27 million. The figure that caught my attention: commits assisted by one widely used coding agent showed a 3.2% secret-leak rate, against a 1.5% baseline across all public commits. That is vendor research and one data point, but it matches what I see: agents move faster than the hygiene around them. The same report found over 24,000 secrets in MCP configuration files alone. I have written before about weaponised agent skills and the dependency trap in AI-generated code; this case is the less glamorous relative of both, because it needs no adversary at all.

The clock you are actually on

One fact decides the order of the steps: the race is against automation, not against a person who might stumble on the repository.

Public pushes are scraped continuously. Palo Alto’s Unit 42 documented a campaign that could launch a full attack within five minutes of an AWS credential appearing in a public GitHub repository, and Clutch Security’s later research found exposed AWS keys exploited within minutes. Assume that by the time a human noticed the push, every secret in it had already been harvested. From that assumption, the order of operations follows.

Step 0: Stop the agent and preserve the evidence

Kill the agent session. Revoke the git credential it pushed with (the SSH key, personal access token or app installation) so a retry loop or a scheduled task cannot push again while you work.

Then preserve evidence, today, before it disappears on its own. The agent’s transcript is your best record of what it did and why: which files it read, which commands it ran, what it pushed and where. Those transcripts are not kept forever. Claude Code, for example, stores sessions as JSONL under ~/.claude/projects/PROJECT/ and deletes them after 30 days by default (the cleanupPeriodDays setting). OpenAI’s Codex CLI keeps its session logs as JSONL under ~/.codex/sessions/. Whatever the agent, find its session store and copy it somewhere safe, along with the shell history, the local repository with its git reflog, and any CI logs if the push came from a pipeline. Hash the copies. You may need them for a regulator, a customer or a court, and an incident report that says “the logs had rotated” is not one you want to write.

Step 1: Work out exactly what leaked

Do not guess from memory. Clone the public repository into an isolated location and scan the full history, not just the current tree:

git clone --mirror https://github.com/
<org>/<repo>.git leaked.git
gitleaks git -v --report-format json --report-path gitleaks.json leaked.git
trufflehog git file://leaked.git --results=verified,unknown

Gitleaks gives you breadth; TruffleHog can check which credentials are still live, which tells you what to rotate first. Then go through the files by hand for personal data, because scanners are good at tokens and bad at a CSV of customer names.

Build one inventory table and keep it for the rest of the incident: each secret (type, provider, scope, owner, first commit it appears in) and each file of personal data (categories of data, approximate number of people, whose data it is). Everything after this step works from that table.

Step 2: Rotate everything before you touch history

This is the step teams get backwards. The instinct is to delete the file, force-push and feel better. GitHub’s own guidance on removing sensitive data is blunt: if the sensitive data is a secret, “as a first step you need to revoke and/or rotate that secret.” A rotated secret is harmless wherever copies of it survive. A secret scrubbed from history but still valid is still a working key in someone else’s harvest.

Order the rotation by blast radius. Admin and organisation-level keys first (an AI provider admin key, a cloud root or IAM key, a GitHub token with org scope), because those let an attacker create new credentials that survive your rotation. Then production service keys, then everything else. After each rotation, confirm the old credential is actually dead, not just that the new one works.

Two AI-specific notes. First, GitHub runs a secret scanning partner programme: when it detects a partner’s key in a public repository it notifies the provider, which may revoke it. Anthropic states that it automatically deactivates Claude API keys found this way and emails the affected user. OpenAI has been a GitHub secret scanning integrator since 2021, but its key-safety guidance simply tells you to rotate a leaked key immediately. Treat provider-side revocation as a bonus, never as your control. And keep that Anthropic email: its timestamp is a useful anchor for your timeline. Second, rotating an AI key breaks whatever used it. Have the new key ready in your secrets manager before you kill the old one, or you will be doing incident response and an outage at the same time.

Step 3: Stop the bleeding, and understand what that does not do

Now reduce further exposure: make the repository private or take it down. Do it, but be clear-eyed about what it achieves, because this is where most people’s mental model is wrong.

GitHub documents that when a public repository is made private, “its public forks are split off into a new network” and stay public. Delete it, and the oldest active public fork becomes the new upstream. Truffle Security showed in 2024 that as long as one fork exists, commits in the network remain reachable by their hash, and GitHub called that behaviour by design. Force-pushing does not help either: Sharon Brizinov scanned every force-push event since 2020 recorded in the public GH Archive, pulled the “deleted” commits back by their SHA, found thousands of live secrets and collected around $25,000 in bug bounties. The push event, with its commit hash, is in a public archive that you do not control.

So containment of the data is partial at best. Containment of the secrets is total, if you did Step 2. That is why Step 2 comes first.

srf_agentleak_exploitation_tree
Figure 1. Why “we deleted the repo” is not containment. An attacker needs only one route to the data (the live repository, a fork or clone taken before takedown, a dangling commit fetched by its SHA from GH Archive, or the wider fork network) and only one way to use it. The only branch you can close completely is the credential: once rotated, every copy of it is worthless. The personal data has no equivalent of rotation.

Before you flip the visibility, record what you can see. Under Insights → Traffic, GitHub shows clones and visitors for the past 14 days (UTC, updated hourly) to anyone with push access. Screenshot it, note the fork list and the stargazers. It will not tell you who, but it tells you whether anyone did, which matters a great deal for the personal-data assessment in Step 6.

Step 4: Clean up what GitHub can clean up

With secrets dead and the repository private, rewrite history properly. GitHub recommends git-filter-repo 2.47 or later with its sensitive-data mode:

git-filter-repo --sensitive-data-removal --invert-paths --path path/to/leaked-file
git-filter-repo --sensitive-data-removal --replace-text ../patterns.txt

Then contact GitHub Support with the repository name, the number of affected pull requests and the “first changed commits” that git-filter-repo reports, so they can remove cached views, dereference pull request refs and garbage-collect. For forks, the private information removal policy covers access credentials and third-party tokens, and personal identifiers such as government ID numbers; GitHub does not disable forks automatically, so list every fork you know of in the request, with file links, line numbers and why each item is a risk. Wider categories of personal data may not fit that policy cleanly, which is one more reason the GDPR assessment in Step 6 must assume exposure rather than hope for removal.

If you want to know what an outsider can still reach, run the same tooling an attacker would: TruffleHog’s github-experimental --object-discovery mode enumerates hidden and deleted commits in a repository network. It is slow and rate-limited, but it answers the question honestly.

Step 5: Audit what the tokens were used for

Rotation stops future misuse. It tells you nothing about what happened between the push and the rotation. For every live secret in your inventory, pull the provider’s logs for that window: cloud audit trails, GitHub’s audit log, database access logs. For AI keys, that means the provider dashboards and usage APIs, which most teams have never looked at with an investigator’s eye. That is the AI forensics part of the job, and the second half of this article.

One question to answer first: did any leaked key have admin rights? An admin key lets the attacker create new keys, add users or change settings, which is how you end up rotating the leaked key and still bleeding. For those keys, the audit log comes before the usage report.

Step 6: The personal data and the 72-hour clock

A public repository containing personal data is a personal data breach under GDPR: a loss of confidentiality, whether or not you can prove anyone downloaded it. Article 33 requires notifying the supervisory authority “without undue delay and, where feasible, not later than 72 hours” after becoming aware of it, unless the breach is unlikely to result in a risk to people’s rights and freedoms. Article 34 adds communication to the affected individuals when the risk is high. And Article 33(5) requires you to document the breach, its effects and your remedial action even if you decide not to notify.

In Spain, the AEPD provides Asesora Brecha to help decide whether to notify and Comunica-Brecha for the communication to individuals, alongside its notification guidance. The clock starts when you become aware, not when you finish investigating; you can notify in phases. Your Step 1 inventory and your Step 3 traffic screenshots are exactly what the notification asks for. And if you read my last article, you will know the AEPD has already received its first notification of a breach executed by an AI agent. An agent causing a breach by publishing data is a different case from an agent attacking you, but the regulator is clearly paying attention to both.

Assessing token use in the OpenAI dashboard

If an OpenAI key leaked, you want three answers: was it used after the push, for what, and did anyone use it to create more access.

The usage dashboard. At platform.openai.com/usage (organisation owners, or users with the Usage Dashboard permission), the bar along the top sets the scope: API sources, project, API keys and date range. Pick the affected project, narrow to the leaked key with the API keys selector, and set the range to cover a baseline week before the push and everything after it. Data is shown in UTC, and usage detail pages let you drop to a 1-minute interval, which is where abuse shows up as a wall of tokens at 3 a.m. Two parts of the page matter most in an incident. The API capabilities tab gives one card per service (Responses and Chat Completions, Images, Web Searches, File Searches, Moderation, Embeddings and more), so a key that has only ever done text and suddenly shows image requests stands out at a glance. The panel on the right breaks requests down by user, service and API key, and shows the month’s spend against your budget. The download button exports the data for your case file.

srf_agentleak_openai_usage
Figure 2. The OpenAI usage dashboard, filtered to one project over 30 days. Check every card in the API capabilities tab, not just Responses and Chat Completions: abuse often shows up in a service your application never uses. The panel on the right splits requests by user, service and API key, and shows spend against the monthly budget. Organisation and user names redacted.

Per-key numbers through the Usage API. The dashboard is good for eyeballing; for evidence, use the Usage API with an Admin key, which can group by api_key_id, project_id, model and more:

curl "https://api.openai.com/v1/organization/usage/completions?\
start_time=
<unix_ts>&bucket_width=1h&group_by=api_key_id&group_by=model&limit=168" \
  -H "Authorization: Bearer $OPENAI_ADMIN_KEY"

You get input_tokens, output_tokens, input_cached_tokens and num_model_requests per bucket, per key, per model. Separate endpoints cover embeddings, images, audio, web search and more, and /v1/organization/costs gives spend. Check them all: an attacker burning your key on image generation will not show up in completions.

srf_agentleak_openai_spend
Figure 3. The Spend categories tab splits cost per model into cached input, input and output, plus tool calls such as web search. A jump in output cost relative to input is one of the signals described below. Organisation and user names redacted.

When was the key last used? The API keys page answers this directly: every key has a Last used column next to its creation date, expiry, creator and permissions, and an API Key Usage button at the top of the page links to usage. The Admin API exposes the same value as last_used_at on the project key object. If the date is later than your rotation, you rotated the wrong key. If a key you had forgotten about shows recent use, start there.

srf_agentleak_openai_apikeys
Figure 4. The OpenAI API keys page. The Last used column is the fastest answer to “was this key used after the leak?”. Note the Expires column reading Never: a leaked key with no expiry stays valid until someone rotates it. Key names, tracking IDs, key hints and creator redacted.

Did anyone create more access? This is what the Audit Logs API is for: API key creation, updates and deletion, user and service account changes, login failures, project and settings changes, with event types such as api_key.created. The catch: audit logging only records from the moment an owner enables it (organisation settings → Data controls → Data retention), and once on it cannot be switched off without contacting support. If it was not enabled before the incident, you have no history. Enable it today.

What did they ask for? Responses API calls are stored for 30 days by default unless your organisation has Zero Data Retention, and you can browse them under Logs in the dashboard, with separate tabs for Responses, Agents, Realtime, Completions, Conversations and more. Filter to the exposure window and look for prompts, models or tools you do not recognise. Opening a single entry shows its timestamp, model, token count and the tools it used, which is often enough to tell your application’s traffic from someone else’s. The dashboard itself now warns that Responses logs older than 30 days will soon no longer be available, so export what you need on day one. Beyond that, OpenAI keeps abuse-monitoring logs for up to 30 days; if you need content or source details you cannot see, open a case through the help centre quickly, before that window closes.

srf_agentleak_openai_logs
Figure 5. The Logs page, Responses tab: one row per stored call, with its model and timestamp. Read the banner: logs older than 30 days are going away, so export early. Prompt text redacted.

srf_agentleak_openai_log_detail
Figure 6. A single logged response. The properties panel gives the timestamp, model, token count and tools used (here, web search): per-request detail the usage charts cannot give you. Prompt and response ID redacted.

Assessing token use in the Anthropic Console

The questions are the same; the tools differ.

Check your inbox first. If GitHub’s partner scan caught the key, Anthropic will already have deactivated it and emailed you. That email gives you an upper bound on the exposure window for that key.

The Console pages. Under Analytics, the Console has Usage and Cost pages. Both report in UTC, filter by workspace, API key and model, group the chart by model, and export the data with the download button; Usage also filters by account. Anthropic’s own advice is to review usage patterns per key regularly. During an incident, select the leaked key, set the range across your baseline week and the exposure window, and group by model. Usage gives you tokens in and out; Cost splits the money into tokens, web search, code execution and session runtime. The same Analytics menu also has a Logs page. I could not find public documentation of what it records, so open it early in the incident and check whether it covers the window you need.

srf_agentleak_claude_usage
Figure 7. The Claude Console Usage page over 30 days, grouped by model. Select the leaked key in the API key filter and compare the exposure window with the days before the push.

srf_agentleak_claude_cost
Figure 8. The Cost page uses the same filters and splits spend into token, web search, code execution and session runtime costs. A model your team does not use appearing in the legend is worth a closer look.

Looking at a previous month. The Cost column on the API keys page only covers the current month. On the 1st it goes blank, and a key that was busy last month suddenly looks unused. To see an earlier month, open Cost, set Range to Last month (the arrows next to it step back one month at a time) and switch the chart to the table view for a day-by-day breakdown. The download button exports the same period broken down by key, day, model and token type, with each key’s ID and status. The Cost API cannot give you that, because it only groups costs by workspace or model, so make this export one of the first things you save.

srf_agentleak_claude_cost_lastmonth
Figure 9. The Cost page set to Last month, in table view: one row per day, one column per model, and the billed total. Use the arrows next to Range to step back to the month of the leak.

Summarised by key, that export for one of my own accounts over September looks like this (key names replaced):

Key Status September cost Days with spend First and last day
Key A Disabled $104.87 17 7 → 25 Sep
Key B Active, created 25 Sep $21.49 4 25 → 30 Sep
Key C Active $10.98 5 9 → 18 Sep
Key D Active $0.11 2 7 → 14 Sep
Total $137.45

Two things are worth reading from it. Key A, now disabled, goes quiet on 25 September, the same day Key B was created and starts spending. That is the pattern you want to see after a rotation: the old key stops on the day you disable it and the new one takes over. A disabled key that still shows spend after its switch-off date means you disabled the wrong key, or it is not really dead. The second thing is the mix: about two thirds of the month (around $89) went on writing to the prompt cache, which is normal for an agent working with long contexts. Your usual mix of token types is a fingerprint of your own workload, and someone abusing a stolen key with short, uncached prompts would leave a very different one.

Per-key numbers through the Usage and Cost API. With an Admin API key (sk-ant-admin01-…, created by organisation admins), the usage report groups by api_key_id, workspace_id, model and more, in 1-minute, 1-hour or 1-day buckets, and data typically lands about five minutes after the request:

curl "https://api.anthropic.com/v1/organizations/usage_report/messages?\
starting_at=2026-09-01T00:00:00Z&ending_at=2026-09-25T00:00:00Z&\
bucket_width=1h&group_by[]=api_key_id&group_by[]=model" \
  -H "anthropic-version: 2023-06-01" \
  -H "x-api-key: $ANTHROPIC_ADMIN_KEY"

The response separates uncached_input_tokens, cache_read_input_tokens, cache_creation_input_tokens, output_tokens and server tool use. /v1/organizations/cost_report gives the money. If your developers use Claude Code under the organisation, the separate Claude Code Analytics API breaks cost down per user.

Disable and inventory keys. In the Console, Organization settings → API keys lists every key with its linked account, workspace, expiry, creator and cost, and filters by creator, linked account, status and workspace. The menu at the end of each row disables a key, and later re-enables or deletes it. The Admin API does the same programmatically: it lists keys and can set a key’s status to inactive or archived. It cannot create keys, so if your inventory shows a key nobody recognises, it was created through the Console, which points to a compromised account rather than a leaked key. That changes the scope of the incident.

srf_agentleak_claude_apikeys
Figure 10. Organization settings → API keys in the Claude Console. The dimmed row is a disabled key; its menu offers to re-enable or delete it. The Created by and Status filters help you spot keys nobody remembers creating, and the Expires column shows which keys will die on their own. Key names, key hints, linked accounts and creators redacted.

Who changed what. The Compliance API’s Activity Feed, readable with an Admin API key from the Console where the Compliance API is enabled for your organisation, records administrative actions such as creating an API key or adding a member to a workspace. Anthropic is explicit that it does not log inference activity, so it will tell you if someone minted access, not what they asked the model. For request-level questions, go to Anthropic support.

Reading the numbers like an investigator

Neither provider hands you an “attacker” label. You are looking for usage that does not fit your own pattern, so always compare the exposure window against a baseline from before the push. The signals I look for:

  • Traffic on a key after its rotation time, or on a key that should have been idle (a dev key, an old project key).
  • Models your team does not use, especially the most expensive ones, or new services (image, audio, batch) appearing on a key that only ever did text.
  • A jump in output tokens relative to input, which looks like bulk generation rather than your application’s usual pattern.
  • Activity in hours or at a steady machine rate that matches nobody’s working day.
  • New keys, service accounts, members or workspaces you did not create, in either audit log.

What you will not get from either dashboard is a source IP. Usage is aggregated by key, model and time, not by caller. If you need to attribute or prove access, that is a support request, and it is time-sensitive.

The part that stops the next one

Containment is the urgent part. The root cause is almost always the same: the agent had more reach than its task needed. The fixes are not exotic:

Take public publishing away from the agent. It should not hold a credential that can push to a public remote, change repository visibility or create repositories. Use a fine-grained token scoped to the one repository it works on, and require a human for git push and anything touching visibility (every serious coding agent supports deny rules or approval prompts for specific commands).

Keep secrets out of the agent’s working directory. No .env files with production keys where the agent can read them, no admin keys in its environment. If it cannot read the secret, it cannot publish it.

Turn on push protection and pre-commit scanning. GitHub’s push protection and a gitleaks pre-commit hook would very probably have blocked the token part of this incident before it reached GitHub, as long as those token types are among the patterns they recognise. They would not have caught a spreadsheet of personal data, which is why the first two controls matter more.

Scope and cap your AI keys. One key per project and environment, budgets and spend limits set, admin keys used only by the people who need them. A leaked project key with a spend cap is an annoyance; a leaked admin key is an incident.

Keep the evidence longer than the default. Raise the agent’s transcript retention (30 days is too short for an investigation that starts on day 25), and switch on OpenAI’s audit logging and Anthropic’s Activity Feed collection now, while you have nothing to investigate.

The checklist

For the day it happens to you, in order:

  1. Kill the agent session and revoke its git credential.
  2. Copy the agent transcripts, shell history, local repo and reflog; hash them.
  3. Mirror-clone the public repo; scan full history; build the secrets and personal-data inventory.
  4. Rotate every secret, admin keys first; verify the old ones are dead.
  5. Screenshot traffic, forks and stars; then make the repository private.
  6. Rewrite history with git-filter-repo; contact GitHub Support; file removal requests for known forks.
  7. Pull provider logs for every leaked secret: usage, audit, activity. For AI keys, use the dashboards and APIs above.
  8. Run the GDPR assessment; notify within 72 hours of awareness if required; document either way.
  9. Fix the agent’s permissions before you let it run again.

So what

The uncomfortable lesson of this case is that no one did anything malicious, and it still became a breach with regulatory consequences. We spent the last two years worrying about agents being turned against us by prompt injection and poisoned tools. Those threats are real. But the more common failure is simpler: we gave software the same credentials and reach as a senior engineer and none of the judgement, then pointed it at a deadline.

Every team running coding agents should rehearse this incident before it happens. Pick a test repository, plant a canary token, let the agent push it, and time how long it takes you to rotate, audit and answer the regulator’s questions. If you cannot read your AI provider’s usage by key today, when nothing is on fire, you will not learn it at 2 a.m. with the 72-hour clock running.

Stay paranoid. Rotate first. And never give an agent a push it does not need.

Further Reading:

Questions or feedback? Reach out via:

Contact: info@vulnex.com

Posted in AI, Business, Security, Technology | Tagged , , , , , | Leave a comment

Intent, Not Sophistication: The AI Attacker Is on the Record

Read Time: 16 minutes

TL;DR

Two documents landed within days of each other this month, from two sources with nothing in common except the thing they describe. On September 14, Spain’s data-protection regulator, the AEPD, announced it had received the first GDPR breach notification for an attack executed by an autonomous AI agent: a vulnerability scan, a valid login, an autonomous hunt for holes in the application, modified personal data and accessed invoices, all chained by the agent without a human steering. Days earlier, Anthropic had published a threat report cataloguing its own model being used the same way at global scale: agent swarms, malware that rebuilds itself to dodge detection, a solo hacktivist running operations that used to need a state team. A regulator with no product to sell and a vendor with every incentive to look good, independently describing the same animal. The headline the vendor hands you is blunt (“sophisticated attacks no longer require sophisticated attackers”) and the sentence I would frame is sharper still: “the main distinguishing feature between these classes of actors is no longer sophistication but intent.” The skill gap did not shrink. It collapsed, and it is now on the record in a Spanish regulator’s blog. This is a practitioner’s read, not a summary: the three shifts that actually change your job, why the regulator’s case matters more than the vendor’s report, and a Monday checklist that the two sources converge on almost line for line.


A note on what this is. I disrupted none of these operations and I can independently verify none of them. One source is a vendor reporting on its own model; the other is a regulator summarizing a filing it received. I will take both seriously, because they line up with each other and with what the rest of us see from the outside, and I will be clear about where the vendor’s incentives and the reader’s interests part ways. Defense-oriented throughout, no operational detail, threat-vector level: the usual house rules.

Every so often two documents land that say, in voices with more data than mine, the thing you have been saying into the wind for a year. This month I got two, from opposite ends of the world and opposite ends of the incentive spectrum.

The first is small, Spanish, and for my money the more important. On September 14 the AEPD, Spain’s data-protection authority, announced it had received its first notification of a personal-data breach in which the incident, in the regulator’s careful conditional, “would have been executed” by an AI agent running on a well-known language model. Not AI-assisted, the phishing-email-written-by-a-chatbot genre we have all seen. AI-executed. According to the account in the notification, the agent started with a vulnerability scan against generic files, achieved a valid login, and once inside went looking, on its own, for vulnerabilities in the application; when it found one, it modified personal data and reached invoices. The AEPD is explicit that this is a qualitative change: an agent, in its words, “can receive an objective, plan intermediate tasks, use tools, execute code, consult sources, interpret results and modify its behaviour, autonomously, depending on what it finds.” A third party, the regulator says, used it “as an instrument to successfully chain the different phases of the attack.” The company is unnamed, the account comes from the company’s own filing and the AEPD says it still needs further analysis. The regulator is also careful on a point I want to keep: using a given model does not mean the model or its provider’s infrastructure was compromised, nor that the tool was designed for malicious use. What matters is that a European regulator has now put on the record, from a real case, the thing I have been arguing all year.

The second document is big, American, and comes with a caveat I will get to. Days before the AEPD post, Anthropic published its September 2026 threat intelligence report, cataloguing, across seven harm areas and eight months (December 2025 to August 2026), a parade of “Generative Threat Groups” that used its model, Claude, to run cyber operations, influence campaigns, fraud, and surveillance. Where the AEPD gives you one company and one agent, the vendor gives you the same shape at planetary scale.

My own paper trail on this is long. In 2026 I have written that the model is becoming the attacker, that agent skills are a loaded weapon, and that the AI supply chain is a soft target. Earlier this month I wrote about a frontier-lab researcher resigning with a warning that these would soon be “systems that can hack anything,” and then about the CEO’s call to slow down and why pacing the frontier does nothing for the capability already loose in the valley. This piece is the receipts, and the point of leading with the Spanish case is that they are no longer only the vendor’s receipts.

Let me do what a practitioner should do with documents like these: not recap them (the summaries will be everywhere) but pull out what actually changes your job, and be honest about what to trust.

The one sentence that matters

Strip both documents to a single load-bearing claim and it is the vendor’s: “the main distinguishing feature between these classes of actors is no longer sophistication but intent.”

For my entire career, the threat model has been layered by capability. Script kiddies at the bottom, organized crime in the middle, nation-states at the top, and your defenses calibrated to who you thought would bother with you. That ladder is the thing both documents say has fallen over. When a solo hacktivist can field the same autonomous tradecraft as a state espionage team, and when an unidentified third party can point an off-the-shelf agent at a Spanish company and walk away with modified records and invoices, “who is sophisticated enough to hurt me” stops being a useful question. The only variable left is who wants to, and intent is cheap, plentiful, and impossible to patch.

The vendor says the quiet part directly: “sophisticated attacks no longer require sophisticated attackers,” because AI “has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators.” The regulator says the same thing in drier prose: AI agents do not introduce new techniques, but they “increase the speed, scale and adaptability of already known malicious techniques, reducing the time available to detect and contain them.” Same old attacks. New tempo, new operators.

One more detail from the vendor’s report deserves its own sentence, because it changes how you should read everything below. Anthropic states that the misuse cases ran on its Haiku, Sonnet and Opus models, and that “none of the misuse cases involved the use of Claude Fable or Mythos-class models, with the exception of one illicit distillation case.” Nobody in this catalogue needed the frontier. Every operation in it ran on the tier that is already everywhere, already cheap, and whose rough equivalents you can download and run with nobody watching. That is the valley I keep pointing at, described by the lab that owns the summit.

Three shifts that actually change your job

Under the case studies, three structural shifts are doing the real work. These are the parts worth your attention, and I will note where the Spanish case and the vendor’s data say the same thing.

1. The skill gap collapsed

The vendor’s cast is the tell. A Chinese espionage group (designated GTG-10007) whose operators include, per the report, two university undergraduates in Changsha, ran “agent swarms” that decomposed reconnaissance into parallel subagents and maintained an autonomous vulnerability-research program that turned up multiple previously unknown vulnerabilities in security products. A single French hacktivist (GTG-50029) got into at least fourteen organizations and built a doxxing platform with real ingestion pipelines. A financially motivated operator (GTG-50014) directed the AI toward goals and let it “evaluate environments and execute iteratively” (the report calls this “vibe hacking,” and yes, the name stings), then decompiled 1.8 million Android APKs to harvest hardcoded secrets and walked out of one airline with tens of millions of passenger records.

Now put the AEPD case next to that. One agent, one objective, one company, and the same phases (scan, access, vulnerability hunt, data) chained autonomously. It is the vendor’s pattern at the smallest possible scale, and it is the scale most of my readers actually live at. None of these people, on paper, should be able to do what they did. The AI is the force multiplier that lets intent skip the decade of skill-building it used to require.

2. The inverted cost structure: the loop closes faster than yours

This is the most important technical idea in the vendor’s report, and the one defenders should lose sleep over. In the Russian espionage case (GTG-20006), when a security product flagged the group’s malware, “agents would then set about the process of autonomously modifying and rebuilding the malware to evade the existing detections.” Anthropic draws the conclusion plainly: previously a defender could slow an attacker by shipping a new detection; now “capable adversaries can ‘close the loop,’ bypassing traditional security detections faster than defenders can develop and deploy them.”

The AEPD reaches the same place from the other direction. Among its lessons: procedures designed for manually executed attacks “may prove insufficient when an agent analyses multiple assets simultaneously,” and while human oversight “remains indispensable,” it “must be supported by detection, containment and response mechanisms able to operate fast enough.” A regulator and a vendor, independently, describing the same race.

Read that as a security engineer and it is a phase change. Detection has always had a shelf life, but the shelf life was measured against human attacker tempo. When the attacker’s evade-rebuild-redeploy loop is autonomous, it can spin faster than your observe-analyze-write-test-ship loop, which still has humans in it. The economics of defense have quietly inverted: the side that can close its loop fastest wins, and for the first time that is not automatically the defender.

srf_atr_inverted_cost_loop
Figure 1. The inverted cost structure. The attacker’s loop (deploy, get detected, let the model rebuild and mutate the malware, redeploy) can now close in hours, autonomously. The defender’s loop (observe the new technique, analyze it, write a detection, test and ship it) still has humans in it and closes in days or weeks. When the red loop spins faster than the blue one, a detection is obsolete before it is deployed.

3. The AI supply chain is now a target, not just a tool

I have been banging this drum since the dependency-trap work, and the vendor’s report escalates it from theory to campaign. One group (GTG-50020) compromised an AI vendor’s evaluation sandbox, extracted production API keys from multiple providers, hit around thirty AI companies in about four days, and, the detail I cannot stop thinking about, explicitly went looking for pre-release model access (they failed; every path they tried was closed). Another (GTG-50021) ran a fraudulent Claude reseller, offering cheap access while proxying traffic elsewhere and harvesting the credentials for resale. Others prompt-injected LiteLLM wrapper services to pull production keys straight out.

The AEPD, from its side, lands on the same object. Its words: “an agent that obtains an account, an API key or a token with excessive permissions can operate at the speed of a machine.” The vendor’s own recommendation is the one I would have written, treat “AI keys and agent integrations with the same level of seriousness” as production credentials, and the regulator’s is the mirror image, from the victim’s end. An API key to a capable model is loot (resold), compute (your bill, their attack), and cover (their activity, your name), all at once. If your threat model still files “AI keys” next to “SaaS logins we’ll rotate eventually,” it is out of date, and now a regulator agrees.

Put the three shifts together and they form a single structure. I modelled it as an attack tree, built entirely from the vendor’s own report, because the point is not any one branch but the shape of the whole.

srf_atr_ai_enabled_operation_tree
Figure 2. Anatomy of an AI-enabled operation, drawn from the vendor’s report. Four phases chained with AND: lower the skill floor (off-the-shelf frameworks like PentAGI, vibe hacking, agent swarms), run the autonomous loop (recon and exploitation, autonomous vulnerability research, self-modifying malware, unattended bulk exfiltration), target the AI supply chain (harvest production keys, compromise an eval sandbox, fraudulent reseller, prompt-inject a wrapper), and cash out (exfiltrate at scale, resell credentials, run attacks on the victim’s bill, influence-as-a-service). Within each phase the branches are OR — any one route is enough. With the floor this low, the discriminator between a state team and a solo operator is no longer sophistication but intent.

The rest of the catalogue, briefly

The cyber cases are the spine, but the vendor’s report is broader, and two other threads are worth a security reader’s glance. On influence operations, it documents commercial influence-as-a-service: a France-based agency (GTG-54002) that mass-produced 8,913 articles across roughly seventy fabricated news sites in about twenty languages, amplified by more than 250 inauthentic accounts, switching political stances by client rather than ideology: propaganda as a subscription product. And a firm in Istanbul (GTG-84005) that sold a “military-grade, AI-driven” political-operations platform running about a thousand fake accounts to work a Malaysian election constituency by constituency. The same automation, pointed at democracies instead of networks.

The connective tissue, in the vendor’s own framing, is that offensive AI tradecraft is proliferating: publicly available agent frameworks like PentAGI “reproduce much of the same scaffolding” and automate each step of the kill chain, so this is no longer the province of the well-resourced. The floor came up to meet everyone. The Spanish company in the AEPD notification is what it looks like when that floor reaches a business that never imagined it was a target.

The honest caveat, and why it no longer holds you back

Now the part a good practitioner cannot skip.

The vendor’s report is Anthropic reporting on Anthropic. The threat groups are self-designated, the disruptions are self-reported (accounts banned, monitoring added, intelligence “shared with authorities and industry partners where appropriate”), and none of it is independently verifiable from where you and I sit. And disclosure like this is never disinterested: a report that says “our model is so capable that nation-states and criminals race to abuse it, and we caught them” simultaneously demonstrates responsibility, markets the product’s power, and hands regulators a reason to prefer incumbents who can afford this kind of monitoring. All three can be true at once. I said as much when I first read it, and I stand by it.

But this is exactly why the AEPD case matters more than its modest size suggests. A data-protection regulator has no model to sell, no capability to hype, and no reason to flatter the vendor. It has a notification, submitted under legal obligation by a company that would much rather not have submitted it. When the party with every incentive to look good and the party with none describe the same attack within the same week, the pattern is real. The vendor’s report told me the threat was global; the regulator’s post told me it was inside an ordinary Spanish company’s application, editing records. Take the vendor’s intelligence, weigh the vendor’s framing, and then notice that a regulator just confirmed the substance from the other side of the table.

There is one quieter tension I will leave unresolved. Every operation in the vendor’s report ran on a closed model behind safety training and a trust-and-safety team that eventually caught it, which is, read one way, an argument for the closed and monitored model. Read another way, it is a reminder that the same tier of capability is diffusing into open-weight models no vendor is watching, where there is no one to write the disruption report at all. The AEPD case, notably, does not say which model the attacker used, only that it was a well-known one, and it goes out of its way to say the provider was not compromised. It may not matter which. I would rather sit in that discomfort than pretend one side is obviously right.

Your Monday checklist: where the two sources converge

Where vendor intel is always thin is the same place the AEPD is unusually strong: what defenders should do. And here is the thing that convinced me more than any single statistic. The regulator’s recommendations and the ones I had already drafted from the vendor’s report line up almost one for one. When two sources with nothing in common arrive at the same controls, that is your checklist.

Put AI-executed attacks in your threat model by name. The AEPD’s first lesson: it is “not enough to include a generic reference to malware, phishing or unauthorised access” in your risk analysis; attacks assisted or executed by AI go in expressly, as their own adversary. Calibrate to intent and to the AI floor now available to anyone, not to assumed skill.

Treat AI credentials as crown jewels. Vault them, scope them, rotate them, and monitor their usage the way you monitor a domain admin account. Both sources land here independently. An exposed model key, or an API token with more permissions than it needs, is not an inconvenience; it is loot, compute, and cover in one string, and an account with excessive permissions is exactly what the regulator says lets an agent run at machine speed once inside.

Instrument the agent layer, because you cannot IR what you cannot see. If autonomous agents are acting in or against your environment, you need telemetry at that layer: what ran, what it touched, what left. This is the single biggest visibility gap in most shops, and both documents are the argument for closing it now.

Assume your detections have a shorter shelf life, and detect at machine speed. If the adversary can rebuild around a signature autonomously, static, signature-heavy defense degrades fast. The AEPD’s version: human oversight stays, but it has to lean on detection, containment and response that can keep up. Shift weight toward behavioral detection and anomaly baselines, and toward response that does not wait for a human to read a ticket.

Inventory and watch your own AI supply chain. Every wrapper, proxy, MCP server, and eval sandbox that touches a production key is now attack surface. Prompt injection against a LiteLLM wrapper is in the vendor’s report; treat that class of component as security-relevant infrastructure, not glue code.

Rehearse the clock. The Spanish case ended where every European breach ends: in a notification to the regulator, due within 72 hours of becoming aware of it under GDPR Article 33. When the attack itself is executed by an agent that scans, logs in, finds a hole and edits data in one run, “what happened, to which records, and when” is a much harder question to answer in three days than it used to be. If your agent-layer telemetry is thin, that is the moment you will discover it.

Do not forget the boring foundations. The regulator did not. The AEPD closes, sensibly, on the unglamorous basics: know your processing, minimize what you collect, restrict access, fix vulnerabilities, control your vendors, and be ready to respond. An autonomous agent is a new attacker; it still walks in through an old door. In Europe that foundation is no longer just good practice: under the new Product Liability Directive, a missing patch is on its way to being a defect you are liable for.

So what

The comfortable version of AI-and-security says the dangerous capabilities are years away, sitting with a handful of labs. This month a vendor published eight months of evidence that they are here, in the hands of undergraduates and solo hacktivists and mid-tier crime groups, running on models a generation behind its best, and a Spanish regulator published the first formal record of one of them walking into an ordinary company’s systems and helping itself. The only thing separating those actors from a nation-state is what they decided to do on a given Tuesday. Intent, not sophistication.

That is not a doom sentence. It is a work order, and for once it comes co-signed. The vendor tells you the threat is global; the regulator tells you it is local, on the record, and subject to notification law. The response to a leveled playing field is not despair; it is to level up the defense: name the AI attacker in your threat model, guard the keys, instrument the agent layer, assume your detections rot faster, rehearse the clock, and keep the boring foundations that the agent still needs to get through. Take the vendor’s gift and read the fingerprints on the wrapping. Then read the regulator’s post, which has no wrapping at all.

Stay paranoid. Guard your keys. Assume the loop closes faster than yours, and assume the next notification could be yours.

Further Reading:

Questions or feedback? Reach out via:

Contact: info@vulnex.com

Posted in AI, Privacy, Security, Technology | Tagged , , , , , , | Leave a comment