Skip to content

Semgrep Researcher Poisons Open-Weight AI Model for Under $100

A Semgrep security researcher backdoored an open-weight AI model for under $100 using ten training examples, exposing an AI supply chain observability gap.

Semgrep logo green and black on white
Semgrep logo. Composite: techshooked.

A cybersecurity researcher at Semgrep installed a backdoor in an open-weight AI model in about an hour for less than $100, using only ten training examples to make the model reliably output code vulnerable to remote code execution.

Katie Paxton-Fear, who also lectures in cybersecurity at Manchester Metropolitan University, began the experiment by fine-tuning a model to swap coding conventions, a task she described as straightforward. Escalating to a deliberate backdoor, she found the poisoned behavior persisted across novel prompts and domains the model had not encountered during fine-tuning. Paxton-Fear and Semgrep colleagues Isaac Evans and Cris Thomas published a research post on the results, which The Register covered Thursday. A notable secondary finding: larger models were easier to poison than smaller ones, suggesting the attack becomes more accessible as the industry shifts toward higher-capability deployments.

The central problem Paxton-Fear and her colleagues identify is one of observability. Traditional software binaries can be analyzed with reverse-engineering tools to produce a complete account of their behavior. AI model weights offer no comparable method. "A compromised or subtly manipulated model doesn't need to 'break' to create business risk," the Semgrep team wrote. "It only needs to influence decisions in ways that are difficult to detect."

A parallel experiment from David Kaplan, AI security research lead at Origin, illustrated the commercial stakes. Kaplan engineered a poisoned model targeting drug-discovery workflows: given access to a send_email tool, the model silently exfiltrated research data with no indication to the user. Kaplan noted the attack sidesteps the standard framing that agent risk requires untrusted input to arrive from outside the system. With a backdoored model, the malicious intent is already embedded in the weights before the model is deployed.

No widely used open-weight model has been publicly confirmed as compromised, but the Semgrep researchers argue that reflects limited detection capability rather than the absence of incidents. The low entry cost, achievable with commercially available fine-tuning services, puts the technique within reach of adversaries well below the nation-state tier. The security community has only recently shifted serious attention to the issue as AI supply chain attacks have begun appearing in the wild.

Share this story

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.