Knowlegic
Technology

Anthropic Will Pay You 35,000 Dollars to Break Its Own AI

The companies building the world's most advanced AI models are now running a second business on the side: paying strangers to prove those models can be broken, and the going rate keeps climbing.

Knowlegic Editorial TeamSeptember 15, 20265 min read0 views
Share
Anthropic Will Pay You 35,000 Dollars to Break Its Own AI

Somewhere right now, a freelance security researcher is trying to trick Claude into describing how to synthesize a dangerous chemical, not because they want the answer, but because Anthropic will pay them up to 35,000 dollars if they succeed in a genuinely new way.

That's not an isolated bounty. OpenAI runs a nearly identical program for its own models, xAI runs one for Grok through HackerOne, and Google launched its own AI Vulnerability Reward Program in 2025, paying up to 30,000 dollars for a single serious flaw. Every major AI lab now pays people specifically to break its flagship product before someone with worse intentions does it for free.

The pay has been climbing all year, and it's climbing because the labs building these models have quietly concluded that a stranger finding the flaw for a fee is far cheaper than a stranger finding it for a headline.

The Business of Being Wrong First

A "jailbreak" is a prompt or technique that talks an AI model into ignoring its own safety rules, describing something dangerous, or acting outside its intended limits. For years, people did this for fun, for internet clout, or to make a point. What changed in 2026 is that the labs started treating those same people as a workforce.

Anthropic's Model Safety Bug Bounty Program went fully public in 2026, opening beyond an invite-only, NDA-bound group to anyone who wants to try. The top reward, up to 35,000 dollars, is reserved specifically for a universal jailbreak: a technique that doesn't just trick the model once, but reliably breaks its defenses across an entire high-risk category, like chemical or biological weapons information.

OpenAI runs the same logic under a different name. Its dedicated program for high-risk domains raised its own top reward for a universal break from 25,000 to 50,000 dollars in mid-2026, and backed the broader effort with an annual bounty pool in the millions.

Think of it like a bank hiring its own safecrackers, on contract, specifically to keep trying to open the vault, and paying a bonus scaled to how badly the vulnerability they find could have been used against a real customer.

Did You Know?

The reward isn't flat. A minor policy bypass, tricking a model into a mildly inappropriate response, might earn a researcher a few hundred dollars. A universal jailbreak against a genuinely dangerous capability, the kind of flaw that could theoretically help someone cause real harm, can pay into six figures at the top end across different labs' programs. The price tag is effectively a public signal of how seriously each lab ranks the underlying risk.

A Labor Market Built Entirely on Finding Flaws

This has stopped being a side project for curious hackers. Full-time red teamers, people whose entire job is professionally trying to break frontier models, average around 129,000 dollars a year in the US according to Glassdoor's own salary data, with top earners clearing well over 200,000. Contract specialists doing the same work hourly can bill 100 to 200 dollars an hour, with senior researchers on the most sensitive engagements clearing more.

The same agentic capability that makes AI agents useful is exactly what red teamers are paid to stress-test: an agent that can take independent action is also an agent that can be talked into taking the wrong one. That dual nature is why this labor market exists at all.

Estimates of the market's size vary by how narrowly it's defined, but analysts agree on the shape: several market-research firms put the broader AI red-teaming industry in the low billions of dollars in 2026, growing somewhere between 28 and 36 percent a year for the rest of the decade. A market that size, built around paying people to find what's broken, didn't exist in any organized form five years ago.

Interesting Fact: HackerOne, one of the platforms labs use to run these programs, has documented AI-specific submissions growing fast enough that it now tracks "AI red teaming" as its own reporting category, separate from traditional software bug bounties. The distinction matters because a flaw in an AI model rarely looks like a flaw in ordinary code. There's no broken line to point to, just a pattern of words that produces an answer the model was never supposed to give.

Why the Labs Are Paying, Not Just Testing Internally

It would be simpler, in theory, for a lab to only use its own employees to look for these flaws. The labs have decided that's not enough, and the reasoning is mostly financial rather than sentimental.

An outside researcher, motivated by a real cash reward, will try inputs and angles no internal team has time to imagine, because the internal team already has a mental model of how their own system is supposed to work. A stranger doesn't share that blind spot. Paying for the discovery, rather than waiting to be embarrassed by it publicly, is simply the cheaper failure mode.

That calculation scales with how much the labs are already spending elsewhere. Companies pouring tens of billions of dollars into the chips and infrastructure behind these models treat a bounty pool in the low millions as a rounding error against the cost of a single serious safety failure making headlines.

Knowlegic Perspective

It's tempting to read all of this as a safety story: labs being responsible, hiring watchdogs, closing gaps. That's true as far as it goes, but it undersells what's actually happened. A market this well funded, with published pay bands and full-time career paths, only exists because the people running these companies have calculated, in cold financial terms, exactly what an undiscovered flaw could cost them.

That calculation is also a form of power. Whoever sets the bounty decides what counts as a serious flaw worth six figures and what counts as a minor one worth a few hundred dollars, and that decision shapes where a generation of security talent spends its time. The labs aren't just buying safety. They're buying the attention of the very people most capable of proving them wrong.

Anthropic didn't start paying 35,000 dollars for a broken safety rule because it suddenly became more cautious. It started paying because it had already priced, to the dollar, what a rule breaking on its own would cost instead.

Sources & References

Enjoyed this?

Get notified when a new Knowlegic story worth knowing is published.

By subscribing you agree to our Privacy Policy.