Back to Intelligence

Google Suspends Open Source Vulnerability Rewards Program After Flood of AI-Generated Bug Reports — What Defenders Must Learn

SA
Security Arsenal Team
October 7, 2026
6 min read

Google has suspended its Open Source Vulnerability Rewards Program (OSVRP) after being flooded with low-quality, AI-generated vulnerability submissions. This is not a vulnerability disclosure story in the traditional sense — there is no CVE, no zero-day, no exploit chain. But make no mistake: this is a security operations story, and it carries a warning that every organization running a vulnerability disclosure program (VDP), bug bounty, or even a security@ inbox needs to internalize right now.

The AI-generated report problem — widely called "AI slop" in the research community — has been building since LLMs became mainstream. Projects like cURL and the Linux kernel have publicly complained about plausible-sounding but fabricated vulnerability reports consuming maintainer time. Google pausing one of its reward programs is the highest-profile acknowledgment yet that the economics of vulnerability triage have shifted against defenders. If Google's triage machinery can be overwhelmed, yours can too.

What Actually Happened

Per reporting from Infosecurity Magazine, Google paused the Open Source Vulnerability Rewards Program — the bounty track covering vulnerabilities in Google's open-source projects (think components and libraries Google maintains and ships publicly) — due to a surge of AI-generated submissions. The core dynamic:

  • LLMs lower the cost of report generation to near zero. A submitter can point an LLM at a public repository, ask it to "find vulnerabilities," and receive polished, technically formatted write-ups describing buffer overflows, race conditions, injection flaws — many of which are hallucinated, untestable, or describe behavior that is not actually a vulnerability.
  • The reports look legitimate. This is the dangerous part. Unlike spam, AI-generated reports use correct terminology, reference real files and functions, and include convincing PoC code. A triager cannot dismiss them at a glance; each one consumes real analyst time to disprove.
  • Bounty economics create perverse incentives. Because rewards are paid on valid findings, low-effort submitters play a numbers game: flood the queue, hope something sticks. The program operator bears the entire cost of filtering.

The result is a denial-of-service condition against human triage capacity — not against infrastructure. Legitimate researchers' valid reports sit in queues longer, real vulnerabilities take longer to confirm and patch, and maintainer/PSIRT burnout accelerates.

Why This Matters Beyond Google

Three defensive realities follow from this event:

1. Your disclosure pipeline is an attack surface. Any channel that accepts external input and routes it to a scarce human resource can be resource-exhausted. AI makes this trivially scalable. Organizations subject to regulatory disclosure expectations (PCI-DSS 6.3.3, HIPAA security incident procedures, CISA's push for VDPs across federal agencies) cannot simply shut the door — they must engineer the funnel.

2. Open-source dependency risk doesn't pause with the program. The OSVRP existed to incentivize eyes on Google's open-source code. Suspending rewards doesn't reduce the vulnerability count — it reduces the incentive for skilled researchers to look. Defenders consuming Google-maintained open-source components should assume a temporary reduction in externally reported findings and compensate with their own SCA/audit coverage, not relax it.

3. AI-generated findings will increasingly hit YOUR products. If you ship software or operate a PSIRT, expect an increase in AI-authored reports against your own attack surface — some valid, most not, all expensive to adjudicate. The time to build triage policy is before the flood, not during it.

Executive Takeaways

Because this story is about program integrity rather than a specific technical exploit, the defensive value is procedural. These are the recommendations we're giving clients running disclosure programs, bounties, or PSIRT intake:

1. Implement a proof-of-concept gate. Require a working, reproducible PoC — executable exploit, failing test case, or captured request/response pair — before a report enters human triage. AI slop typically fails at reproduction. A hard PoC requirement filters the majority of hallucinated findings at near-zero analyst cost. Google's core VRP and most mature programs already do this; formalize it if you haven't.

2. Publish explicit AI-generated submission policy. State plainly in your VDP policy that bulk AI-generated reports without researcher validation are out of scope and may result in account suspension. cURL's maintainers and HackerOne have both moved in this direction. Ambiguity invites volume; clarity reduces it.

3. Add automated pre-triage scoring. Before a human reads a report, run automated checks: Does the referenced file/function exist in the current tree? Does the claimed vulnerable code path compile and execute as described? Does the PoC actually trigger? Simple CI-based validation harnesses can auto-reject a large fraction of fabricated claims. This is the same philosophy as email spam filtering — layered, automated, upstream of humans.

4. Rate-limit and reputation-weight intake. Per-reporter submission caps, cooling-off periods after repeated invalid reports, and priority queuing for researchers with validated track records preserve triage capacity for signal over noise. Bounty platforms support this natively; self-hosted VDPs should implement equivalents.

5. Protect your open-source maintainers. If your organization maintains open-source projects, route security reports through a dedicated PSIRT or security team rather than directly to maintainers. Maintainer burnout from AI slop is a real retention and security risk — exhausted maintainers ship worse code and respond slower to genuine issues.

6. Watch the downstream effect on vulnerability intelligence. With one major incentive program paused, expect a temporary dip in publicly reported findings for affected open-source projects. Do not interpret a quiet feed as a safe dependency. Maintain your SBOM, keep SCA tooling current, and treat reduced external reporting as a reason to increase internal audit cadence on critical open-source dependencies — not decrease it.

The Bigger Picture

The security industry spent two decades building disclosure ecosystems that assume human researchers on both sides of the transaction. That assumption is now broken. LLMs generate infinite, plausible-looking input; human triage capacity remains finite. Programs that survive this shift will be the ones that treat their intake pipeline like any other production system: engineered against abuse, load-tested against volume, and instrumented so that analysts spend their time on findings that matter.

Google pausing a program is a pressure valve, not a solution. The organizations that come out ahead will use this moment to rebuild their disclosure intake for the AI era — automated validation first, human expertise reserved for what automation can't resolve.

Related Resources

Security Arsenal Red Team Services AlertMonitor Platform Book a SOC Assessment pen-testing Intel Hub

Is your security operations ready?

Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.