Mantis

Google's open-source toolkit of agent skills that find, reproduce and patch security vulnerabilities in a codebase, from threat model to verified fix and report.

Mantis turns a coding agent into a security review team. It's a set of agent skills, one per step of a security audit, plus a reference harness that chains them: read the code's history, build a threat model, hunt for flaws, throw out the false alarms, prove the real ones with a working reproducer, patch them, and write the report.

It comes from Google, but the README is clear that it's a demonstration, not a supported product. The first lines are a warning in capitals: run it only in isolated environments, and have a security expert check every finding before reporting it.

Maker Google, open source, "not an officially supported Google product"
Released June 2026, actively updated
Popularity About 1,950 GitHub stars
Language Python
Licence Apache 2.0
Works with Gemini CLI, Antigravity CLI, Google ADK, and other coding agents through the skills format
Needs Docker, ideally gVisor for sandboxing, and a model API, such as Gemini on Vertex AI

What it does

Mantis splits a security review into about 20 skills. The main pipeline runs them in order:

Step Skill What it does
0 mantis-history Reads the version history for past vulnerabilities not to repeat
0 mantis-structural-index Builds an index of the code for fast navigation
1 mantis-architecture Writes a knowledge base of how the code is built
2 mantis-threat-model Develops a living threat model from that knowledge base
3 mantis-plan Maps the attack surface and plans what to scan
4 mantis-researcher Hunts for security flaws, in parallel
5 mantis-dedupe Merges duplicate findings
6-7 mantis-review, mantis-critic Check each finding and drop false positives and issues that can't happen in production
8 mantis-reproduce Writes a proof of concept and runs it in a sandbox
9 mantis-chain Combines findings into multi-step attacks
10 mantis-patch Writes a minimal fix and checks that it blocks the reproducer
11 mantis-calibrate Rates severity against a fixed rubric, so not everything comes out "critical"
12-13 mantis-reflect, mantis-report Records lessons for the next round and writes a report for humans

After a review, mantis-advise uses the threat model and everything learned to guide new code, so a later change doesn't remove a guard that once blocked a bug.

For business people

The problem: AI writes code faster than security teams can review it. A classic scanner flags 1,000s of possible issues, most of them noise, and someone has to sort them by hand.

What Mantis changes: it doesn't stop at "this might be a bug". It tries to prove each finding with a working reproducer, patches it, and checks the patch. And it ranks severity on a fixed scale, so the report surfaces the few real risks first.

Where it fits: security reviews of in-house software, audits of new code before release, and teams that build with AI agents and want a security check that keeps up.

Costs: the toolkit is free. The model calls aren't, and a full review of a large codebase makes a lot of them. The harness has spend limits, which you can switch off for long unattended runs.

The risks, from the README itself:

  • It runs code it generated itself. Use it only on isolated machines with no access to production, sensitive data or internal networks.
  • Findings can be wrong. A security expert must verify every finding before anyone reports it. Google asks users not to send unverified AI reports to open-source maintainers.
  • No reproducer doesn't mean safe, and a reproducer doesn't mean exploitable everywhere.

For technical people

Install the skills into your coding agent:

npx skills add google/mantis

Or run the full ADK reference harness:

cd reference && ./install.sh
gcloud auth application-default login            # if using Vertex AI
python3 scripts/configure.py --auto              # detect models and sandboxes
python3 scripts/configure.py --test --probe      # preflight check
./run.sh path/to/code                            # review a file or a repository
./run.sh path/to/code --focus "look for IDOR"    # steer the planner in plain language
./run.sh path/to/code --objective "Audit for SSRF in webhook handlers"
  • Sandboxing: reproducers run in Docker, ideally under gVisor (runsc) with no network. gVisor runs on Linux. For frontier models, the README recommends a further sandbox layer with monitoring for escape attempts.
  • Research graph synthesis: --objective makes the harness design a new agent graph for that specific goal, instead of running the standard pipeline.
  • Structured findings: the repository ships a schema.json for findings, and the guide describes a snapshot model so stages can be rerun or moved into your own framework.
  • Adaptable: the skills can be tuned for hardware RTL, infrastructure as code, ML pipelines or firmware, and the risk calibration for your own risk tolerance.

Value

My lab apps are written mostly by coding agents, and a few of them face the internet. Nobody does a security review on them today. The pipeline shape is what I'd borrow first: a threat model, then hunting, then a critic that throws findings out, then proof. The same "builder never judges its own work" idea as my council (see The Council ).

The mantis-advise step is the most useful for a solo builder: feed the threat model back to the agent while it writes code, instead of auditing afterwards. Running the full pipeline would need a separate, isolated Linux box with no route to my home network.

Further Reading

NicAI
Written by NicAI, Nic's AI assistant, for his personal knowledge base. Researched and drafted by the model, not hand-written by Nic. Verify anything you plan to act on.