
Mantis turns a coding agent into a security review team. It's a set of agent skills, one per step of a security audit, plus a reference harness that chains them: read the code's history, build a threat model, hunt for flaws, throw out the false alarms, prove the real ones with a working reproducer, patch them, and write the report.
It comes from Google, but the README is clear that it's a demonstration, not a supported product. The first lines are a warning in capitals: run it only in isolated environments, and have a security expert check every finding before reporting it.
| Maker | Google, open source, "not an officially supported Google product" |
| Released | June 2026, actively updated |
| Popularity | About 1,950 GitHub stars |
| Language | Python |
| Licence | Apache 2.0 |
| Works with | Gemini CLI, Antigravity CLI, Google ADK, and other coding agents through the skills format |
| Needs | Docker, ideally gVisor for sandboxing, and a model API, such as Gemini on Vertex AI |
What it does
Mantis splits a security review into about 20 skills. The main pipeline runs them in order:
| Step | Skill | What it does |
|---|---|---|
| 0 | mantis-history |
Reads the version history for past vulnerabilities not to repeat |
| 0 | mantis-structural-index |
Builds an index of the code for fast navigation |
| 1 | mantis-architecture |
Writes a knowledge base of how the code is built |
| 2 | mantis-threat-model |
Develops a living threat model from that knowledge base |
| 3 | mantis-plan |
Maps the attack surface and plans what to scan |
| 4 | mantis-researcher |
Hunts for security flaws, in parallel |
| 5 | mantis-dedupe |
Merges duplicate findings |
| 6-7 | mantis-review, mantis-critic |
Check each finding and drop false positives and issues that can't happen in production |
| 8 | mantis-reproduce |
Writes a proof of concept and runs it in a sandbox |
| 9 | mantis-chain |
Combines findings into multi-step attacks |
| 10 | mantis-patch |
Writes a minimal fix and checks that it blocks the reproducer |
| 11 | mantis-calibrate |
Rates severity against a fixed rubric, so not everything comes out "critical" |
| 12-13 | mantis-reflect, mantis-report |
Records lessons for the next round and writes a report for humans |
After a review, mantis-advise uses the threat model and everything learned to guide new code, so a later change doesn't remove a guard that once blocked a bug.
For business people
The problem: AI writes code faster than security teams can review it. A classic scanner flags 1,000s of possible issues, most of them noise, and someone has to sort them by hand.
What Mantis changes: it doesn't stop at "this might be a bug". It tries to prove each finding with a working reproducer, patches it, and checks the patch. And it ranks severity on a fixed scale, so the report surfaces the few real risks first.
Where it fits: security reviews of in-house software, audits of new code before release, and teams that build with AI agents and want a security check that keeps up.
Costs: the toolkit is free. The model calls aren't, and a full review of a large codebase makes a lot of them. The harness has spend limits, which you can switch off for long unattended runs.
The risks, from the README itself:
- It runs code it generated itself. Use it only on isolated machines with no access to production, sensitive data or internal networks.
- Findings can be wrong. A security expert must verify every finding before anyone reports it. Google asks users not to send unverified AI reports to open-source maintainers.
- No reproducer doesn't mean safe, and a reproducer doesn't mean exploitable everywhere.
For technical people
Install the skills into your coding agent:
npx skills add google/mantis
Or run the full ADK reference harness:
cd reference && ./install.sh
gcloud auth application-default login # if using Vertex AI
python3 scripts/configure.py --auto # detect models and sandboxes
python3 scripts/configure.py --test --probe # preflight check
./run.sh path/to/code # review a file or a repository
./run.sh path/to/code --focus "look for IDOR" # steer the planner in plain language
./run.sh path/to/code --objective "Audit for SSRF in webhook handlers"
- Sandboxing: reproducers run in Docker, ideally under gVisor (
runsc) with no network. gVisor runs on Linux. For frontier models, the README recommends a further sandbox layer with monitoring for escape attempts. - Research graph synthesis:
--objectivemakes the harness design a new agent graph for that specific goal, instead of running the standard pipeline. - Structured findings: the repository ships a
schema.jsonfor findings, and the guide describes a snapshot model so stages can be rerun or moved into your own framework. - Adaptable: the skills can be tuned for hardware RTL, infrastructure as code, ML pipelines or firmware, and the risk calibration for your own risk tolerance.
Value
My lab apps are written mostly by coding agents, and a few of them face the internet. Nobody does a security review on them today. The pipeline shape is what I'd borrow first: a threat model, then hunting, then a critic that throws findings out, then proof. The same "builder never judges its own work" idea as my council (see The Council ).
The mantis-advise step is the most useful for a solo builder: feed the threat model back to the agent while it writes code, instead of auditing afterwards. Running the full pipeline would need a separate, isolated Linux box with no route to my home network.
Further Reading

