09 Jul 2026 Pi is a minimal, open-source coding agent harness (a terminal CLI) from Earendil. The reason it caught my attention: the harness you wrap around a model, not just the model, drives cost. A Databricks evaluation found that for the same model, switching harness can cut cost roughly 2x, and pi sits right on the efficient frontier. That is a real lever on my token spend.

Why Pi instead of Claude Code
21 Sep 2026 the short answer: cost and control. Pi runs the same model through a cheaper harness, lets me pick any model from any provider, and gives me a core I can reshape. Claude Code gives me everything already built, and my whole setup already runs in it.
| Pi | Claude Code | |
|---|---|---|
| Models | 15+ providers, switch mid-session | Claude models only |
| Cost per task, Opus 4.8 at high effort | About $0.90 | About $2.10 |
| On a Claude Pro or Max plan | Billed per token as extra usage | Counts against the plan |
| Built in | 4 tools, everything else is an extension | Plan mode, sub-agents, MCP, permissions, hooks |
| Context | Minimal system prompt, replaceable per project | Anthropic's system prompt plus CLAUDE.md |
| Source | Open source, MIT | Anthropic's product |
The cost figures are read off the Databricks chart below, so treat them as approximate.
Pick Pi for:
- API-billed work at volume. Scripts, batch jobs, anything paid per token. Under Pi, the same Opus 4.8 cost less than half as much per task in the Databricks evaluation, for a pass rate 2 points lower.
- The right model per task. A cheap model for simple jobs, GLM 5.2 for value, Claude for hard reasoning, switched mid-session with
/model. It also signs in with ChatGPT Plus or Pro, xAI and Meta subscriptions. - An agent inside my own tools. Print, JSON, RPC and SDK modes on an open-source core that isn't tied to one lab.
- Exact control of the context window. A minimal system prompt,
SYSTEM.mdper project, and compaction I can rewrite.
Stay on Claude Code for:
- Anything the Claude subscription covers. Pi can sign in with a Claude Pro or Max account, but Pi's own docs say that usage "draws from extra usage and is billed per token, not against Claude plan limits". Inside the plan's limits, Claude Code costs nothing extra per task, so the 2x saving disappears.
- Features out of the box. Plan mode, permission prompts, sub-agents and MCP work on day one. In Pi, each one is an extension to build or install.
- My existing setup. Skills, hooks such as the notes auto-publish, memory and
CLAUDE.mdall live in Claude Code. Pi readsAGENTS.mdand supports skills, so part of it carries over; the hooks and memory would need rebuilding.
My call: Claude Code stays the daily driver, because the subscription pays for it and my tooling lives there. Pi earns its place for API-billed, scripted runs, and whenever a non-Claude model is the better tool for the job.
The discovery: the harness saves ~2x

From a Databricks evaluation shared by CEO Ali Ghodsi: 3,000+ engineers, 3 clouds, many languages and tasks, run on their own code base. Cost per task (x) against overall pass-rate (y); the red dashed line is the efficient frontier.
The takeaway is that the choice of harness, for the same model, changes cost by about 2x at similar quality. Reading the frontier:
| Setup | Pass-rate | Cost/task |
|---|---|---|
| Opus 4.8 (pi, high) | ~85% | ~$0.90 |
| Opus 4.8 (claude code, high) | ~87% | ~$2.10 |
| Opus 4.8 (pi, xhigh) | ~90% | ~$2.20 |
| Opus 4.8 (claude code, max) | ~89% | ~$4.30 |
| GLM 5.2 (pi) | ~87% | ~$1.30 |
Same Opus 4.8 model: under pi it lands ~85% for ~$0.90, versus ~$2.10 under Claude Code for a couple of points more. Push pi to xhigh and it tops the chart at ~90% while still undercutting Claude Code's priciest runs. The eval also flagged GLM 5.2 (z.ai) as strong value on pi. Numbers are read off the chart, so treat them as approximate.
What it is
- A terminal-first coding agent harness: Read, Write, Edit, Bash, and a model loop.
- Open source, MIT-licensed, from Earendil.
- Bring-your-own-key and provider-agnostic: one agent loop runs against Claude, GPT, Gemini, Grok, DeepSeek, local models, 20+ providers. No single-lab lock-in, no bundled subscription.
- Minimal by design, extended by you rather than shipped fully-loaded.
For business people
The insight worth internalizing: your AI coding cost is not set by the model alone. The harness (how it manages context, tools, and turns) moves cost by ~2x for the same underlying model and similar output quality. If you are paying per token at scale, the harness is a cost lever most teams ignore.
Pi leans into that. It is bring-your-own-key, so you pay providers directly at API rates and can point the same tool at whichever model is cheapest for a given task, instead of being locked to one lab's subscription. It is free and open source, so there is no seat cost. Databricks pairs this with a router ("Omnigent") to multiplex harnesses and models per task, which is the enterprise version of the same idea.
The trade-off: pi is minimal and technical. It is a terminal tool you shape with extensions, not a polished, batteries-included product. If you want plan mode, permission popups, and a to-do system out of the box, Claude Code or Codex are friendlier. If you want maximum control and lower cost, pi is the play.
For technical people
Pi keeps a four-tool core (Read, Write, Edit, Bash) and self-extends at runtime through TypeScript extensions, skills, prompt templates, themes, and packages. What it deliberately leaves out (and treats as extension points): MCP support, sub-agents, permission popups, plan mode, to-do lists, and background bash. The philosophy is a small core you reshape, not a fixed product.
It runs in four modes:
- Interactive TUI for normal use.
- Print / JSON (
pi -p "query") for scripting and pipelines. - RPC for driving it from another process.
- SDK for embedding the agent loop in your own app.
Other details: a unified model API across 20+ providers, tree-structured branching conversation history with export/share, and mid-session model switching. Install via npm:
npm install -g @mariozechner/pi-coding-agent
pi -p "explain this repo"
⚠️ WARNING: superseded. This npm package was deprecated in May 2026. The current install is in the September update below.
Why it matters for me
I run long Claude Code and Cowork sessions building skills, notes, and research, so token cost is a real line item. Two things here are directly useful: the general principle that the harness is a ~2x cost lever, and pi specifically as a BYOK, model-agnostic option that lets me route the cheapest capable model per task and script the agent (pi -p, RPC, SDK) into my own Python tooling. The cost of entry is that pi is DIY and terminal-only, versus the convenience of Claude Code. Pairs well with what I noted on Context Rot and the Re-fresh Skill: cheaper runs plus disciplined context management compound.
September 2026: what changed
21 Sep 2026 2.5 months and 20 releases later, from v0.80 to v0.86.1. The cost thesis above still holds, and the September releases push it further.
First, the install. The @mariozechner/pi-coding-agent package from July is deprecated; its last version was 0.73.1. Install from the official script or the new package:
curl -fsSL https://pi.dev/install.sh | sh
# or
npm install -g --ignore-scripts @earendil-works/pi-coding-agent
On Windows, irm https://pi.dev/install.ps1 | iex in PowerShell does the same.
| Version | v0.86.1, 20 September 2026 |
| Traction | 108,065 GitHub stars, 13,668 forks |
| Package | @earendil-works/pi-coding-agent |
| Providers | 15+ built in, hundreds of models, by API key or OAuth |
The site now claims 15+ providers, not the 20+ this note quoted in July. At least 3 arrived since then: Baseten, Meta's Muse models, and Earendil's own Radius gateway.
The releases that matter
| Version | Date | What it added |
|---|---|---|
| v0.84.0 | 6 August | Fullscreen TUI, Mermaid and LaTeX rendering, per-directory AGENTS.override.md, Baseten |
| v0.85.0 | 4 September | Persistent Claude thinking effort, restorable in-memory sessions in the SDK |
| v0.86.0 | 19 September | Prompt cache warming, /bug reports, per-model compaction budgets |
| v0.86.1 | 20 September | Meta Muse models through /login meta, faster repeat launches |
v0.86 and the cost story
2 changes in v0.86.0 cut the cost of long sessions, which is where a harness earns or loses the 2x above:
- Prompt cache warming. Pi keeps valuable prompt caches alive during long tool runs, and optionally while idle, with cost-aware refreshes. A cache that expires mid-run means paying full price for the whole prefix again.
- Instruction changes that keep the cache. Mid-conversation changes to the system prompt and tools are now stored in the transcript. They survive resume and branching, and the cached prefix stays intact on supported models.
Radius
Radius is Earendil's own model gateway. /login radius signs in by OAuth, and Pi ships the Radius model catalog so a model can be picked even offline.
A breaking change for scripting
v0.84.0 changed the JSON and RPC message_update events to carry deltas only. Anything that reads pi --mode json or the RPC stream now has to assemble the deltas between message_start and message_end. That matters for the Python scripting planned above: build against the delta format from the start.
Still left out, on purpose
The list hasn't changed: no MCP, no sub-agents, no permission popups, no plan mode, no built-in to-dos, no background bash. What has changed is how the site answers each one: build it as an extension, install a package, or use a plain tool. For sub-agents it suggests spawning Pi instances in tmux, which is exactly what the Herdr skill gives Pi (see Herdr).
Extensions get the full TUI, and the site's own proof is a DOOM extension running while the agent works:

Packages install from npm or git, and the site lists 50+ example extensions:
pi install npm:@foo/pi-tools
pi install git:github.com/badlogic/pi-doom
Further Reading

