
e2e is a testing framework where you describe what a user does in plain English, and an AI agent clicks through the app to do it. You then check the result with normal, exact assertions in the same test. The clever part: once a step passes, e2e records what the agent did and replays it on the next run without calling the model, until the app changes.
It's open source, from TesterArmy, a startup that sells the hosted version of the same idea.

| Maker | TesterArmy, founded in 2026 by Oskar KwaĆniewski and Szymon Rybczak, Warsaw and San Francisco |
| Backing | Y Combinator, and a $1.2M pre-seed with angels including Guillermo Rauch (Vercel) and Charlie Cheever (Expo) |
| Released | July 2026, before version 1.0 |
| Popularity | About 830 GitHub stars |
| Language | TypeScript, installed with npm |
| Licence | Apache 2.0 |
| Platforms | Web (Chromium, Firefox, WebKit through Playwright), iOS and Android simulators and emulators |
| Models | Bring your own: a subscription, an API key or a local model |
What it does
// tests/checkout.e2e.ts
import { test, expect } from 'e2e';
test('a member upgrades to Pro', async ({ app, agent, screen }) => {
await app.open('/settings/billing');
await agent.act('upgrade the workspace to the Pro plan');
await agent.assert('the invoice preview shows a prorated amount');
await expect(screen.getByRole('status')).toContainText('Pro');
});
agent.act: the agent performs a goal written in plain English.agent.assert: the agent checks something on screen that's hard to pin with code.expect(screen...): classic, exact assertions with locators, as in Playwright.- Record and replay: a passing agent step is recorded and replayed without the model on later runs. The agent only steps in again when the app changes.
- No model needed for tests without agent steps.
For business people
The problem: end-to-end tests check that an app works the way a customer uses it. Written by hand, they break every time a button moves. Driven fully by AI, every run costs model calls and can behave differently each time.
What e2e does: it mixes the two. The AI handles the steps that are tedious to script, then the recorded path makes later runs fast, cheap and repeatable. Exact checks stay exact.
Costs: the framework is free. Model costs come from your own provider, and only when the agent has to work out a step again.
The limits:
- Before 1.0. The APIs and config can still change between minor releases.
- Telemetry on by default. The CLI sends anonymous usage data: commands, engines and failure points, not test or app content. You can switch it off.
- Simulators and emulators only for mobile, through the hosted options or your own machine.
How it compares
| Tool | Approach |
|---|---|
| e2e | Open source, AI steps plus exact assertions, record and replay, web and mobile |
| Momentic | Hosted platform, plain-English YAML tests, self-healing, paid per step (see Momentic ) |
| Playwright | The classic: free, fast, every step written in code |
For technical people
Set up a project:
npx e2e init
init asks for an engine (web or mobile) and a model provider, then writes a config and an example test.
| Package | Role |
|---|---|
e2e |
The SDK, runner and CLI |
@e2e-dev/web |
Browser engine through Playwright |
@e2e-dev/mobile |
iOS and Android engine through agent-device |
@e2e-dev/github |
Posts results as a pull request comment |
@e2e-dev/kernel |
Hosted browsers from Kernel |
@e2e-dev/eas |
Hosted iOS simulators and Android emulators from Expo's EAS |
- Docs for agents: the
e2epackage ships the full documentation innode_modules/e2e/docs, so a coding agent can read it offline. - Telemetry off:
npx e2e telemetry disable, or setE2E_TELEMETRY_DISABLED=1. - Built for agents: the repository itself ships agent skills and an MCP config.
Value
It's the same problem as Momentic, with a better answer for a solo builder: open source, my own model, and record and replay so a nightly run of my lab apps doesn't call a model every time. The test file reads like a user story, which a coding agent can write from a prompt like "test that the dashboard loads and the pipeline table shows deals".
I'd start with 1 test per lab app that has a login or a form, run it on every change, and switch telemetry off.
Further Reading

