e2e by TesterArmy

An open-source testing framework for web and mobile apps: an AI agent drives the app from plain-English steps, then replays the recorded actions with no model calls until the app changes.

e2e is a testing framework where you describe what a user does in plain English, and an AI agent clicks through the app to do it. You then check the result with normal, exact assertions in the same test. The clever part: once a step passes, e2e records what the agent did and replays it on the next run without calling the model, until the app changes.

It's open source, from TesterArmy, a startup that sells the hosted version of the same idea.

e2e banner: the e2e wordmark and the line Open Source AI Testing Framework on a dark dotted background, with the TesterArmy helmet mascot in the corner

Maker TesterArmy, founded in 2026 by Oskar Kwaƛniewski and Szymon Rybczak, Warsaw and San Francisco
Backing Y Combinator, and a $1.2M pre-seed with angels including Guillermo Rauch (Vercel) and Charlie Cheever (Expo)
Released July 2026, before version 1.0
Popularity About 830 GitHub stars
Language TypeScript, installed with npm
Licence Apache 2.0
Platforms Web (Chromium, Firefox, WebKit through Playwright), iOS and Android simulators and emulators
Models Bring your own: a subscription, an API key or a local model

What it does

// tests/checkout.e2e.ts
import { test, expect } from 'e2e';

test('a member upgrades to Pro', async ({ app, agent, screen }) => {
  await app.open('/settings/billing');

  await agent.act('upgrade the workspace to the Pro plan');
  await agent.assert('the invoice preview shows a prorated amount');

  await expect(screen.getByRole('status')).toContainText('Pro');
});
  • agent.act: the agent performs a goal written in plain English.
  • agent.assert: the agent checks something on screen that's hard to pin with code.
  • expect(screen...): classic, exact assertions with locators, as in Playwright.
  • Record and replay: a passing agent step is recorded and replayed without the model on later runs. The agent only steps in again when the app changes.
  • No model needed for tests without agent steps.

For business people

The problem: end-to-end tests check that an app works the way a customer uses it. Written by hand, they break every time a button moves. Driven fully by AI, every run costs model calls and can behave differently each time.

What e2e does: it mixes the two. The AI handles the steps that are tedious to script, then the recorded path makes later runs fast, cheap and repeatable. Exact checks stay exact.

Costs: the framework is free. Model costs come from your own provider, and only when the agent has to work out a step again.

The limits:

  • Before 1.0. The APIs and config can still change between minor releases.
  • Telemetry on by default. The CLI sends anonymous usage data: commands, engines and failure points, not test or app content. You can switch it off.
  • Simulators and emulators only for mobile, through the hosted options or your own machine.

How it compares

Tool Approach
e2e Open source, AI steps plus exact assertions, record and replay, web and mobile
Momentic Hosted platform, plain-English YAML tests, self-healing, paid per step (see Momentic )
Playwright The classic: free, fast, every step written in code

For technical people

Set up a project:

npx e2e init

init asks for an engine (web or mobile) and a model provider, then writes a config and an example test.

Package Role
e2e The SDK, runner and CLI
@e2e-dev/web Browser engine through Playwright
@e2e-dev/mobile iOS and Android engine through agent-device
@e2e-dev/github Posts results as a pull request comment
@e2e-dev/kernel Hosted browsers from Kernel
@e2e-dev/eas Hosted iOS simulators and Android emulators from Expo's EAS
  • Docs for agents: the e2e package ships the full documentation in node_modules/e2e/docs, so a coding agent can read it offline.
  • Telemetry off: npx e2e telemetry disable, or set E2E_TELEMETRY_DISABLED=1.
  • Built for agents: the repository itself ships agent skills and an MCP config.

Value

It's the same problem as Momentic, with a better answer for a solo builder: open source, my own model, and record and replay so a nightly run of my lab apps doesn't call a model every time. The test file reads like a user story, which a coding agent can write from a prompt like "test that the dashboard loads and the pipeline table shows deals".

I'd start with 1 test per lab app that has a login or a form, run it on every change, and switch telemetry off.

Further Reading

NicAI
Written by NicAI, Nic's AI assistant, for his personal knowledge base. Researched and drafted by the model, not hand-written by Nic. Verify anything you plan to act on.