Introducing e2e: open source agentic testing for web, iOS, and Android
e2e is an open source TypeScript testing framework that combines Playwright-style locators and assertions with agent steps like agent.act(). The same tests run on web, iOS, and Android, agent steps run on the AI model or subscription you already pay for, and passing agent steps replay from a cache without calling the model.

Today we're open sourcing e2e, a TypeScript testing framework that lets agent steps and exact assertions live in the same test. It runs on web, iOS, and Android with one API, it works with the AI model or subscription you already pay for, and the setup wizard takes most projects from install to a first passing test in about a minute:
npx e2e inite2e comes out of the work we do every day at TesterArmy, where our testing agent tests other teams' apps. We built it based on the knowledge and feedback we got from our customers, and it will be the open foundation that powers TesterArmy in the future.
Here's what you can expect from this first release:
- Agent steps next to exact checks. You write role-based locators and auto-retrying assertions where precision matters, and use
agent.act(),agent.assert(), andagent.extract()where describing a goal is easier than scripting it. Because both live in one test, the report always says exactly what was verified. - The model runs once per step. When an agent step passes with a recorded check, e2e caches its actions and replays them on later runs without calling the model, so a cached step runs at the speed of a scripted test and costs no tokens.
- One suite for web and mobile. Browsers run through Playwright, and iOS simulators and Android emulators run through agent-device, with the same fixtures, locators, and assertions on every platform.
- The model and subscription you already have. Agent steps run on any AI SDK model, including local ones, or on a ChatGPT, GitHub Copilot, or SuperGrok subscription, and we add no markup on tokens.
- Support for coding agents. e2e ships an agent skill and an MCP server, so Claude Code, Cursor, or Codex can explore your app and write tests with real locators.
Why we built it
Testing so many different apps keeps surfacing the same trade-off. An agent can complete a goal like "buy the cheapest item on this list" without a single selector, but when the test passes, it's hard to say exactly what was checked. A scripted end-to-end test tells you precisely what it verified, and you pay for that precision by rewriting a selector every time someone renames a button or the flow in your app changes.
Teams usually pick one approach for the whole suite and live with its costs. e2e lets you make that choice per step instead of per suite.
My favorite thing about e2e is that you can gradually adopt agentic APIs where it makes sense. Migration from frameworks like Playwright to e2e is super simple: you port your tests using the same familiar APIs, then add agent steps where they help.
On top of that, e2e can run exploration bug bashes via e2e explore before you open a pull request, which makes a great verification step in your software factory. More on that is coming soon.
Agent steps and exact checks in one test
If you've written Playwright tests, most of e2e will feel familiar: role-based locators, auto-retrying assertions, and the same test() API. The three agent steps cover the rest. agent.act() carries out a goal you describe, agent.assert() checks a condition you describe in plain language, and agent.extract() reads data off the screen so you can use it later in the test.
import { test, expect } from "e2e";
test("a member upgrades to Pro", async ({ app, agent, screen }) => {
await app.open("/settings/billing");
await agent.act("upgrade the workspace to the Pro plan");
await expect(screen.getByRole("status")).toContainText("Pro");
await agent.assert("the invoice preview shows the Pro price");
});In this test, the agent works out the upgrade flow on its own, so a reworked billing page is far less likely to break it. The expect line then pins down the result with an exact locator, which means a passing run tells you the status really reads "Pro", whatever path the agent took to get there.
What agent steps cost
Agent steps are slower and more expensive than scripted ones on their first run, because each action needs a model call. We designed e2e so that you pay this cost once per step rather than on every run.
When an agent step passes with a recorded check, e2e saves the actions it took to a trace cache. On the next run, it replays those actions directly, without calling the model, so the step runs at the speed of a scripted test and uses no tokens. If the UI changes enough that the replay fails, the runner hands the step back to the live agent to find a new path, and that run costs model calls again. In practice, this means your token spend follows how often your UI changes rather than how often your tests run.
One API for web and mobile
e2e drives browsers through Playwright, and iOS simulators and Android emulators through agent-device. Your tests use the same fixtures, locators, and assertions on every platform, so your web app and mobile app can share one suite and one config:
import type { E2EConfig } from "e2e";
import { createAgent } from "e2e/agent";
import { web } from "@e2e-dev/web";
import { mobile } from "@e2e-dev/mobile";
import { gateway } from "ai";
export default {
targets: [
{ name: "web", engine: web({ url: "http://127.0.0.1:3000" }) },
{ name: "ios", engine: mobile({ platform: "ios", app: "com.example.app" }) },
],
agents: {
default: createAgent({ model: gateway("openai/gpt-6-luna-fast") }),
},
} satisfies E2EConfig;If you'd rather your CI run on infrastructure someone else maintains, e2e has first-class support for hosted browsers from Kernel and mobile simulators from Expo. Your tests stay the same either way, and only the config changes.
Models, keys, and secrets
Agent steps run on any AI SDK model. You can bring your own key through Vercel AI Gateway or OpenRouter, point e2e at a local model server, or sign in with the ChatGPT, GitHub Copilot, or SuperGrok subscription you already have. You pay your provider's price for tokens, with no added markup.
Test credentials stay in environment variables, outside the model's context. The agent can type a password into a login form without the password ever appearing in a prompt, which matters when the model runs on a third-party API.
Working with coding agents
e2e ships an agent skill and an MCP server (Model Context Protocol, the standard most coding agents use to call external tools). With them, Claude Code, Cursor, or Codex can open your app, explore it, write tests with locators taken from the real UI, and read the failure report when something breaks:
npx skills add tester-army/e2eThis is how tests get written in our own repo: the coding agent that built a feature also explores it and adds the regression test in the same change.
See it in action
The two-minute explainer below shows how agent goals and deterministic checks fit together in one test, then sets up e2e from scratch on a Next.js app. It's a quick way to see the whole wizard before you run it on your own project.
You can also watch the e2e explainer on YouTube.
Availability
e2e is available today on npm as e2e, under the Apache 2.0 license. It's still pre-1.0: the core API is the one we use every day, and we expect parts of it to change as more teams run it on their own apps.
Web, iOS, and Android are supported now. Desktop apps and other platforms aren't covered yet. Engines are pluggable, so you can write your own and keep the test API unchanged, and we're working on more platforms ourselves.
To get started, run this in your project:
npx e2e initThe docs cover writing tests, choosing a model, and running in CI, and tester.army/e2e has the full overview. The code lives at github.com/tester-army/e2e. If e2e saves you from rewriting a selector or two, a star helps more people find it, and issues and PRs are very welcome.