plain

From an agent session.
To a test you can keep.

Explore with your coding agent. Keep the passed steps as plain-English YAML. Replay the flow from your terminal.

Get plain

For Claude Code, Codex and other MCP clients.

Explore the appSave the specReplay the flow

Explore the app.
Describe the action.

Your agent sends targets and claims through MCP. Jev picks from structured candidates and judges accessible state.

Passed actions.
Written in plain language.

The flow takes shape as you explore. Only passed actions become steps in the saved spec.

The session ends.
The test stays.

Save writes the passed steps as YAML. Read it. Edit it. Commit it.

Run it again.
Without the agent.

The CLI replays your saved flow. Jev still makes picks and judgments where needed.

demo.playwright.dev / todomvc

A little less to do.

Your test environment.

What needs to be done?
buy milk
write a regression test
2 items left All · Active · Completed

“Add buy milk to the list.”

fill

the new todo input

press

Enter

expect

a todo item named ‘buy milk’ is listed

Illustrative passed steps

add-todo.yamlSaved spec
name: add a todo
url: https://demo.playwright.dev/todomvc
steps:
  - goto: /
  - fill:
      target: "the new todo input"
      value: "buy milk"
  - press: Enter
  - expect: "a todo item named 'buy milk' is listed"
Plain language. A file you own.
Your terminal · illustrative replay
$ npx -p @gabe4coding/plain plain add-todo.yaml

✓ goto
✓ fill
✓ press
✓ expect

Flow replayed · simulated

No coding agent required.
Jev decisions still apply.

Try the illustrative flow yourself
A session, made repeatable.Illustrative demo · local simulation, not a live agent or Jev response.
demo.playwright.dev / todomvc

A little less to do.

Your test environment, not your production app.

What needs to be done?
write a regression test
1 item leftAll · Active · Completed
add-todo.yamlSession in progress
name: add a todo
url: https://demo.playwright.dev/todomvc
steps:
  - goto: /

Start with Explore to add a todo and see the steps become YAML.

Same plain language.
Different engines.

A website. A desktop app. A phone. The language is familiar; the setup belongs to each platform. Not one portable spec.

Browser

A little less to do.

buy milk
write a regression test

The browser.

Playwright handles the actions in Chromium or Chrome. Jev picks targets from structured candidates and judges claims against accessible state.

url: https://demo.playwright.dev/todomvc
steps:
  - goto: /
  - fill: { target: "the new todo input",
            value: "buy milk" }
  - press: Enter

plain · Playwright

Browser spec reference ↗
Desktop Test Fixture

A message in progress.

Hello from the test app.
Preview

Preview before you send.

The desktop.

xa11y works with the native accessibility tree. macOS is tested; Windows and Linux parity is not verified.

app: Desktop Test Fixture
steps:
  - fill: { target: "the Message text field",
            value: "Hello from the test app." }
  - click: the Preview button

plain-computer · xa11y
Native accessibility permissions required.

Desktop setup ↗

Your test app.

buy milk
Add
buy milk

The phone.

Appium drives iOS and Android, React Native apps included. Use a configured Appium server and an installed app on a test device or simulator.

platform: android
device: emulator-5554
app: com.example.todo
steps:
  - tap: the Add button

plain-mobile · Appium
Platform, device and app are explicit.

Mobile setup ↗

Read the evidence.
Not just the story.

The animation above explains the workflow. The repository holds real agent traces, task oracles and measured results.

Repository evidence · 21 September 2026 · revision 0c1698e · not a fresh run of this website.

Six local browser workflows.

The corrected comparison records 108 trials against Playwright MCP 0.0.82. It measures discovery, actions, checks and recovery.

Mean time fell 25–52% by model. API cost ranged from 24% lower to 7% higher. This is not a website-wide guarantee.

The suite uses synthetic interfaces. Cost intervals for two models include parity. A strict baseline failure involved a duplicate save.

About the explainer and save/replay proof

The README links a six-minute explainer, not a recorded agent → save → replay run. Its attachment returned HTTP 404 during verification.

Original README and video reference ↗

The benchmark proves measured agent workflows on local fixtures. It does not prove the complete saved-spec round trip. No new live round trip is claimed here.

Choose who makes
the decision.

Different tools, different responsibilities.
ResponsibilityplainHandwritten PlaywrightPlaywright MCP
Plan the flowYou, or your coding agentYou, in test codeYour coding agent
Pick an elementJev selects a described targetYour explicit locatorAgent chooses a locator or snapshot reference
Check the resultJev judges a claimYour assertionAgent checks tool output or writes an assertion
Keep and replaySave passed MCP steps as YAML. CLI replay still uses Jev.Commit test code. Runner needs no model.Agent can write tests. No plain YAML save/replay contract.

Use conventional tests for exact API contracts, pixel comparisons, custom fixtures and assertions that must not depend on model judgment.

A decision you
can inspect.

Jev is the decision model. It selects a target and judges a claim. Playwright, xa11y or Appium performs the action.

Write specific targets and claims ↗

Uncertain is not passed.

A claim passes at a probability of 0.9 or above and fails at 0.1 or below. Between those thresholds, the result is inconclusive.

passfailinconclusiveerrorskipped

See the state behind the result.

Inspect picks with find, read accessible state with snapshot, and ask about claims with ask. Rejected picks and non-passing claims include debug dumps.

Reports and artifacts ↗

Reuse a pick, not a judgment.

Spec runs can reuse accepted picks only when the candidate list is unchanged, one candidate matches and spatial evidence checks permit reuse. MCP does not use this cache. Claims are still judged.

How the pick cache works ↗

Put it to work.

Install the browser plugin in your coding agent, or run a spec with npx. You need Node 22 or newer, npm, and a TypeSafe or Vercel AI Gateway key.

First, set your key.

All plugins and CLIs read the user env file. Do not commit it.

# ~/.config/plain/.env
TYPESAFE_API_KEY=<your key>

For Vercel AI Gateway, use AI_GATEWAY_API_KEY instead.

Key and environment options ↗
In Claude Code
/plugin marketplace add gabe4coding/plain
/plugin install plain@plain-marketplace
In your terminal
codex plugin marketplace add gabe4coding/plain
codex plugin add plain@plain-marketplace
In your terminal
npx -p @gabe4coding/plain plain todo.yaml

npx installs the package. plain installs Chromium on first run. For other engines, see all plugin instructions.

One file.
Terminal to CI.

Start with the repository’s todo spec. It uses a public demo, needs no login and starts with goto.

Download todo.yaml

Unchanged source spec from this checkout. No live pass is claimed.

Run the first spec.

npx -p @gabe4coding/plain plain todo.yaml

Set the API key first. npx installs the scoped package. plain installs Chromium on its first run.

Then run it in GitHub Actions

Commit todo.yaml at your repository root. Save this workflow as .github/workflows/browser-specs.yml.

name: Browser specs
on: [pull_request]
permissions:
  contents: read
jobs:
  specs:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: gabe4coding/plain@main
        with:
          specs: todo.yaml
          picks: read
          junit: plain-results/junit.xml
          artifacts: plain-results
          artifact-name: plain-results
        env:
          TYPESAFE_API_KEY: ${{ secrets.TYPESAFE_API_KEY }}
Download the workflow

Set TYPESAFE_API_KEY as a repository secret. Fork pull requests do not receive this secret.

Replace @main with a published release tag before routine CI use. A newly merged version can precede its npm publication.

The action uploads JUnit on every run and failure artifacts when a run fails. It does not publish a results check.

Action setup and inputs ↗ · Report formats ↗ · Screenshots and traces ↗

Workflow adapted from the current CI guide and action.yml. The report above is historical benchmark data, not CI output.

Beyond a todo.

Use the source examples with their stated setup. These are repository examples, not newly executed demonstrations.

Forms, selects and checkboxes
Public demo at the-internet.herokuapp.com. No login or hooks. Download the exact spec. Run npx -p @gabe4coding/plain plain components-forms.yaml.
Repeated controls in rows and cards
The corrected benchmark includes catalog browsing and account-row edits. Read the task-level results ↗. Reproduction needs the local workflow harness and provider keys, not a standalone public-page spec.
Spatial claims
Read the local fixture spec ↗. Run from a built checkout with node scripts/e2e.mjs --only spatial --skip-mcp. The runner serves the fixture locally.
Native apps
Mobile examples and exact setup ↗. Appium, the platform driver and a test device or simulator are prerequisites. Preserve the fixture hooks and relative paths. Desktop setup and permissions ↗.

Before you run it.

Does replay still cost model calls?

Yes. The coding agent is absent, but Jev still selects targets and judges claims. Spatial classification can also need a call. Provider charges depend on usage. Measured costs and scope ↗.

What does the cache remove?

Spec runs can reuse accepted picks only when the candidate list and spatial evidence checks allow reuse. MCP does not use the pick cache. Claims are not cached. Cache rules ↗.

What happens when a target is ambiguous?

A weak pick is inconclusive, not passed. Claim probabilities between 0.1 and 0.9 are also inconclusive. Write specific targets and scope claims with within. Thresholds and phrasing ↗.

Can it understand layout and visual appearance?

Jev reads structured candidates and accessible state, not screenshots. Spatial prompts also receive rendered bounds, including reference text, frames and open shadow roots. Bounds do not prove color or image appearance. Missing labels, truncated evidence and unusual controls still limit results. Evidence and limits ↗.

What do I need, and where is my key?

Node 22+, npm and a TypeSafe or Vercel AI Gateway key. All CLIs read ~/.config/plain/.env. Desktop needs accessibility permissions. Mobile needs Appium and a configured device. Requirements and key options ↗.

Is this a replacement for every test?

No. Keep explicit conventional tests for exact assertions and pixel-level checks. Use test environments only. Stop before payment, booking or sending. Usage rules ↗.

Small decisions.
Specific evidence.

Speed and cost depend on the task and model. Read the measured comparisons, not a universal promise. Use test environments only. Stop before payment, booking or sending. Never bypass bot protection or put literal credentials in a spec.