ego (lite) is just a browser, ego is your personal agent across devices.
Join waitlist
PlaywrightCypressEnd-to-end testingTest automationMigration

Playwright vs Cypress: Reliability and Migration

Sep 14, 202618 min read
Last updated Sep 28, 2026
Playwright and Cypress logos facing each other on a blue illustrated background

Short answer: Playwright and Cypress turn known browser workflows into repeatable tests; ego (lite) is a separate visible browser route when an Agent must work through a changing live task. Playwright provides broader control over multiple pages, iframes, browser contexts, and browser engines, while Cypress places more emphasis on its interactive Runner and tightly integrated debugging experience. The two test frameworks fit stable processes that can be described in advance with selectors, actions, and assertions.

This shared assumption matters more than many feature-by-feature comparisons. When a workflow must run repeatedly and the application is controlled by the team, deterministic browser testing is the right fit. But when the website interface, execution environment, or task objective can change during the run, maintaining the test script can sometimes take more effort than completing the task itself, which is where a full Chromium like ego (lite) lets an agent work against the live page, cross-origin iframes and all.

Is Playwright better than Cypress?

For a greenfield end-to-end suite, Playwright is usually the safer default because its browser-context, page, frame, tracing, and parallel-worker primitives cover more application shapes without restructuring the product under test. That recommendation changes when the team depends on Cypress component testing, already has a healthy Cypress suite, or gets more value from the Cypress runner than it would gain from Playwright's broader browser-control model.

ConstraintPrefer PlaywrightPrefer Cypress
Several tabs or popupsPages are first-class objectsUsually redesign the test into one controlled tab
Cross-origin iframeFrame locators can target itOutside cy.origin() support
Component feedback loopSupported, but project fit variesA core Cypress workflow
Existing healthy suiteMigrate only for a measured gapKeep it if constraints are met

How do their execution models differ?

Playwright's test process controls browsers through its automation protocol. A test creates a browser, one or more isolated contexts, and pages inside those contexts. New pages and popups remain addressable, and each context can carry its own cookies, permissions, and storage. The test and the application are separate processes, which makes cross-page orchestration explicit.

Cypress coordinates a Node process, a proxy, and code running with the application in the browser. That architecture powers its live command log, DOM snapshots, time-travel debugging, and direct application feedback. It also explains why Cypress keeps control in one primary browser tab and asks tests to enter a second top-level origin through cy.origin().

Architecture is not a quality score. It is a constraint map. A checkout that embeds a cross-origin payment frame and opens a receipt window places different demands on a runner than a single-origin dashboard with rich component tests.

Which browser workflows can each test?

WorkflowPlaywrightCypress
Multiple tabs/windowsDirect page and popup eventsNo commands in another tab/window; keep the flow in one tab or test the destination separately
Top-level origin changeNavigate and locate normallyUse cy.origin() for commands on the secondary origin
Cross-origin iframeUse frameLocator or contentFrameNot handled by cy.origin()
Network observationEvents, routing, and response waitscy.intercept(), aliases, waits, and stubs
Isolated sessionsSeveral browser contexts in one browserTest isolation resets state between tests; session caching can restore selected setup

The important correction is that Cypress is not simply ‘single-origin.’ Current Cypress supports top-level cross-origin testing with cy.origin(). The narrower boundaries are that cy.origin() does not operate a cross-origin iframe, and Cypress does not execute commands in a different browser tab or window. Those distinctions should drive fixture design and tool choice.

How do waiting and retries differ?

Playwright separates actionability waiting, assertion retry, and whole-test retry. Before actions such as click, it checks conditions such as visibility, stability, event reception, and enabled state. Its web-first assertions retry until they pass or time out. Test retries are a separate runner setting and were disabled in our experiment.

Cypress chains queries and assertions and reruns that linked query chain while the assertion can still succeed. A non-query command such as click executes once; Cypress does not replay every preceding command because a later assertion failed. Whole-test retries are also separate and default to zero unless configured.

What happened in our controlled test?

The earlier article used a two-origin fixture with delayed product data, duplicate labels, a replaced status node, an intentional HTTP 503, a same-origin iframe, a cross-origin payment iframe, and a receipt popup. Its archived screenshot shows the fixture, but the September 11 per-run logs and scripts are not present here. The official pages and cross-origin documentation still explain the browser-control difference; the older screenshot alone cannot establish a pass rate.

Left side of the archived checkout fixture showing email input, duplicate product buttons, and rerender and receipt controls
Native-size crop of the September 11 fixture controls. The earlier per-run logs and scripts are absent, so the image does not establish historical success counts.

Open the full archived fixture screenshot

What changed when the login session changed?

On September 28, 2026, we ran a separate session experiment against our own local Request Desk. Playwright 1.59.1 and Cypress 16.0.0 used the same Chrome 153 browser, input, server, and zero-retry rubric. Each route ran clean, persisted, expired, MFA-required, and revoked states three times, plus one suppressed-write fault. That made 16 runs per route; none belongs to the earlier Playwright 1.63.0 fixture. No model, real account, or third-party login was involved.

The useful task was to enter an authorized session, read a code from a same-origin iframe, fill a request whose owner selector lived in an open shadow root, submit after a button rerender, and verify the saved record after reload. During request-capable states the page also emitted a nonblocking console error and HTTP 503. The Playwright log recorded a separate resource 404 without a URL, so we cannot identify it as a favicon request. We retained those signals; they were not evidence that the request failed. A separate server ledger, not the green Saved message, decided whether the request existed.

Playwright local Request Desk run with iframe guide, Mira as owner, High priority, Saved message, and request ID
Native-size crop from t1-pw-clean, Playwright 1.59.1. The iframe code, fields, and receipt are readable on a phone; the server ledger records the same request.

Open the full Playwright receipt screenshot

Cypress local Request Desk run with the same iframe guide, owner, priority, Saved message, and request ID
Native-size crop from t1-cy-clean, Cypress 16.0.0. The same fixture and input produced the same verified fields without a route-specific success rule.

Open the full Cypress receipt screenshot

For persisted state, Playwright wrote a synthetic cookie to a temporary storageState file and loaded it in a new context; Cypress cached and restored the state with cy.session(). Neither route made a new login request in those six persisted runs. For expired state, both detected the server rejection and logged in again before submitting. The following counts are local case outcomes, not production success rates.

The mechanisms are documented in Playwright's authentication guide and Cypress's cy.session() guide. Our server ledger tested one concrete implementation of those mechanisms; the documentation does not guarantee that any particular live account remains valid.

// Playwright 1.59.1: save and load only the synthetic test state.
await seedContext.storageState({ path: stateFile });
const context = await browser.newContext({ storageState: stateFile });

// Cypress 16.0.0: the second call restores the cached session.
cy.session([runId, state], () => cy.visit(seedUrl));
cy.session([runId, state], () => cy.visit(seedUrl));
cy.visit(appUrl);
Session statePlaywright 1.59.1Cypress 16.0.0Server check
Clean3/3 completed3/3 completedOne login and one write per run
Persisted3/3 completed3/3 completedNo new login; one write per run
Expired3/3 recovered3/3 recoveredExpired detected, then one login and write
MFA required3/3 stopped3/3 stoppedZero submissions and writes
Account revoked3/3 stopped3/3 stoppedZero submissions and writes
Write suppressed1/1 false success detected1/1 false success detectedSaved appeared; zero writes

The 18 runs expected to create a request produced 18 server records. We checked subject, owner, priority, and iframe code against every record, for 72 matching fields out of 72 checked. The fresh page also showed the saved values, as in t1-pw-clean below. MFA and revoked were synthetic stop screens, not a real OTP challenge or account takeover; both routes stopped before a write.

Detail crop of the fresh Playwright request showing its ID, subject, owner, priority, and iframe code
Original-size crop from t1-pw-clean after reload. Swipe horizontally on a phone to read all four fields; the full original and matching server row are retained in the evidence package.

Open the full t1-pw-clean screenshot

Cypress local fixture sign-in page showing mfa_required and no request form
Native-size crop from t1-cy-mfa. The synthetic MFA wall prevented the request; its server ledger contains one login attempt and no submit or saved record.

Open the full Cypress MFA screenshot

Playwright local fixture sign-in page showing account_revoked and no request form
Native-size crop from t1-pw-revoked. Revocation stopped the synthetic login; the server ledger contains no submission or saved record.

Open the full Playwright revoked screenshot

The injected write fault shows why a Toast cannot be the final assertion. In one run per route, the page said Saved but the server intentionally wrote nothing. Both test scripts reloaded the page and detected the missing record; the independent ledger also held zero items. This was an expected fault detection, not a completed request.

Detail crop of a green Saved message with DRAFT below it from a controlled Cypress write fault
Detail crop from fault-cy-suppressed. The UI displayed Saved and DRAFT, but reload and the server ledger found no record, so the test classified the message as false success.

Open the full fault-cy-suppressed screenshot

A separate Cypress 16 pilot failed before its test began because Cypress.env() had been removed. We changed the test to cy.env(), kept a standalone reproduction of the error, and only then started the 32 formal runs. That is a version migration detail, not a failure counted against Cypress in the session table. The experiment plan, exact scripts, 32 raw ledgers, screenshots, failed pilot, and independent review are retained with the first-party evidence for this article.

What did the live-site Playwright pass add?

The archived September 14 Playwright screenshot shows an IKEA US search for desk, the White filter, Price: low to high, and IKEA's 111-item count at that moment. A separate product screenshot shows the TORALD detail page. The screenshot pair does not retain the full set of loaded cards or every navigation step, so we no longer claim a verified ranking of all qualifying desks from this older run.

IKEA search detail from the Playwright screenshot showing 111 items, White filter, and Price low to high selected
Native crop of the archived Playwright view. The page shows 111 items, the White filter, and Price: low to high. The full screenshot includes the controller.

Open the full Playwright and IKEA screenshot

The product-detail image shows TORALD. Its price label is clipped at the edge of this older capture, so we do not use that image alone to establish the full price. The archived images do not independently establish the complete return-state assertions or the full filtered inventory.

Open the full archived Playwright detail screenshot

What happened in the live-site Cypress pass?

The archived Cypress Runner screenshot shows one passing test, the White filter, Price: low to high, and visible command log beside the IKEA results. Its product-detail screenshot shows TORALD at $29.99. The complete script and raw test log are absent from this archive, so these images do not establish all return-state assertions or an inventory ranking.

Native crop of the Cypress IKEA page showing Price low to high selected and White checked
The archived Cypress page has Price: low to high selected and White checked. The full screenshot also shows the Runner and its one passing test.

Open the full Cypress Runner screenshot

The Runner screenshot makes the applied filter and selected sort inspectable. It does not preserve the earlier click attempts or the final spec, so we cannot use this older image to quantify authoring difficulty.

Cropped IKEA TORALD product detail from the Cypress Runner showing the desk and $29.99 primary price
Native crop of the archived Cypress TORALD detail page. The name and $29.99 primary price are visible; the complete return path is not archived.

Open the full Cypress detail screenshot

What changed when we used ego (lite)?

The archived ego (lite) 0.5.0.31 screenshot shows a natural-language IKEA request in Claude Code beside a visible Space with Agent is in control, Take over, and Stop controls. The second image shows the TORALD detail page in that Space. These images show the interaction model, while the complete action history is not part of this older archive.

Cropped ego (lite) Space control bar showing Agent is in control, Take over, and Stop
Native crop of the visible Space controls. The full archived frame also shows the natural-language request in Claude Code.

Open the full Claude Code and ego (lite) screenshot

The later image shows TORALD at $29.99 with the same visible Space controls. It does not preserve the full card inventory or prove every search, filter, sort, and return step, so we do not use it as a completed-task or speed score.

Cropped IKEA TORALD product detail from the ego (lite) Space showing the desk and $29.99 price
Native crop of the archived TORALD detail view. The full screenshot also shows the Space's agent state and takeover controls.

Open the full ego (lite) detail screenshot

That difference is the practical decision point. Playwright and Cypress ask you to encode a durable, repeatable test. ego (lite) lets you start with an outcome in natural language and supervise the agent in a visible Space. That is useful for one-off or changing, authorized browser work where writing and maintaining a full test spec would cost more than the task. It is not a substitute when the result must become a deterministic CI gate.

How do debugging and CI compare?

Cypress's interactive runner is unusually good at showing the command sequence and captured DOM state while a developer works in the browser. Playwright's trace viewer reconstructs actions, DOM snapshots, network activity, console output, attachments, and timing after or during a run. Which feels better depends on whether your team debugs primarily in a live runner or from retained CI artifacts.

Both can parallelize and produce CI artifacts. Compare the open-source runners separately from optional paid dashboards, hosted orchestration, analytics, or test-impact products. Your CI decision should include worker startup, browser caching, sharding, artifact retention, quarantine policy, and how quickly a developer can reproduce a failed job locally.

Which framework should you choose?

Choose Playwright when any of these are non-negotiable:

  • The same journey must coordinate popups, several pages, or several browser contexts.
  • A cross-origin iframe is part of the release-critical flow.
  • Chromium, Firefox, and WebKit coverage should use one integrated runner and API.
  • Trace-first failure analysis and highly isolated parallel workers fit the CI model.

Choose or keep Cypress when these are more important:

  • The application is primarily a single-tab frontend and its critical flows fit the documented origin model.
  • Component testing and an in-browser command log are central to the team's daily feedback loop.
  • The current Cypress suite is healthy, trusted, and cheaper to maintain than replace.
  • The team already has stable fixtures, custom commands, CI dashboards, and debugging habits around Cypress.

Session reuse alone did not separate the frameworks in our local test: both reused a valid session, recovered from an expired one, and stopped at synthetic MFA or revocation. Choose from the browser workflow and maintenance constraints above, then verify the server state after every write. A real account that needs human approval changes the operating workflow; this fixture did not test that handoff.

How do you migrate from Cypress to Playwright?

  1. Inventory capabilities before syntax. Mark every use of cy.origin(), cy.intercept(), custom commands, sessions, tasks, component mounts, plugins, and cloud-only features.
  2. Select five to ten journeys that represent login, data setup, network behavior, frames, downloads, popups, and the slowest CI path.
  3. Rebuild state boundaries. Map Cypress hooks and cy.session() to Playwright fixtures, projects, browser contexts, and storageState without sharing mutable accounts across workers.
  4. Translate intent, not chaining syntax. Prefer role, label, and test-id locators; replace implicit subject chains with named locators and explicit assertions.
  5. Match network semantics. Decide whether each intercept is observing, waiting, stubbing, or mutating, then implement the equivalent route or response wait.
  6. Run both suites against the same build and seed data. Compare uncovered requirements, failure causes, authoring time, triage time, and infrastructure work.
  7. Migrate only after the representative slice meets an agreed reliability and debugging contract. Keep a rollback window instead of rewriting the entire suite at once.
// Cypress
cy.contains("button", "Load products").click();
cy.get('[role="status"]').should("have.text", "Ready");

// Playwright
await page.getByRole("button", { name: "Load products" }).click();
await expect(page.getByRole("status")).toHaveText("Ready");

Where does ego (lite) fit?

The IKEA run makes the choice concrete. Playwright and Cypress are the better tools when a team needs a versioned spec, exact assertions, repeatable fixtures, and a CI result. ego (lite) is the better fit when the work starts as a natural-language goal, the path may change as the real site responds, and a person should be able to watch or take over the authorized browser Space.

In this case, ego (lite) completed the same search, filter, sort, inspection, detail check, and return-state check without a reusable selector-heavy test file. That reduces setup for bounded, one-off work, but it does not remove the need for judgment or turn the run into regression coverage. Keep release assertions, fixtures, isolation, and CI gates in Playwright or Cypress. Do not automate destructive account actions, bypass MFA, or reuse a personal profile where a dedicated test account is the safer boundary.

What are the limits of this comparison?

The earlier Playwright 1.63.0 fixture and September 14 IKEA work are represented here only by archived screenshots; their original per-run logs and complete scripts are unavailable, so we removed the earlier pass counts and working-time ranking. The September 28 session lab is a separate, auditable synthetic-account fixture using Playwright 1.59.1 and Cypress 16.0.0. It did not measure a production suite, WebKit, Firefox, component testing, visual regression, paid cloud services, memory use, or large-scale parallelism. Verify current documentation and rerun the relevant task before a high-cost migration.

Which official sources support the comparison?

For Playwright, verify the current documentation for pages and popups, actionability, network control, and test retries.

For Cypress, use its current documentation for retry ability, cy.origin(), and its documented cross-origin and multi-tab boundaries. The earlier comparison sources and package versions were checked on September 11, 2026.

For the September 28 session lab, we rechecked Playwright's context isolation, Cypress's test isolation, and the Cypress 16 migration guide. The archived experiment includes the exact installed versions and the failed migration pilot.

FAQ

Is Cypress flaky?

Not inherently. Cypress retries queries and assertions, but bad selectors, shared data, external dependencies, and incorrect command boundaries can still create flaky tests.

Can Cypress test multiple domains?

Yes, for supported top-level origin transitions using cy.origin(). That support does not extend to a cross-origin iframe, and Cypress does not run commands in a second tab or window.

Can Cypress reuse a login session?

Yes. cy.session() can cache and restore browser session data. In our own local fixture it restored an authorized synthetic session without another login in three runs. The application still decides whether the session is valid; our expired runs needed a new login, while MFA and revoked runs stopped without submitting.

Does Playwright retry every failed action?

No. Playwright auto-waits before actions and retries web-first assertions. Whole-test retries are a separate configuration. A side-effecting action is not blindly replayed because a later assertion failed.

Should an existing Cypress suite migrate?

Only when a representative pilot proves that a real capability, maintenance, coverage, or debugging gap is worth the rewrite and infrastructure cost.