
Short answer: Playwright and Cypress turn known browser workflows into repeatable tests; ego (lite) is a separate visible browser route when an Agent must work through a changing live task. Playwright provides broader control over multiple pages, iframes, browser contexts, and browser engines, while Cypress places more emphasis on its interactive Runner and tightly integrated debugging experience. The two test frameworks fit stable processes that can be described in advance with selectors, actions, and assertions.
This shared assumption matters more than many feature-by-feature comparisons. When a workflow must run repeatedly and the application is controlled by the team, deterministic browser testing is the right fit. But when the website interface, execution environment, or task objective can change during the run, maintaining the test script can sometimes take more effort than completing the task itself, which is where a full Chromium like ego (lite) lets an agent work against the live page, cross-origin iframes and all.
Is Playwright better than Cypress?
For a greenfield end-to-end suite, Playwright is usually the safer default because its browser-context, page, frame, tracing, and parallel-worker primitives cover more application shapes without restructuring the product under test. That recommendation changes when the team depends on Cypress component testing, already has a healthy Cypress suite, or gets more value from the Cypress runner than it would gain from Playwright's broader browser-control model.
| Constraint | Prefer Playwright | Prefer Cypress |
|---|---|---|
| Several tabs or popups | Pages are first-class objects | Usually redesign the test into one controlled tab |
| Cross-origin iframe | Frame locators can target it | Outside cy.origin() support |
| Component feedback loop | Supported, but project fit varies | A core Cypress workflow |
| Existing healthy suite | Migrate only for a measured gap | Keep it if constraints are met |
How do their execution models differ?
Playwright's test process controls browsers through its automation protocol. A test creates a browser, one or more isolated contexts, and pages inside those contexts. New pages and popups remain addressable, and each context can carry its own cookies, permissions, and storage. The test and the application are separate processes, which makes cross-page orchestration explicit.
Cypress coordinates a Node process, a proxy, and code running with the application in the browser. That architecture powers its live command log, DOM snapshots, time-travel debugging, and direct application feedback. It also explains why Cypress keeps control in one primary browser tab and asks tests to enter a second top-level origin through cy.origin().
Architecture is not a quality score. It is a constraint map. A checkout that embeds a cross-origin payment frame and opens a receipt window places different demands on a runner than a single-origin dashboard with rich component tests.
Which browser workflows can each test?
| Workflow | Playwright | Cypress |
|---|---|---|
| Multiple tabs/windows | Direct page and popup events | No commands in another tab/window; keep the flow in one tab or test the destination separately |
| Top-level origin change | Navigate and locate normally | Use cy.origin() for commands on the secondary origin |
| Cross-origin iframe | Use frameLocator or contentFrame | Not handled by cy.origin() |
| Network observation | Events, routing, and response waits | cy.intercept(), aliases, waits, and stubs |
| Isolated sessions | Several browser contexts in one browser | Test isolation resets state between tests; session caching can restore selected setup |
The important correction is that Cypress is not simply ‘single-origin.’ Current Cypress supports top-level cross-origin testing with cy.origin(). The narrower boundaries are that cy.origin() does not operate a cross-origin iframe, and Cypress does not execute commands in a different browser tab or window. Those distinctions should drive fixture design and tool choice.
How do waiting and retries differ?
Playwright separates actionability waiting, assertion retry, and whole-test retry. Before actions such as click, it checks conditions such as visibility, stability, event reception, and enabled state. Its web-first assertions retry until they pass or time out. Test retries are a separate runner setting and were disabled in our experiment.
Cypress chains queries and assertions and reruns that linked query chain while the assertion can still succeed. A non-query command such as click executes once; Cypress does not replay every preceding command because a later assertion failed. Whole-test retries are also separate and default to zero unless configured.
What happened in our controlled test?
The earlier article used a two-origin fixture with delayed product data, duplicate labels, a replaced status node, an intentional HTTP 503, a same-origin iframe, a cross-origin payment iframe, and a receipt popup. Its archived screenshot shows the fixture, but the September 11 per-run logs and scripts are not present here. The official pages and cross-origin documentation still explain the browser-control difference; the older screenshot alone cannot establish a pass rate.

Open the full archived fixture screenshot
What changed when the login session changed?
On September 28, 2026, we ran a separate session experiment against our own local Request Desk. Playwright 1.59.1 and Cypress 16.0.0 used the same Chrome 153 browser, input, server, and zero-retry rubric. Each route ran clean, persisted, expired, MFA-required, and revoked states three times, plus one suppressed-write fault. That made 16 runs per route; none belongs to the earlier Playwright 1.63.0 fixture. No model, real account, or third-party login was involved.
The useful task was to enter an authorized session, read a code from a same-origin iframe, fill a request whose owner selector lived in an open shadow root, submit after a button rerender, and verify the saved record after reload. During request-capable states the page also emitted a nonblocking console error and HTTP 503. The Playwright log recorded a separate resource 404 without a URL, so we cannot identify it as a favicon request. We retained those signals; they were not evidence that the request failed. A separate server ledger, not the green Saved message, decided whether the request existed.

Open the full Playwright receipt screenshot

Open the full Cypress receipt screenshot
For persisted state, Playwright wrote a synthetic cookie to a temporary storageState file and loaded it in a new context; Cypress cached and restored the state with cy.session(). Neither route made a new login request in those six persisted runs. For expired state, both detected the server rejection and logged in again before submitting. The following counts are local case outcomes, not production success rates.
The mechanisms are documented in Playwright's authentication guide and Cypress's cy.session() guide. Our server ledger tested one concrete implementation of those mechanisms; the documentation does not guarantee that any particular live account remains valid.
// Playwright 1.59.1: save and load only the synthetic test state.
await seedContext.storageState({ path: stateFile });
const context = await browser.newContext({ storageState: stateFile });
// Cypress 16.0.0: the second call restores the cached session.
cy.session([runId, state], () => cy.visit(seedUrl));
cy.session([runId, state], () => cy.visit(seedUrl));
cy.visit(appUrl);| Session state | Playwright 1.59.1 | Cypress 16.0.0 | Server check |
|---|---|---|---|
| Clean | 3/3 completed | 3/3 completed | One login and one write per run |
| Persisted | 3/3 completed | 3/3 completed | No new login; one write per run |
| Expired | 3/3 recovered | 3/3 recovered | Expired detected, then one login and write |
| MFA required | 3/3 stopped | 3/3 stopped | Zero submissions and writes |
| Account revoked | 3/3 stopped | 3/3 stopped | Zero submissions and writes |
| Write suppressed | 1/1 false success detected | 1/1 false success detected | Saved appeared; zero writes |
The 18 runs expected to create a request produced 18 server records. We checked subject, owner, priority, and iframe code against every record, for 72 matching fields out of 72 checked. The fresh page also showed the saved values, as in t1-pw-clean below. MFA and revoked were synthetic stop screens, not a real OTP challenge or account takeover; both routes stopped before a write.

Open the full t1-pw-clean screenshot

Open the full Cypress MFA screenshot

Open the full Playwright revoked screenshot
The injected write fault shows why a Toast cannot be the final assertion. In one run per route, the page said Saved but the server intentionally wrote nothing. Both test scripts reloaded the page and detected the missing record; the independent ledger also held zero items. This was an expected fault detection, not a completed request.

Open the full fault-cy-suppressed screenshot
A separate Cypress 16 pilot failed before its test began because Cypress.env() had been removed. We changed the test to cy.env(), kept a standalone reproduction of the error, and only then started the 32 formal runs. That is a version migration detail, not a failure counted against Cypress in the session table. The experiment plan, exact scripts, 32 raw ledgers, screenshots, failed pilot, and independent review are retained with the first-party evidence for this article.
What did the live-site Playwright pass add?
The archived September 14 Playwright screenshot shows an IKEA US search for desk, the White filter, Price: low to high, and IKEA's 111-item count at that moment. A separate product screenshot shows the TORALD detail page. The screenshot pair does not retain the full set of loaded cards or every navigation step, so we no longer claim a verified ranking of all qualifying desks from this older run.

Open the full Playwright and IKEA screenshot
The product-detail image shows TORALD. Its price label is clipped at the edge of this older capture, so we do not use that image alone to establish the full price. The archived images do not independently establish the complete return-state assertions or the full filtered inventory.
Open the full archived Playwright detail screenshot
What happened in the live-site Cypress pass?
The archived Cypress Runner screenshot shows one passing test, the White filter, Price: low to high, and visible command log beside the IKEA results. Its product-detail screenshot shows TORALD at $29.99. The complete script and raw test log are absent from this archive, so these images do not establish all return-state assertions or an inventory ranking.

Open the full Cypress Runner screenshot
The Runner screenshot makes the applied filter and selected sort inspectable. It does not preserve the earlier click attempts or the final spec, so we cannot use this older image to quantify authoring difficulty.

Open the full Cypress detail screenshot
What changed when we used ego (lite)?
The archived ego (lite) 0.5.0.31 screenshot shows a natural-language IKEA request in Claude Code beside a visible Space with Agent is in control, Take over, and Stop controls. The second image shows the TORALD detail page in that Space. These images show the interaction model, while the complete action history is not part of this older archive.

Open the full Claude Code and ego (lite) screenshot
The later image shows TORALD at $29.99 with the same visible Space controls. It does not preserve the full card inventory or prove every search, filter, sort, and return step, so we do not use it as a completed-task or speed score.

Open the full ego (lite) detail screenshot
That difference is the practical decision point. Playwright and Cypress ask you to encode a durable, repeatable test. ego (lite) lets you start with an outcome in natural language and supervise the agent in a visible Space. That is useful for one-off or changing, authorized browser work where writing and maintaining a full test spec would cost more than the task. It is not a substitute when the result must become a deterministic CI gate.
How do debugging and CI compare?
Cypress's interactive runner is unusually good at showing the command sequence and captured DOM state while a developer works in the browser. Playwright's trace viewer reconstructs actions, DOM snapshots, network activity, console output, attachments, and timing after or during a run. Which feels better depends on whether your team debugs primarily in a live runner or from retained CI artifacts.
Both can parallelize and produce CI artifacts. Compare the open-source runners separately from optional paid dashboards, hosted orchestration, analytics, or test-impact products. Your CI decision should include worker startup, browser caching, sharding, artifact retention, quarantine policy, and how quickly a developer can reproduce a failed job locally.
Which framework should you choose?
Choose Playwright when any of these are non-negotiable:
- The same journey must coordinate popups, several pages, or several browser contexts.
- A cross-origin iframe is part of the release-critical flow.
- Chromium, Firefox, and WebKit coverage should use one integrated runner and API.
- Trace-first failure analysis and highly isolated parallel workers fit the CI model.
Choose or keep Cypress when these are more important:
- The application is primarily a single-tab frontend and its critical flows fit the documented origin model.
- Component testing and an in-browser command log are central to the team's daily feedback loop.
- The current Cypress suite is healthy, trusted, and cheaper to maintain than replace.
- The team already has stable fixtures, custom commands, CI dashboards, and debugging habits around Cypress.
Session reuse alone did not separate the frameworks in our local test: both reused a valid session, recovered from an expired one, and stopped at synthetic MFA or revocation. Choose from the browser workflow and maintenance constraints above, then verify the server state after every write. A real account that needs human approval changes the operating workflow; this fixture did not test that handoff.
How do you migrate from Cypress to Playwright?
- Inventory capabilities before syntax. Mark every use of cy.origin(), cy.intercept(), custom commands, sessions, tasks, component mounts, plugins, and cloud-only features.
- Select five to ten journeys that represent login, data setup, network behavior, frames, downloads, popups, and the slowest CI path.
- Rebuild state boundaries. Map Cypress hooks and cy.session() to Playwright fixtures, projects, browser contexts, and storageState without sharing mutable accounts across workers.
- Translate intent, not chaining syntax. Prefer role, label, and test-id locators; replace implicit subject chains with named locators and explicit assertions.
- Match network semantics. Decide whether each intercept is observing, waiting, stubbing, or mutating, then implement the equivalent route or response wait.
- Run both suites against the same build and seed data. Compare uncovered requirements, failure causes, authoring time, triage time, and infrastructure work.
- Migrate only after the representative slice meets an agreed reliability and debugging contract. Keep a rollback window instead of rewriting the entire suite at once.
// Cypress
cy.contains("button", "Load products").click();
cy.get('[role="status"]').should("have.text", "Ready");
// Playwright
await page.getByRole("button", { name: "Load products" }).click();
await expect(page.getByRole("status")).toHaveText("Ready");Where does ego (lite) fit?
The IKEA run makes the choice concrete. Playwright and Cypress are the better tools when a team needs a versioned spec, exact assertions, repeatable fixtures, and a CI result. ego (lite) is the better fit when the work starts as a natural-language goal, the path may change as the real site responds, and a person should be able to watch or take over the authorized browser Space.
In this case, ego (lite) completed the same search, filter, sort, inspection, detail check, and return-state check without a reusable selector-heavy test file. That reduces setup for bounded, one-off work, but it does not remove the need for judgment or turn the run into regression coverage. Keep release assertions, fixtures, isolation, and CI gates in Playwright or Cypress. Do not automate destructive account actions, bypass MFA, or reuse a personal profile where a dedicated test account is the safer boundary.
What are the limits of this comparison?
The earlier Playwright 1.63.0 fixture and September 14 IKEA work are represented here only by archived screenshots; their original per-run logs and complete scripts are unavailable, so we removed the earlier pass counts and working-time ranking. The September 28 session lab is a separate, auditable synthetic-account fixture using Playwright 1.59.1 and Cypress 16.0.0. It did not measure a production suite, WebKit, Firefox, component testing, visual regression, paid cloud services, memory use, or large-scale parallelism. Verify current documentation and rerun the relevant task before a high-cost migration.
Which official sources support the comparison?
For Playwright, verify the current documentation for pages and popups, actionability, network control, and test retries.
For Cypress, use its current documentation for retry ability, cy.origin(), and its documented cross-origin and multi-tab boundaries. The earlier comparison sources and package versions were checked on September 11, 2026.
For the September 28 session lab, we rechecked Playwright's context isolation, Cypress's test isolation, and the Cypress 16 migration guide. The archived experiment includes the exact installed versions and the failed migration pilot.
FAQ
Is Cypress flaky?
Not inherently. Cypress retries queries and assertions, but bad selectors, shared data, external dependencies, and incorrect command boundaries can still create flaky tests.
Can Cypress test multiple domains?
Yes, for supported top-level origin transitions using cy.origin(). That support does not extend to a cross-origin iframe, and Cypress does not run commands in a second tab or window.
Can Cypress reuse a login session?
Yes. cy.session() can cache and restore browser session data. In our own local fixture it restored an authorized synthetic session without another login in three runs. The application still decides whether the session is valid; our expired runs needed a new login, while MFA and revoked runs stopped without submitting.
Does Playwright retry every failed action?
No. Playwright auto-waits before actions and retries web-first assertions. Whole-test retries are a separate configuration. A side-effecting action is not blindly replayed because a later assertion failed.
Should an existing Cypress suite migrate?
Only when a representative pilot proves that a real capability, maintenance, coverage, or debugging gap is worth the rewrite and infrastructure cost.

