ArtStroy logo
ArtStroy qa · ai · engineering
Programming · October 8, 2026 · 9 min read

Playwright 1.64: WebMCP, Locator.within(), and Smarter Runs

Playwright 1.64 adds WebMCP tool testing, locator.within(), video fps/VP9, project defaults, --shuffle, and breaking visibility rules. What to adopt first in a real suite.

Playwright 1.64.0 dark dashboard hero with WebMCP, within(), video, shuffle, and default projects panels

A release that only adds locators is easy to skim. Playwright 1.64 is messier in a useful way: pages can register tools through experimental WebMCP, locators can be composed with within(), video artifacts get real styling controls, and the runner finally admits that not every configured browser should run on every local npx playwright test.

This continues the version track after Playwright 1.63. Locks are already covered there. Here the interesting parts are how tools on a page become testable, how nested locators stop fighting tables, and which config defaults change what “green” means in CI.

npm install -D @playwright/test@1.64.0
npx playwright install

WebMCP: test the tools the page exposes

WebMCP is an experimental browser API. A page registers tools (think navigator.modelContext), and Playwright reaches them through page.webmcp / frame.webmcp instead of only driving the UI that wraps those tools.

Chromium needs an explicit launch flag today:

import { chromium } from '@playwright/test';

const browser = await chromium.launch({
  args: ['--enable-features=WebMCP'],
});
const page = await browser.newPage();
await page.goto('https://example.com');

for (const tool of await page.webmcp.tools()) {
  console.log(tool.name, tool.description);
}

const result = await page.webmcp.callTool('add', { a: 2, b: 40 });
console.log(result.content[0].text); // "42"

Firefox uses prefs (dom.modelcontext.enabled and the testing flag). WebKit does not implement it yet. Treat that as a product constraint: WebMCP coverage is Chromium-first until the others catch up.

Why care as Automation QA? Because agent-facing UIs will ship capabilities that never appear as a button. If the contract is “call search_catalog with this schema,” asserting the tool result is tighter than clicking through a chat chrome that will change next sprint.

Playwright MCP and playwright-cli surface the same tools by default (webmcp_* for agents; webmcp-list / webmcp-call on the CLI). Opt out with --no-webmcp when you do not trust page-defined schemas in an agent session. That sits next to the workflow discussion in Driving a Browser With an Agent Is Still Specification: agents still need a contract; WebMCP just makes the contract addressable.

locator.within(): say the nesting in reading order

You already know how to scope with a parent locator. within() flips the sentence so it matches how you describe the UI: “this Save button, within that settings dialog.”

const saveButton = page.getByRole('button', { name: 'Save' });
const dialog = page.getByTestId('settings-dialog');

await saveButton.within(dialog).click();

The win is not shorter code. It is reusable locator pieces. Keep saveButton and dialog as named constants in a page object, compose at the call site, and avoid rebuilding every combined selector as a one-liner.

Relative helpers resolve per parent. That is the table case people keep reinventing with fragile CSS:

// Third cell of every row — not "the third cell somewhere in the table"
const thirdColumn = page
  .getByRole('cell')
  .nth(2)
  .within(page.getByRole('row'));

await expect(thirdColumn).toHaveText(['Apple', 'Banana', 'Cherry']);

I would reach for within() anywhere a component owns a repeated child pattern: columns, card grids, list rows with an action in each. Keep classic parent.locator(...) when the parent is the only thing you care about naming.

Video that looks like a demo, not a blur

Video used to be “on or off, default decorations.” 1.64 treats it like a product artifact.

import { defineConfig } from '@playwright/test';

export default defineConfig({
  use: {
    video: {
      mode: 'retain-on-failure',
      size: { width: 1280, height: 720 },
      fps: 25,
      show: {
        actions: {
          style: {
            point: 'width: 16px; height: 16px; border-radius: 50%; background: #e11',
            highlight: 'outline: 2px solid #0ea5e9; background: rgba(14,165,233,.12)',
            title: 'font-size: 14px; font-family: system-ui',
          },
        },
      },
    },
  },
});

What changed in practice:

  • fps — custom frame rate. Firefox and WebKit top out at 25 fps; do not set 60 and expect the same on every project.
  • style — CSS for the action point, highlight, and title. Replaces fontSize (deprecated).
  • Cursor — stays at the last action, survives navigations, moves on an eased path. Failures are easier to narrate in a bug report.
  • VP9 — replaces VP8. Less CPU, smaller files at similar quality. Watch CI artifact budgets; retention policies may need a revisit after the codec change, not before.

Same options land in Playwright MCP and playwright-cli, so agent-recorded sessions can match your suite’s visual language.

Test runner: defaults, shuffle, and quieter WebP

Four runner changes matter for day-to-day CI more than they look on a changelog.

Projects that exist but do not run by default

testProject.default: false keeps Firefox and WebKit in the config without paying for them on every local run:

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  projects: [
    { name: 'chromium', use: devices['Desktop Chrome'] },
    { name: 'firefox', use: devices['Desktop Firefox'], default: false },
    { name: 'webkit', use: devices['Desktop Safari'], default: false },
  ],
});
npx playwright test                     # chromium only
npx playwright test --project=firefox
npx playwright test --project="*"       # all three — use this on CI

Reporters and global setup can read fullConfig.filteredProjects to see what actually survived --project. That is the hook for “skip expensive seed data when only chromium is selected.”

Shuffle to find order bugs

npx playwright test --shuffle
npx playwright test --shuffle 271828182

The first command prints a shuffle seed (for example 271828182). Pass that seed back to reproduce the same order when you file a ticket. I would run shuffle on a nightly job, not on every PR, until the suite is clean of shared state.

Locks on a whole describe

1.63 introduced named locks on individual tests. 1.64 adds the same idea on test.describe.configure():

test.describe.configure({ lock: 'user-settings' });

Use this when an entire file or group shares one resource. For the mental model and the “lock held for the whole file in default mode” trap, start with the 1.63 locks write-up and treat this as the group-level shortcut.

Unnamed screenshots as WebP

export default defineConfig({
  expect: {
    toHaveScreenshot: { type: 'webp' },
  },
});

Named .png snapshots still follow the filename. This config only flips the default for unnamed ones. It pairs with the WebP direction from 1.62.1 without forcing a mass rename.

Smaller APIs worth knowing

  • page.getByRef('e2') — locate by the aria ref from page.ariaSnapshot({ mode: 'ai' }). Useful when an agent or AI helper already speaks in those refs and you want a deterministic follow-up assertion.
  • page.content({ includeShadow: true }) — open shadow roots serialize as declarative Shadow DOM. Handy when a bug only exists under a web component and a plain HTML dump lied to you.
  • signCount on virtual WebAuthn credentials — set, read, and persist through storage state. Passkey flows that assert counter bumps can stop faking that detail.
  • Cookies on APIRequestContext — addCookies(), cookies(), clearCookies() mirror browser context. Setup that was awkward for API-only projects gets a normal cookie jar. If you already reconstruct HTTP history with playwright-api-logger, cookie control on the request context closes another gap between UI and API fixtures.

Breaking changes that will surprise a green suite

Do not upgrade on Friday without reading these four.

1. screen comes from device descriptors. Spreading ...devices['Desktop Chrome'] now feeds real emulated window.screen and media queries. Layout tests that assumed the old screen size can flip. Opt out:

use: {
  ...devices['Desktop Chrome'],
  screen: undefined,
},

2. JSX in tests follows tsconfig.json. Playwright honors jsx, jsxFactory, jsxFragmentFactory, and jsxImportSource, defaulting to the automatic React runtime. Component-test files that relied on Playwright’s previous defaults may need a tsconfig tweak, not a Playwright config one.

3. --update-snapshots=missing no longer fails the run. Missing baselines are written and the run can stay green, so CI can create new shots and verify existing ones in one pass. The old “write missing and fail” behavior is the new 'default' mode of updateSnapshots. Update any job that treated a non-zero exit as “we generated baselines.”

4. Elements inside hidden iframes are hidden. If the iframe has visibility: hidden (or is otherwise not visible), actions, isVisible(), and toBeVisible() treat children as hidden too. Tests that clicked into a pre-mounted but invisible frame will start failing for the right reason.

Bundled browsers jump to Chromium 156, Firefox 157, and WebKit 27.2 (also checked against Chrome 155 / Edge 155). Expect visual and timing drift on top of the API breaks.

What I would adopt first

1. --project="*" on CI, default: false locally. Cheapest win. Developers stay on Chromium; the pipeline still owns the matrix.

2. locator.within() for tables and repeated cards. Refactor page objects that already split “row” and “cell” locators. Leave one-off selectors alone until they break.

3. Nightly --shuffle with a logged seed. Fix the first order-dependent failures before making shuffle a PR gate.

4. WebMCP only where the product ships tools. Write contract tests against callTool for those surfaces. Do not force WebMCP onto a classic CRUD app that has nothing to register.

5. Re-check visual and iframe tests after the screen + hidden-iframe breaks. Those two will produce the loudest red on upgrade day.

Video CSS and VP9 are polish. Turn them on when you already attach videos to tickets; they do not unblock a failing suite.

The question this release is answering

1.63 asked how to keep thousands of tests parallel without lying about shared state. 1.64 asks a different one: can the runner, the locators, and the browser meet the app where it is going — tools on the page, nested UI, and agent-shaped workflows — without making every local run pay for the full matrix.

WebMCP is the headline if you build agent-facing products. For everyone else, within(), project defaults, and shuffle are the parts that will show up in next week’s PR comments. Upgrade when you can spend an afternoon on the screen and iframe behavioral changes. Skip the afternoon and those “random” failures will look like flaky product bugs.


Technical details checked against the Playwright 1.64 release notes. External references kept to the official docs.