Skip to main content
You make two choices before you write any automation. They’re independent, but the first constrains the second:
  1. How you drive the browser — the control surface your code or model uses to act on the page.
  2. Where the loop runs — the machine your decision-making code runs on, relative to the browser.

1. How you drive the browser

Kernel browsers accept four control surfaces. Pick by what’s driving the page, not by what you already know. For agents, start with playwright execution with a computer use fallback: script the deterministic steps, and hand the page to a computer use model when a step doesn’t respond to a selector.

Why the choice matters on hardened sites

CDP is what Playwright and Puppeteer speak, and anti-bot vendors scan for its signatures. Computer controls carry no CDP connection, so there’s no protocol fingerprint to leak. That makes them the stronger option on sites with aggressive detection, and it’s why managed auth drives logins with coordinate-based input rather than CDP. How much this matters is site-specific, so test before you commit — see bot anti-detection.

Control surface examples

Kernel’s Computer Controls API exposes OS-level mouse, keyboard, and screen primitives — the surface a computer-use model already knows how to drive (screenshot, click, type, key, scroll, drag). No CDP or WebDriver connection required, so there’s no protocol fingerprint to leak. Ideal for Claude, OpenAI, or Gemini computer-use loops.

2. Where the loop runs

Your loop is whatever decides the next action: a script, an agent, or a model. It can run in three places.
Connect to cdp_ws_url or webdriver_ws_url from wherever your code already runs. Any CDP client works, and there’s no lock-in.Costs: a network round trip per action, disconnects to handle, screenshot and DOM bandwidth, and the CDP fingerprint above. It’s fine for low-frequency or deterministic work, and it hurts most in a vision loop.

Where computer use fits

A computer use agent answers the first question, not the second — it still has to run its loop somewhere. Because every turn ships a screenshot instead of a small script, running that loop off-platform costs far more than it does for a Playwright-driven agent: you pay image bandwidth and a round trip on every step. That makes computer use the strongest case for running your loop next to the browser. Model inference stays with the model vendor either way.

Putting it together

Why computer use for agents

Kernel’s computer controls are built to match how computer-use models were trained — the same primitives the model emits (screenshot, click at coords, type, key, scroll, drag) map 1:1 onto the API. There’s no harness translating model output into framework calls.
  • Native fit. Screenshot, click, type, key, scroll, drag — the primitives the model already speaks.
  • Faster screenshots. Captures bypass CDP, which removes the largest source of latency in a vision loop.
  • Better against bot detection. No CDP connection means no CDP fingerprint to leak. Pairs naturally with stealth mode and residential proxies.
  • Human-like input. OS-level events with Bézier-curve mouse paths, variable typing speed, and configurable mistype rate.
  • Not DOM-limited. Screenshots capture the full VM, so the agent can see and interact with native dialogs, canvas elements, iframes, and PDFs — not just things you can address with a selector.

Why playwright execution over a direct CDP connection

If you’re reaching for Playwright, prefer the execution API over connectOverCDP. Same Playwright API you already know, none of the setup.
  • Run from anywhere. No playwright package to version-pin, no Chromium download, no CDP connection to manage. Send the code, get the result.
  • Co-located with the browser. Code runs in the same VM as the browser — no network hop between your script and the page, fewer flakes.
  • Patchright by default. Hardened against bot detection out of the box.
  • Full Playwright API. page, context, and browser are all in scope. Anything Playwright can do — DOM queries, file uploads, full-page screenshots — works here.
  • Returns values. return from your code and the result comes back in the response. Easy to use as an agent tool.

Computer use + playwright execution

Computer controls drive the browser the way a person would — they don’t speak the programmatic API surface. Anything you’d reach for the DOM or Playwright client for (reading text and attributes, page.goto, file uploads, cookie or storage access, switching tabs) belongs on the playwright execution side. When computer use is driving, expose playwright execution to the agent as a tool it can call for structured data or a programmatic action. For the full pattern in the other direction — playwright execution first, computer use when a step doesn’t respond to a selector — see playwright with computer use fallback.

Going deeper