Model
The model decides the next action from what it’s given: the system prompt, the conversation so far, tool results, and screenshots. Claude, GPT, and Gemini are general-purpose models. Computer use models are trained to act on screenshots with clicks and keystrokes, so they can drive a page without selectors. The model never touches the browser. It returns text or a tool call, and something else carries it out.Agent framework or harness
The agent framework runs the loop: it sends context to the model, runs the tool call the model returns, feeds the result back, and repeats until the task is done. “Harness” usually means the loop plus everything configured around it: the system prompt, the tools, and the skills. You either build with a framework, such as the Claude Agent SDK, the Vercel AI SDK, or Browser Use, or you use a finished agent that already has one, such as Claude Code or Replit, and give it a browser.System prompt
The system prompt is the instruction set the harness sends with every request to the model: its role, its rules, and the format its answers should take. Because it’s sent on every turn, it’s the place for guidance that applies to every task, such as “log out before you finish” or “never submit a payment without confirming the total”.Tools
A tool is a function the model can ask the harness to call. Each tool has a name, a description, and an input schema. The model only chooses the tool and its arguments; the harness runs it and returns the result. For a browser agent, tools look like “navigate to this URL”, “click at these coordinates”, “take a screenshot”, or “run this Playwright code”. KERNEL provides several ready-made ones:- Playwright execution runs a snippet of Playwright code inside the browser’s vm and returns the result.
- Computer controls click, type, scroll, and take screenshots at the operating-system level.
- The MCP server exposes KERNEL’s API as tools to any client that speaks the Model Context Protocol (MCP), such as Claude, Cursor, or Codex.
Skills
A skill is a packaged set of instructions and reference files that an agent loads only when a task needs it. That’s the difference from a system prompt, which is sent on every turn. Tools let an agent do something; skills teach it how and when to do it well. KERNEL publishes Agent Skills that teach coding agents the KERNEL CLI, the SDKs, bot detection, and authentication, so they don’t have to re-read the docs every session.Browser automation framework
A browser automation framework turns an action like “click the submit button” or “fill in this form” into browser protocol commands, and sends them to the browser over the Chrome DevTools Protocol (CDP) or WebDriver BiDi.- Deterministic: Playwright and Puppeteer do exactly what your code says, using selectors you write.
- AI-assisted: Stagehand and Agent Browser add natural-language actions on top, so the model can describe what to do instead of naming a selector.
Browser infrastructure
Browser infrastructure is where the browser runs and what it carries. KERNEL creates each chromium browser in its own vm, and around it handles what agents need on real websites: stealth and proxies, profiles, logins, payments, and live view, replays, and telemetry for debugging. Everything above this layer is interchangeable. You can switch models, frameworks, or automation libraries and keep the same browsers, profiles, and credentials.Which piece to change
For the objects you’ll work with in KERNEL’s API, such as browsers, browser pools, apps, and invocations, see KERNEL objects.