A browser is a computer the agent already knows how to use. pixelpi is the thin harness that lets it — six tools, raw CDP, any model.
A new primitive
Established
define → call → parse
A contract. Works only where someone built the endpoint.
New
see → click → read
A pair of hands. Works anywhere a person could.
Six tools
An essay by Harsh Joshi · June 2026
The default way to give a model new abilities is to wire up tools. You write a schema, you call an API, you parse the response. It works, and it scales to exactly the surfaces someone bothered to build an endpoint for. Everything else stays out of reach.
But there is already a universal interface, and the model already reads it. The browser renders the world's software into one consistent surface of text, links, and fields. A person uses it without a manual. A model can too — the page is just a document, and reading documents is the one thing these models are unambiguously good at.
So pixelpi is defined as much by what it leaves out. No Playwright, no headless framework wrapping a framework. No vision model squinting at pixels. No cloud service in the loop. Six tools, not thirty — look, act, fill, nav, eval, store — sitting directly on raw Chrome DevTools Protocol.
The expensive part of driving a browser was never the clicking. It was the looking — dumping a full DOM into the context window on every step. So look() returns a compact, structured view instead, around 107x cheaper than a raw-DOM dump on heavy pages. The agent sees what matters and nothing else, and the whole loop stays under 1,100 tokens of prompt and tools.
The bet is simple. You do not need a new protocol for every site, and you do not need a bigger model to read one. You need to hand the agent the interface humans already won, and get out of the way. The page is the prompt.