MCP-first browser primitives for AI agents — real eyes and hands on the web. Local-first. Vendor-agnostic. Yours to own.
PixelCheck provides 14 tools with basic JSON Schema definitions and short descriptions. Most tools have input schemas with parameter types, but descriptions are terse (often under 50 chars), parameter annotations are minimal, and output schemas are not documented. Tools follow a clear naming convention (verb_noun pattern), but lack the depth of documentation and error-handling guidance expected for production-grade agent tools. The server implements a TypeScript-based tool registry (ToolRegistry, ToolDefinition pattern visible in server.ts), but individual tool files show incomplete parameter documentation and no explicit output schema definitions in the visible source. Composition is reasonable (distinct concerns per tool), but descriptions don't guide LLMs toward when/why to call each tool.
execute a sequence of actions
full audit pipeline against a URL
run the critic calibration gate
A/B page comparison
holistic page-health diagnosis
diagnose / self-heal the environment
autonomous agent run with a goal
schema-bound structured extraction
read the most recent audit summary
Tool descriptions are uniformly terse (15 - 50 chars). Pattern baseline is 194 chars average; these descriptions omit WHEN to use each tool, prerequisites, and how to chain with others. LLMs cannot disambiguate similar tools (audit_url vs diagnose, see vs explore_url) based on current descriptions alone.
Input parameter descriptions are absent or generic. E.g. 'url' param in 'see' tool has description 'URL to navigate to' but lacks context: what happens to navigation state? Is it stateless? Do cookies persist? Parameter baseline is 72 chars avg; most are 10 - 20 chars here.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 46 | <=2025-11-25 | v2 |
rubric-driven page critic
self-describe the MCP server
enumerate installed personas
enumerate installed scenarios
one-shot navigation snapshot
No output schemas are visible in the provided source. Tool definitions show input schemas (InputSchema with properties/types) but do not document what fields or types are returned. LLMs cannot plan downstream tool calls or extract required IDs without knowing response structure.
No error handling or recovery guidance. Tools like 'explore_url' and 'act' (WRITE risk) do not document failure modes, retryability, or what the LLM should do if navigation fails, action execution fails, or the page is unreachable.
Parameter 'instruction' in 'act' and 'judge' tools is extremely vague. Baseline guidance: describe expected format, range, constraints. No hints on valid action syntax, goal phrasing, or rubric format. LLMs will guess and fail.
The 'extract' tool accepts a 'schema' parameter of type 'object' with description 'JSON schema for the extracted data', but no constraints on what that schema can contain, size limits, complexity, or format validation. LLMs may pass malformed schemas.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). WRITE tools 'explore_url', 'act', and 'doctor' are not explicitly marked as destructive or non-idempotent, forcing LLMs to reason about side effects and retry safety without guidance.
Composition gap: tools like 'see', 'act', and 'extract' operate on a browser instance, but it is unclear how state is managed across calls. Does each 'see' start fresh? Can 'act' modify the page prepared by 'see'? Parameter docs omit session/context info.