A browser automation tool that uses AI to process natural language commands and execute browser actions through chromedp, with support for fingerprinting, stealth mode, and API server interface
This MCP server exhibits foundational deficiencies across definition quality. Tool descriptions are present but minimal (6-20 chars in most cases). Input schemas are defined with JSON Schema but descriptions for parameters are sparse or absent. The server claims 11 tools but provides only file paths without explicit tool registration visible in the source code. Naming follows verb_noun conventions (navigate, click, type_text) but lacks clarity in some cases (extract_links vs search_google). No output schemas are documented. Error handling is minimal, the API server returns generic HTTP status codes without actionable recovery guidance. Security posture is weak: the AI client appears to directly call APIs without credential isolation. The codebase shows a functional HTTP API but lacks the polish and documentation required for production agent tooling.
Click on an element
Extract all links from the page or specific selector
Extract text content from the page
Get text content from elements
Navigate to a URL
Press a keyboard key (Enter, Tab, Escape, etc.)
Take a screenshot
Scroll the page
Tool descriptions are critically short (6-30 characters), below the 34-character p10 baseline for production tools. Descriptions like 'Click on an element' and 'Wait for a specified duration' lack context for LLM selection, prerequisites, or when to use vs alternatives.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract returned data structure. navigate returns {'url': string} per code, but this is inferred, not documented in tool definition.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Search Google for information
Type text into an input field or textarea
Wait for a specified duration
Parameter descriptions are missing or minimal. The 'selector' parameter (used in click, type_text, extract_links, extract_text, get_text) lacks a description explaining expected CSS selector syntax, required format, or examples. Parameter descriptions are 0 characters.
No input validation or constraints on numeric/enum parameters. The 'seconds' parameter in wait accepts any number; the 'direction' enum in scroll is defined but no validation guidance. No minimum/maximum bounds.
Error handling is generic. The api/server.go returns HTTP status codes (400, 500) with minimal messages ('Invalid request', error.Error()). No recovery guidance. No categorization (retryable, user-fixable, fatal). LLMs cannot self-correct or plan recovery.
No tool annotations present (readOnlyHint, destructiveHint, idempotentHint). Tools are marked with Risk labels in the specification (WRITE, READ_ONLY) but not visible in schema output or handled via tool metadata.
Ambiguous tool names create LLM confusion. Both 'extract_links' and 'search_google' produce similar outcomes; 'extract_text' and 'get_text' appear to do the same thing. No documented distinction in when to call each.
Tool definitions inferred from file locations and brief descriptions; no explicit tool registration code visible. Tools are listed with filenames (ai/lmstudio.go, ai/ollama.go) but the registration mechanism is not shown in provided source. Cannot verify schema validation or runtime behavior.