A high-pass filter that sits between chaotic web pages and clean LLM context. Converts 100KB HTML into ~2KB Semantic Markdown with integer IDs for zero-hallucination element targeting.
The server defines 4 tools with explicit schemas and descriptions. However, significant issues reduce quality: (1) descriptions are verbose but lack LLM-optimized concision; (2) parameter descriptions are present but some lack detail on constraints and valid values; (3) output schemas are not documented, the server returns synthesized markdown but the underlying SemanticMap structure is not exposed to the LLM; (4) error handling is basic (generic error messages with limited recovery guidance); (5) no tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk profiles. The naming is clear and action-oriented (get_*, perform_*, etc.), and input schemas use proper Zod validation. The tools are well-composed (each does one thing), but the implementation lacks patterns for pagination, result limiting, and structured output documentation that would help LLMs reason about what to expect.
Re-extract the semantic map without navigating to a new page. Useful for refreshing the page representation after external changes or to verify element IDs.
Capture a PNG screenshot of the current page
Navigate to a URL and extract a filtered semantic map of the page. The task_intent guides what gets included: 'read content' prunes nav/footers, 'fill form' focuses on inputs, 'find X' keeps navigation elements. Returns compact Semantic Markdown with integer IDs for every interactive element.
Execute an action on the current page and return only the semantic diff (what changed). Actions: click/type/select/hover on elements by ID, scroll, press_key, go_back, go_forward, navigate to URL. Returns a concise diff instead of the full page, plus the updated semantic map if significant changes occurred.
No documented output schema. Tools return synthesized markdown and SemanticMap objects, but LLMs are not told what fields to expect in the semantic map structure (nodes, landmarks, totalFilteredNodes, etc.). This forces LLMs to infer the structure and makes downstream tool chaining error-prone.
Missing tool annotations. All four tools have clear risk profiles (read-only vs write), but none are annotated with readOnlyHint or destructiveHint. perform_action modifies state but lacks a destructiveHint flag, preventing safety-aware clients from enforcing approval gates.
Error handling lacks recovery guidance. Error messages like 'Error: ${msg}' or 'No active session' do not tell the LLM what to do next. E.g., when id not found, the error says 'call get_current_state to refresh', but similar guidance is inconsistent across tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Verbose descriptions for LLM consumption. get_semantic_view description is ~280 chars; perform_action is ~230 chars. These exceed the recommended 200-char target, wasting tokens. Descriptions are informative but could be condensed (e.g., 'Navigate to a URL and extract a semantic map filtered by task intent' vs the current longer version).
Parameter 'task_intent' lacks enum constraints or examples of valid values. Description says 'e.g., find the pricing table' but these are examples, not constraints. An LLM might invent task intents not handled by the semantic processor.
Missing pagination/result limiting on semantic map output. If a page contains thousands of nodes, the full list is returned. No documented limit, cap, or pagination guidance. This could exceed context windows on large pages.
get_screenshot description is too brief (25 chars: 'Capture a PNG screenshot of the current page'). This fails to clarify when to use it vs get_semantic_view, or what format/size to expect.
perform_action parameter 'action' enum values lack individual descriptions. An LLM must infer what 'wait' does, whether 'hover' triggers interactions, and how 'scroll' differs from 'navigate'. Each enum value should be documented.
No confirmation pattern for destructive actions. perform_action allows 'click', 'type', 'select' without dry-run or preview. An agent could accidentally submit forms or delete data. Recommend a dry_run parameter or confirmation step.