A Model Context Protocol server built with Swift that provides tools for fetching web pages, searching within pages, and searching the web using various search engines.
MCP-WebReader demonstrates good tool design with strong naming conventions, comprehensive parameter schemas, and well-structured descriptions that guide LLM usage. All three tools follow verb_noun naming patterns (fetch-page, search-page, search-web) and include detailed descriptions (180-250 chars each) explaining use cases and synergistic relationships. Input schemas are complete with typed parameters and descriptions. However, there are notable gaps: output schemas are not documented (only mentioned generically in descriptions), error handling guidance is minimal, and tool annotations (readOnlyHint, idempotentHint) are missing. The fetch-page tool's description is particularly strong, explaining when to use renderJS and how tools work together. Search patterns (literal text matching in search-page, enum-based engine selection in search-web) show good constraint design. The server implements HTTP method flexibility and custom headers, which is sophisticated. Overall, the tools are production-usable but lack the polish and guardrails of an A-grade implementation.
Fetch web page content at a given URL and return paginated text. Use `renderJS: true` for JavaScript-heavy sites (SPAs, Google Search, Reddit). Use `renderJS: false` (default) for faster fetching of static content. Returns cleaned text with HTML stripped. Works symbiotically with search-page: use search-page to find specific content and get character positions, then use fetch-page with offset/limit parameters to retrieve the full context around those positions. Can also be used with search-web results to read content from discovered URLs.
Search for a literal text within a web page and return all match positions with context. Use this to find specific information on a page, then use fetch-page with the returned positions to retrieve full content. Shares cache with fetch-page for efficiency. Use `renderJS: true` for JavaScript-heavy sites.
Search the web using a search engine (Google, DuckDuckGo, Bing, Brave) or a custom search URL template. Returns links with context from search results. Works symbiotically with other tools: use search-web to discover relevant URLs, then use fetch-page to read content from those URLs or search-page to search within them. Uses caching for efficiency.
Output schemas not documented. Tools return structured data (StructuredContentOutput with metadata and content arrays, visible in ToolSupport.swift) but the exact response schema is never declared to the LLM. LLMs cannot plan downstream operations or extract specific fields without knowing the response structure.
Missing tool annotations for risk classification. All three tools are READ_ONLY (fetching and searching), but this is declared only in the evaluation metadata, not in the tool definition via readOnlyHint. LLMs cannot reliably infer idempotency and safety without explicit annotations.
Error handling lacks recovery guidance. The ContentError enum (contentError, missingArgument, mismatchedType, other) exists but tool descriptions do not explain what errors can occur or what the LLM should do (retry, ask user, use alternative tool). fetch-page with renderJS=true may fail for certain sites; search-web may return empty results, neither explains fallback strategies.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 73 | 2025-06-18+ | v2 |
search-page and search-web accept customHeaders and httpMethod parameters with minimal constraint documentation. customHeaders could be abused (e.g., Host header injection); httpMethod allows DELETE/PUT on read-only operations. No validation guidance or security notes in parameter descriptions.
Pagination for search-web is underspecified. The tool returns 'links with context' but does not declare how many results are returned, whether there is a next_cursor, or whether results are ordered by relevance. An agent cannot safely iterate or know when all results are exhausted.
fetch-page limit parameter (default 2500 chars) lacks guidance on typical page sizes and how to handle truncation. If a user needs the full page, should they increase limit? Should the tool return a 'truncated: true' flag? Tool description mentions pagination but does not explain truncation behavior.