MCP server for fetching and parsing documentation from URLs or local files. Provides tools to list available documentation sources and fetch documentation content with domain-based access control.
The server provides 3 tools with descriptions and basic schemas, but has significant quality gaps. Tool naming follows verb_noun conventions (list_doc_sources, fetch_docs, get_docs), which is appropriate. However, descriptions vary in quality, some are verbose and repetitive, parameter descriptions are minimal or missing, and input schemas lack proper type constraints. The tools are read-only with no mutation semantics, which is good for safety, but parameter validation and error guidance are weak. Output schemas are not documented. The server lacks pagination support despite potentially returning large documentation files. STDIO transport caps the overall score significantly.
Fetch and parse documentation from a given URL or local file. Use this tool after list_doc_sources to: 1. First fetch the documentation file (langgraph.txt or mcp.txt) from a source 2. Analyze the URLs listed in the documentation file 3. Then fetch specific documentation pages relevant to the user's question Args: url: The URL or file path to fetch documentation from. Can be: - URL from an allowed domain - A local file path (absolute or relative) - A file:// URL (e.g., file:///path/to/langgraph.txt or file:///path/to/mcp.txt) Returns: The fetched documentation content converted to markdown, or an error message if the request fails or the URL is not from an allowed domain.
Get documentation for LangGraph or MCP. Always fetch the overview first to get a list of available URLs: - Use "langgraph_overview" for LangGraph documentation - Use "mcp_overview" for MCP documentation Args: url: The URL to fetch. Must start with https://raw.githubusercontent.com/ or be one of the overview options.
List all available documentation sources. This is the first tool you should call in the documentation workflow. It provides URLs to documentation files (langgraph.txt or mcp.txt) or local file paths that the user has made available. Returns: A string containing a formatted list of documentation sources with their URLs or file paths
Inconsistent parameter schemas across tools. list_doc_sources has no parameters (correct for discovery), but fetch_docs and get_docs both accept a 'url' string with minimal type information. Neither schema declares constraints (format, pattern, enum, minLength) that would guide the LLM toward valid inputs. Parameter descriptions exist but are vague (e.g., 'The URL to fetch' in get_docs lacks guidance on valid prefixes beyond the inline text).
Output schemas are not documented. The tool descriptions mention what is returned (e.g., 'formatted list', 'documentation content converted to markdown', 'fetched documentation content') but do not specify the structure of the response (type, fields, format). LLMs cannot plan multi-step workflows or extract specific fields without knowing the response schema. Documentation tools should return structured metadata (source name, URL, fetch timestamp, content length) alongside the actual documentation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
fetch_docs description is overly verbose and repetitive. It repeats 'langgraph.txt or mcp.txt' and explains a three-step workflow that should be implicit. The description is ~280 characters, well above the 194-char baseline for A-grade tools. This dilutes clarity and wastes tokens. Simplify to: 'Fetch documentation from a URL or local file path. Supports HTTP(S) URLs from allowed domains or file:// paths. Returns content converted to markdown.'
No pagination or result limits documented. If a documentation file is large (e.g., 100KB+ of markdown), fetch_docs will return it all in one response, potentially exhausting token budgets. The rubric baseline caps results at 20-50 items for list operations. For fetch_docs, consider adding optional max_length or chunk_size parameters and documenting a response size limit.
Error messages in the source code are generic. E.g., 'Error reading local file: {str(e)}' and 'Error: URL not allowed. Must start with one of the following domains: ...' are returned as plain strings, not structured error objects. The rubric requires errors to guide recovery ('Try search_users() with a partial name' style). When a URL is disallowed, the error lists allowed domains, good, but doesn't suggest alternatives like checking list_doc_sources first.
get_docs uses a specialized string 'langgraph_overview' or 'mcp_overview' as input rather than a formal enum. While the description mentions these as valid options, they are not declared as enum values in the schema. This forces the LLM to know the literal strings rather than discovering them via schema validation. The schema should declare enum: ['langgraph_overview', 'mcp_overview', 'https://raw.githubusercontent.com/...'] to make valid inputs explicit.
Tool composition lacks idempotency guarantees. All three tools are read-only, which is safe, but fetch_docs has side effects via HTTP requests (timeouts, network failures). The description does not mention retry semantics or whether repeated calls with the same URL are guaranteed to return the same content. This matters for agent reliability.