MCP server for web scraping and fetching official documentation content from multiple libraries and frameworks
DocBridge MCP has a single tool 'get_docs' with a clear action verb, but suffers from multiple critical quality gaps. The tool description (139 chars) is adequate length but lacks WHEN to use guidance and prerequisites. The input schema is present with two string parameters ('query' and 'library'), but parameters lack type validation constraints (no enums for the supported library set, no format for query). The output is completely undocumented, no schema definition provided, users cannot know what fields to expect from the response. Error handling is minimal (ValueError for unsupported library) but returns are bare strings without recovery guidance. The tool performs external web scraping with timeouts and UTF-8 encoding edge cases, but none of these constraints or failure modes are documented for the LLM. The description mentions 7 supported libraries in prose ('supports langchain, openai, chromadb, pinecone, uvicorn, fastapi, llama-index') rather than as a formal enum constraint, forcing the LLM to guess valid values or hallucinate new ones.
Search to get official latest documentation content based on query and library name supports langchain, openai, chromadb, pinecone, uvicorn, fastapi, llama-index.
Output schema completely undocumented. Tool returns concatenated strings from multiple URLs but LLM has no schema to guide parsing or downstream chaining. Response structure is free-form text with embedded 'Source: URL' labels, no formal fields, types, or count information.
Library parameter should be an enum, not free-form string. Currently: library: {type: string, description: '...e.g. fastapi'}. Should declare: library: {type: string, enum: ['langchain', 'openai', 'chromadb', 'pinecone', 'uvicorn', 'fastapi', 'llama-index'], description: 'Supported library name...'}. This prevents LLM hallucination of unsupported values.
Query parameter description is minimal (49 chars). Does not explain expected format, length constraints, or example query patterns. Baseline is 72 chars average. Should explain: 'Natural language search query (10-200 chars) describing the feature or API you need documentation for. E.g., "async database connection", "file upload endpoint".'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 38 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Error responses are bare exceptions without recovery guidance. ValueError for unsupported library returns only '❌ Library {library} documentation URL not found'. Should return: 'Library "xyz" not supported. Available libraries: langchain, openai, chromadb, pinecone, uvicorn, fastapi, llama-index. Did you mean one of these?'
No timeout or failure mode documentation. Tool calls httpx with 30-second timeout and falls back to Playwright for JS-heavy pages, but these behaviors are not documented in the tool description. LLM cannot reason about when the tool might timeout or what to do if content fetch fails.
Response limit not documented. Tool concatenates results from all matching URLs in web search, potentially producing very large outputs (multiple full documentation pages). No mention of max response length or truncation behavior. Baseline pattern requires stating limits in description.
Serper API key exposed indirectly: error message reveals external dependency ('❌ SERPER_API_KEY not found in .env file'). While the key itself is not in parameters (good), the failure mode exposes implementation detail. Should return generic: 'Search service unavailable. Please try again later.'