MCP server for RAG and web crawling with DocMonitor. Integrates web crawling and RAG into AI agents and AI coding assistants with intelligent document type detection (OpenAPI, sitemaps, text files, web pages).
The server provides 5 tools with explicit descriptions and input schemas. Most tools have verb-prefix naming (check_, monitor_, search_, get_) which is good. However, several critical gaps exist: parameter descriptions are minimal or absent, output schemas are not documented, error handling is unspecified, and tool compositions lack clear chaining documentation. The descriptions are adequate in length (ranging from ~80-180 chars) but lack WHEN/WHY guidance and prerequisite hints that LLMs need for tool selection. No tool annotations are present (readOnlyHint, destructiveHint, idempotentHint), making it harder for agents to understand mutation semantics. The 'monitor_documentation' tool has a WRITE risk but its description does not explicitly state the side effect (storing data in vector DB). Input parameters have type declarations but descriptions are sparse or generic.
Check for changes in a document by comparing the latest version with the previous version.
Get change history and differences for a monitored documentation URL between versions.
Get list of all currently monitored documentation URLs with their status and metadata.
Add a documentation URL for monitoring, crawl and index its content with intelligent document type detection (OpenAPI, sitemaps, text files, web pages) and store in searchable vector database.
Search through crawled documentation using semantic search with optional metadata filtering. Returns relevant chunks with similarity scores.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract data without knowing response structure.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. Agents cannot infer mutation semantics or retry safety from schema alone.
Parameter descriptions are minimal or missing context. Example: 'status' in get_monitored_urls lists values as free-form text; should be enum. 'match_count' and 'limit' lack min/max bounds, allowing agents to pass absurd values.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 14 | - | v1 |
Descriptions lack WHEN/WHY guidance and prerequisite hints. No guidance on tool ordering, e.g., 'Call monitor_documentation() first to index content before searching.' Agents must infer logical flow.
Error handling is unspecified. No guidance on recovery: what if URL is unreachable? Crawl fails? Indexing times out? Agents receive no actionable error recovery hints.
Tool composition not documented. No clear chaining IDs in responses. E.g., does check_document_changes return a version_id that get_document_changes accepts? Breaks tool chains.
No pagination documentation. search_documentation accepts 'match_count' but no mention of next_cursor or result offset. Large result sets could blow context window.
Parameter 'metadata_filter' in search_documentation has type 'object' but no schema or example. Agents cannot know what keys/values are valid.