An MCP server providing document processing, web fetching, and web search capabilities over HTTP using streamable transport
The server defines 6 tools with reasonable descriptions and clear purpose statements. However, critical gaps exist: NO input schemas are visible in the source code for ANY tool. The code shows tool registration via mcp.AddTool with Name and Description fields, but no InputSchema definitions. This is a hard blocker, per the rubric, 'If a tool has NO input schema at all: its schema score MUST be 0.' Descriptions are generally good (100-300+ chars, exceeding the 10-1024 char baseline and the 194 char median), but parameter descriptions are missing entirely from the code. The tools are well-named (verb_noun pattern: get_document_info, read_document_smart, read_document_by_page, read_document_by_line, fetch, search) and cover a coherent domain (document reading + web fetching + search). However, without visible input schemas, the server cannot pass schema validation or guide LLM parameter selection, a critical blocker for production use.
Fetch web resource content from URL and return content in specified format. ## When to Use Use this tool when you need to: - Get raw content from a web page - Access API endpoints to get JSON data - Download HTML/text/Markdown content - Quickly obtain web resources without complex processing Don't use this tool when you need to: - Extract specific information from a web page (should use specialized extraction tools) - Analyze or summarize web page content (should fetch first then analyze) ## Features - Supports four output formats: text (plain text), markdown (Markdown format), html (HTML format), json (JSON format) - Automatically handles HTTP redirects - Automatically extracts plain text from HTML (text format) - Automatically converts HTML to Markdown (markdown format) - Sets reasonable timeout to prevent long waits - Limits response size (maximum 5MB) to prevent memory overflow ## Usage Tips - text format: Suitable for getting plain text content or extracting text from HTML - markdown format: Suitable for content that needs formatted rendering - html format: Suitable for scenarios requiring raw HTML structure - json format: Suitable for JSON data returned by API endpoints - Set appropriate timeout based on website speed (default 30 seconds, maximum 120 seconds)
Get basic information about a document, including file type, size, page count, and metadata. Supported formats: .docx, .pdf, .xlsx, .pptx, .txt, .csv, .md, .rtf. This is the first step before reading a document, helping you understand document structure and decide how to read it.
Read document content by line range. Supports specifying page and line range. Parameters: - start_line: Starting line (0-based) - end_line: Ending line (-1 means to end) - page_index: Page index (-1 means first page) Best for: Reading specific lines or paragraphs
NO input schemas visible for any tool. Tools are registered with mcp.AddTool(s, &mcp.Tool{Name, Description}) but InputSchema field is missing or not implemented. This prevents LLMs from understanding parameter types, required fields, enums, or constraints, a critical blocker for tool selection and parameter binding.
Parameter descriptions embedded only in tool descriptions, not formalized in schema. LLMs cannot reliably parse prose descriptions for parameter guidance. The rubric requires 'Every parameter needs a description explaining what it controls' and 'Describe the expected format, range, and allowed values directly in the parameter description.' Without formal schemas, this is impossible.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Read document content by page range. Supports multi-page documents like PDF, PPTX. Parameters: - start_page: Starting page (0-based) - end_page: Ending page (-1 means to end) Returns detailed information for each page (page number, line count). Best for: Reading specific pages or sections
Intelligently read document content, automatically handling large documents. Features: - Automatically adapts to context limits (default 50000 characters) - Supports sampling mode (uniform sampling) or truncation from start - Automatically cleans extra spaces and blank lines - Provides suggestions when document is too large Best for: First-time document reading, quick content overview
Search for information using DuckDuckGo search engine. ## When to Use Use this tool when you need to: - Search for the latest information on the internet - Find materials on specific topics - Get search results within a specified time range ## Usage Tips - Provide clear, specific search keywords for better results - You can use the time_range parameter to limit search results to a specific time period
fetch tool accepts 'format' parameter with implicit enum values (text/markdown/html/json) documented only in the description prose. Without a formal enum in the schema, LLMs may invent other formats or misinterpret the constraint.
search tool's 'time_range' parameter has implicit enum values (d/w/m/y/empty string) documented in prose only. Without a formal enum schema, LLMs may pass invalid values like 'past-month' or 'monthly'.
read_document_by_page and read_document_by_line both accept integer parameters (start_page, end_page, start_line, end_line) with implicit ranges (0-based, -1 means end) documented only in descriptions. Without formal schema constraints (minimum, maximum, or description text parsing), LLMs may pass negative numbers, floats, or out-of-range values.
NO output schemas documented for any tool. The rubric states 'Document the output schema. LLMs need to know what fields to expect so they can plan downstream tool calls.' Without output documentation, LLMs cannot reason about result structure or chain tools effectively.
NO error handling guidance. Tools do not document what errors are possible, whether they are retryable, or what the LLM should do on failure. Per pattern:recovery-guide, 'Error responses must tell the LLM what to do next.'
fetch tool sets a 5MB response size limit and mentions it in the description, but no timeout error recovery guidance is provided. If a fetch times out or exceeds the limit, the LLM has no guidance on what to do (retry, use a different tool, ask user).