MCP server for full-text search across PDF document collections
A well-structured PDF search server with clear tool purposes and comprehensive descriptions. All 4 tools are properly named with action verbs (search, read_page, read_page_image, stats). Descriptions are detailed and context-aware (194-500 chars), explaining when to use each tool and how they interact. Input schemas are complete with proper types and constraints. Error handling is graceful (PdfSearchError converted to plain text). Main gaps: (1) output schemas are not formally documented in the code; (2) tool composition could be tighter (read_page and read_page_image are separate when they could guide the user to the right one automatically); (3) error responses are plain-text instead of structured with recovery guidance.
Read the full text of a specific page from an indexed PDF. Use after search() to read the complete page content around a match. If the result contains garbled text, broken symbols, or unreadable formulas, use read_page_image() instead — it renders the page as a PNG that preserves formulas, diagrams, and tables exactly. For tables and dense data, crop to the relevant region to read values reliably.
Render a PDF page (or cropped region) as a PNG for visual inspection. Use instead of read_page() when text extraction misses formulas, diagrams, or tables. Workflow: 1. First call: render the full page (no region, default dpi) to orient yourself. Do NOT raise dpi — default 140 already fills the vision model's 1568 px input limit. Higher DPI just gets downscaled. 2. ALWAYS crop before reading values. Tables, formulas, and dense data are NOT reliably readable at full-page scale. Call again with region to crop the area of interest. DPI auto-scales to fill 1568 px for the crop — do NOT set dpi manually, it is computed automatically. stdio transport: the PNG file path on the first line — open it with your file-reading tool to view. Full-page renders append a crop-advisory line after the path; when passing the result to a file reader, use only the first line. http transport: the rendered PNG as inline image content (no file access needed). Full-page renders include the crop advisory as an additional text item. On error: a plain-text message describing the problem.
Search indexed PDFs using FTS5 full-text search. Supports FTS5 syntax: phrases ("exact phrase"), AND (implicit), OR, NOT, prefix (term*), NEAR(term1 term2, 10), parentheses. Terms with special characters (dots, hyphens, colons, slashes, ...) are auto-quoted — FTS5 treats them as token separators. You can also quote them yourself: "4200-3", "v2.1". German ß↔ss / ä↔ae / ö↔oe / ü↔ue variants are expanded automatically. When no results match all terms (implicit AND), the query is automatically relaxed: first by dropping the term least represented in the corpus, then by OR-ing all terms. A note at the top explains what was searched. Structured queries (explicit operators, NEAR, parentheses) are never relaxed.
Output schemas not formally documented in code. read_page returns str, read_page_image returns str or Image union (implicit based on transport), search returns str (formatted list), stats returns str. Descriptions mention output structure but code lacks explicit return type hints or schema registration.
Error handling returns plain-text messages instead of structured errors with recovery guidance. PdfSearchError exceptions are converted to strings (str(e)) without classification (retryable/user-fixable/fatal) or actionable next steps. Pattern: 'User not found. Try search_users()' is not implemented.
Tool composition: read_page and read_page_image are separate, requiring the agent to reason about which to use. The descriptions guide the choice (text extraction fails → try image), but this could be automated. Consider a single 'read_page' tool that offers both modalities or a smarter default.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Show PDF search index statistics (file count, page count, DB size, renderer).
read_page_image transport-specific behavior (stdio path vs http inline image) is handled via module-level _HTTP_TRANSPORT flag. This creates implicit coupling between tool registration and runtime transport selection. Cleaner approach: negotiate in capabilities and adapt at runtime.
search() limit parameter clamped to _MAX_LIMIT=50 but does not inform the user if they requested a higher limit. Silent clamping can surprise agents. Better: return a note 'Limiting to 50 results (requested 100)' or error with guidance 'Max limit is 50. Try refining your query.'