An MCP server for extracting text from documents (PDF, TXT, images) using OCR and AI, with chat capabilities for querying uploaded documents
This MCP server has significant quality gaps across naming, descriptions, parameter documentation, and schema completeness. While the server does register three tools with input schemas (using Zod), the definitions lack production-grade rigor. Tool names are somewhat descriptive but lack clarity around intent. Descriptions are present but generic and under-optimized for LLM selection. Parameter schemas exist but lack comprehensive documentation of constraints, ranges, and dependencies. Critical security issues exist around stateful session management and lack of permission gating. Error handling is minimal, no recovery guidance in tool responses.
Ask questions about previously uploaded document text
Extracts text from PDF, TXT, or image (base64 input) and stores it
Run predefined checks on previously uploaded document text
sessionId parameter used pervasively for stateful session management, creating implicit dependencies between tool calls. Sessions are stored in-memory server-side (this.sessionData), violating MCP's stateless-per-request design. Protocol requires each request to be self-contained with _meta; reliance on Mcp-Session-Id is a removed pattern.
No input validation on fileBase64 parameter in extractTextFromFile, accepts arbitrary strings with no size limit. Can exhaust memory with multi-MB base64 payloads. Missing constraint documentation (e.g., max 10MB).
Tool descriptions are generic and under-optimized for LLM selection. 'Extracts text from PDF, TXT, or image' (35 chars) is too brief, lacks WHEN to use, what happens on error, or what the tool returns. Baselines: A+ tools average 194 chars, min 50-100 for clarity. All three descriptions fall well below this.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
fileType parameter accepts free-form string (pdf, txt, png, jpg) without enum constraint. LLMs may hallucinate invalid types (jpeg, PDF, JPEG, TIFF). Should be an enum: ['pdf', 'txt', 'png', 'jpg'].
No output schema documented for any tool. Handler returns { content, sessionId, RawText } for extractTextFromFile, but LLMs cannot see what fields to expect. No pagination or result-limiting for future scalability.
askAboutUploadedText description says 'Ask questions about previously uploaded document text' (50 chars) but does NOT specify: What AI model answers? What context is used? How are errors (no docs uploaded) handled? Lacks dependency hints.
runPredefinedChecks description is vague ('Run predefined checks'), what checks? What formats output? No mention of what happens if sessionId has no uploads or if checks fail.
Error handling is minimal. If sessionId doesn't exist, uploading returns success without validation. If OCR fails on image, no retry guidance. Tool responses lack structured error codes, recovery suggestions, or user-fixable vs. fatal classification.
No permission gating. Any client can call any tool without authorization checks. No scope declarations (e.g., 'read:documents', 'write:sessions'). An untrusted agent can access all uploaded documents from all sessions.
GEMINI_API_KEY is injected via process.env, good. However, sessionId is client-supplied and optional (randomUUID() generated server-side). LLM can invent session IDs, accessing or poisoning other sessions' data. No user authentication or session ownership verification.
Tool composition: extractTextFromFile AND stores the file AND generates a summary AND adds to chat history. This bundles text extraction, AI summarization, and history management in one tool, splits concerns. Should be separate: extractText → summarizeText → addToHistory.
askAboutUploadedText does not document what LLM context is passed (chat history? full documents? summary?). Parameter 'question' has no format constraint. Should clarify: 'A natural-language question (1-1000 chars) about the uploaded document text.'