High-fidelity document understanding and hardware control plane with 14+ OCR engines, layout analysis, table extraction, form detection, and direct WIA scanner control
OCR-MCP exhibits significant gaps across definition quality. While 6 tools are explicitly defined with schemas and descriptions, the quality is highly uneven. The primary issue is tool composition: all 6 tools use a generic 'operation' enum pattern that conflates 12+ distinct sub-operations into single tool interfaces (e.g., process_document tool handles 'process_document', 'process_batch', 'analyze_layout', 'extract_tables', 'detect_forms', 'analyze_reading_order', 'classify_type', 'extract_metadata', 'assess_quality', 'validate_accuracy', 'compare_backends', 'analyze_image_quality'). This violates the single-responsibility principle (pattern:tool) and forces LLMs to manage complex conditional logic instead of simple verb_noun tool names. Descriptions are present but generic (100-200 chars) and lack WHEN/WHY guidance. Parameters have types and descriptions but many conditionally-required parameters lack cross-parameter dependency documentation. Output schemas are entirely undocumented, LLMs cannot plan downstream calls. Error handling has no recovery guidance. The execute_agentic_workflow tool references deprecated Sampling (SEP-1577) which is removed in current spec.
Goal-oriented autonomous orchestration via sampling (SEP-1577). AI-orchestrated multi-step document processing with LLM decision-making, adaptive strategy selection, and iterative refinement
Document corpus management with SQLite indexing, full-text search, document retrieval, and metadata tracking. Index processed documents for rapid retrieval and analysis
Image preprocessing, enhancement, format conversion, and PDF composition. Supports operations like deskew, denoise, contrast enhancement, format conversion, PDF merging, and annotation
Batch processing orchestration, workflow automation, system health monitoring, and performance optimization. Configure and execute multi-step document processing pipelines
Hardware scanner control and document acquisition. Discover scanners, list capabilities, configure scan settings, perform scans, and monitor device status
Execute document processing operations including OCR, layout analysis, and quality assessment. Supports 12+ operations: process_document, process_batch, analyze_layout, extract_tables, detect_forms, analyze_reading_order, classify_type, extract_metadata, assess_quality, validate_accuracy, compare_backends, analyze_image_quality
All tools use a generic 'operation' enum pattern that bundles 12+ distinct actions into one tool. process_document handles process_document, process_batch, analyze_layout, extract_tables, detect_forms, analyze_reading_order, classify_type, extract_metadata, assess_quality, validate_accuracy, compare_backends, analyze_image_quality. This violates single-responsibility principle and forces LLMs into conditional reasoning instead of direct tool dispatch.
Tool names are generic verbs + nouns that do not map to actual operations. 'process_document' could mean any of 12 sub-ops. LLMs cannot infer intent from the name alone. Names should be verb_noun pairs matching actual operations: 'extract_tables', 'classify_document_type', 'detect_forms', etc.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 32 | 2.14.3+ | v1 |
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract required fields (e.g., document_id for manage_corpus indexing, backend_name for comparison results). Responses are completely unstructured from the agent's perspective.
Descriptions lack WHEN/WHY guidance and are too generic. 'Image preprocessing, enhancement, format conversion, and PDF composition' does not tell the LLM when to call manage_image vs process_document. Descriptions should answer: What does it do? When should I use it instead of similar tools? What does it return?
Many parameters are conditionally required based on 'operation' enum value, but cross-parameter dependencies are not documented. E.g., process_document with operation='extract_tables' requires 'detect_tables'=true and 'table_region', but other operations ignore these. Undocumented dependencies cause silent misuse.
execute_agentic_workflow references 'sampling (SEP-1577)', which is deprecated and removed from the current MCP spec (2026-07-28). The tool should not rely on server-initiated sampling. Integrate LLM decision-making directly via the LLM provider API or use MRTR (Multi Round-Trip Requests) for user input instead.
No error handling guidance. Tools do not describe what errors are retryable, user-fixable, or fatal. No recovery suggestions. E.g., if OCR fails on a document, should the agent retry with a different backend, enhance the image, or request user input?
operate_scanner tool exposes hardware control but does not document security/permission requirements. Scanner access is sensitive, should require explicit authorization. No mention of audit logging for scan operations.
manage_corpus tool supports import/export and update operations but lacks idempotency documentation. Calling 'index_document' twice with the same document_path should be idempotent (update, not duplicate), but this is not stated.
No batch variants offered for common repetitive operations. manage_workflow has 'execute_batch' but process_document and manage_image do not. Agents calling these in loops waste tokens and latency.