Agent-first, provider-neutral multimodal OCR CLI for images, PDFs, URLs, JSON schemas, and agentic extraction with Gemini, Kimi, Muse, and OpenRouter.
The server defines 3 tools for OCR document processing with reasonable schema completeness. All tools have descriptions and input schemas with type definitions. However, several moderate gaps exist: (1) Tool naming could be more verb-focused and action-oriented ('re_ocr_region' is passive; 'refine_region' or 'extract_region_text' would be clearer). (2) Parameter descriptions lack clarity on expected behavior and constraints, e.g., 'focus' in re_ocr_region doesn't explain the expected format or output. (3) Output schemas are not documented, making it unclear what these tools return. (4) Error handling and recovery guidance are absent. (5) No per-parameter validation rules or examples. The tools operate on a domain (OCR extraction) where composition and chaining are likely, but there's no documentation of which tools call each other or what IDs they exchange. The schema completeness is above average for parameters, but the lack of output documentation and weak error handling prevents a higher score.
Analyze the overall structure and layout of the document.
Extract all currently visible structured fields in one batch using canonical field names.
Re-process a specific region of the document with higher focus to improve accuracy.
Tool names lack clear action verbs. 're_ocr_region' is passive and unclear; 'refine_region_ocr' or 'extract_region_with_focus' would better signal the action. 'analyze_document_structure' is generic, does it segment, classify, or measure layout?
Output schemas are not documented. It is unclear what these tools return. Tool descriptions mention results implicitly (e.g., 'Extract all currently visible structured fields') but no explicit return structure is defined. LLMs cannot plan downstream calls without knowing output shape.
Parameter 'focus' in re_ocr_region lacks constraint guidance. Description says 'Specific text or type of content' but doesn't specify format, allowed keywords, or examples. This invites hallucinated values from the LLM.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
No error handling or recovery guidance. If OCR fails, confidence drops below threshold, or a region is malformed, the tools provide no actionable error messages. Agents cannot self-correct or fall back to alternative strategies.
Tool composition and chaining not documented. For extract_fields_batch input, 'field_name' references canonical field names, but there's no tool or discovery mechanism to list valid field names. analyze_document_structure recommends an extraction strategy, but doesn't explain how to invoke the strategy or which tool to call next.
Parameter descriptions are generic. 'The confidence score of the extraction' doesn't explain whether 0.5 is acceptable, what the LLM should do if confidence is low, or how this interacts with confidence_threshold in re_ocr_region.
Normalized coordinate system (0-1 range) is used without a diagram or explanation of the origin (top-left vs bottom-right) or how to map real pixel coordinates to normalized space.