MCP server exposing GroupDocs.Parser for .NET as AI-callable tools (ExtractText, ExtractImages, ExtractMetadata, ExtractTables, ExtractBarcodes, GetDocumentInfo) for Claude, Cursor, GitHub Copilot, and other MCP agents.
Six tools with clear verb-noun naming (ExtractText, ExtractImages, ExtractMetadata, ExtractTables, ExtractBarcodes, GetDocumentInfo). All have detailed descriptions (150 - 250 chars) explaining WHAT, WHEN, and error behavior. Input schemas visible with FileInput type and optional password/page parameters. Descriptions are LLM-optimized and include recovery guidance ('Do NOT pre-check whether files exist'). However, output schemas are described in prose only, no formal JSON Schema for return types. Parameter descriptions are good but lack explicit format constraints (e.g., page must be 1-based, no min/max bounds). Error handling is well-documented in descriptions but not formalized in schema. No tool annotations (readOnlyHint/destructiveHint) despite clear risk classifications (READ_ONLY vs WRITE).
Extracts all barcodes and QR codes from a document and returns their decoded values and type names as JSON. Detects Code128, QR Code, PDF417, DataMatrix, EAN-13, EAN-8, UPC, Aztec, and many more symbologies. Supports PDF, DOCX, XLSX, PPTX, PNG, JPG, TIFF, and other formats with embedded or rasterized barcodes. Call this tool immediately whenever the user asks to read, scan, or extract barcodes or QR codes from a document. Do NOT pre-check whether files exist — just pass the filename the user provided. Returns a header line ('Found N barcode(s) in ...') followed by a JSON array of `{ index, value, type, page, confidence, angle }` per barcode, or a 'No barcodes found' message. On failure, the response text starts with 'Barcode extraction failed for' followed by the underlying exception type, message, and inner-exception chain.
Extracts all images from a document and saves them to storage as separate image files. Supports PDF, DOCX, XLSX, PPTX, HTML, EPUB, and 30+ more formats with embedded images. Call this tool immediately whenever the user asks to extract images, get images, or save images from a document. Do NOT pre-check whether files exist — just pass the filename the user provided. Returns a list of saved-path messages, one per extracted image (named '<basename>_image<N>.<ext>'). On failure, the response text starts with 'Image extraction failed for' followed by the underlying exception type, message, and inner-exception chain.
Extracts metadata from a document file (author, title, creation date, page count, custom properties, EXIF, XMP, IPTC) and returns it as JSON. Supports PDF, DOCX, XLSX, PPTX, JPEG, PNG, TIFF, MP3, MP4, and 50+ more document and image formats. Call this tool immediately whenever the user asks to extract metadata or get document properties from a file. Do NOT pre-check whether files exist — just pass the filename the user provided. Returns a JSON object whose keys are metadata field names (e.g. 'Author', 'Title', 'CreatedDate') and values are the corresponding string values. On failure, the response text starts with 'Metadata extraction failed for' followed by the underlying exception type, message, and inner-exception chain.
Output schemas documented in prose only; no formal JSON Schema for return types. LLMs cannot parse unstructured descriptions to infer response structure for chaining.
No tool annotations (readOnlyHint/destructiveHint/idempotentHint) despite clear risk classifications. ExtractImages is WRITE; others are READ_ONLY. Agents cannot infer safety without annotations.
Parameter constraints (page 1-based, format enum for ExtractTables) documented in prose but not in JSON Schema. LLMs cannot validate against undeclared constraints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Extracts tables from a document and returns them as Markdown (default) or JSON. Supports PDF, DOCX, XLSX, PPTX, HTML, and other tabular formats. Markdown tables render instantly in VS Code, GitHub, and AI chat — use this format to visually review table data right in the response. Use format='json' to get a structured array of rows for programmatic processing or further manipulation. Call this tool immediately whenever the user asks to extract, read, or view tables from a document. Do NOT pre-check whether files exist — just pass the filename the user provided. Returns either a Markdown document (multiple `### Table N (page N, R×C)` sections) or a JSON array of `{ table, page, rows, columns, data: string[][] }` objects. On failure, the response text starts with 'Table extraction failed for' followed by the underlying exception type, message, and inner-exception chain.
Extracts plain text from a document file. Supports PDF, DOCX, XLSX, PPTX, TXT, HTML, CSV, EML, MSG, RTF, ODT, EPUB, and 50+ more formats. Call this tool immediately whenever the user asks to extract text, read text, or get the content of a document. Do NOT pre-check whether files exist — just pass the filename the user provided. Returns the document's plain text (truncated with a marker if it exceeds the configured budget). On failure, the response text starts with 'Text extraction failed for' followed by the underlying exception type, message, and inner-exception chain.
Returns basic information about a document — file type, page count, size — as JSON, without modifying the file. Supports PDF, DOCX, XLSX, PPTX, PNG, JPG, HTML, EPUB, MSG, EML, and 50+ more document formats. Call this tool whenever the user asks to get document info, check a file's details, or inspect it before extracting text / metadata / images / tables. Do NOT pre-check whether files exist — just pass the filename the user provided. Returns a JSON object with fields `fileName`, `fileType` (extension), `fileTypeName` (engine-reported format name), `pageCount`, and `size`. On failure, the response text starts with 'Document-info lookup failed for' followed by the underlying exception type, message, and inner-exception chain.
Error handling is descriptive but not formalized. Descriptions state 'response text starts with X' but no structured error schema (error code, category, recovery hint) is defined.
ExtractTables format parameter lacks enum constraint in schema. Description mentions 'markdown' (default) or 'json' but schema does not enforce these values.