MCP server for document parsing and text/image/barcode extraction using GroupDocs.Parser Cloud API
The server defines 9 tools with HTTP transport and reasonable structure. Tool names follow verb_noun conventions (parser_extract_text, file_upload, etc.). Descriptions are present and range from 80 - 250 characters, meeting basic adequacy. However, parameter descriptions are sparse or generic, input schemas lack enums and range constraints, and error handling is minimal. Output schemas are partially documented in docstrings but not in formal schema declarations. Security is well-handled (no secrets in parameters, credentials via environment variables). The server passes baseline quality but misses several patterns from Arcade's 54 patterns for production-grade agent tools.
Download a file from Cloud storage and return it as base64 content. Args: path (str): Cloud file path (e.g., "/docs/sample.pdf"). Returns: DownloadedFile: Object containing path, filename, size, base64 data.
Download a file from Cloud storage and save it directly to a local path. This method avoids base64 encoding and is intended for cases where the MCP server is executed locally (for example on 127.0.0.1 or localhost) and has direct access to the desired filesystem path. Args: path (str): Cloud file path (e.g., "/docs/sample.pdf"). local_path (str): Absolute path on the MCP server host where the file should be stored. Returns: str: Success message with target local path.
Upload a file stream to GroupDocs Cloud storage. Args: file_stream (bytes | str): File contents as raw bytes or a base64-encoded string. cloud_path (str): Destination path in Cloud storage (e.g., "/uploads/sample.pdf"). Returns: str: Success message with target path.
Upload a local file (by path) to GroupDocs Cloud storage. This is useful for environments where the LLM cannot send raw/base64 file bytes, but can reference a file that the MCP server can read from the local filesystem. Args: local_path (str): Absolute path to a file that is readable by the MCP server process. cloud_path (str): Destination path in Cloud storage (e.g., "/uploads/sample.pdf"). Returns: str: Success message with target path.
Parameter descriptions are minimal or absent. Example: 'file_stream' in file_upload is described as 'File contents as raw bytes or a base64-encoded string' but lacks guidance on when to use which format, maximum size, or what happens if the format is invalid. LLMs cannot infer format requirements without explicit constraints.
No enum constraints for cloud_path or local_path parameters. These accept free-form strings with no validation guidance (e.g., no character restrictions, path format rules, or examples of valid formats). LLMs will generate arbitrary paths that may fail silently or cause unexpected behavior.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 77 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Extract barcodes from a document in Cloud storage. Args: path (str): Cloud path of the document to analyze. Returns: list[Barcode]: List of detected barcode entries.
Extract images from a document in Cloud storage. Args: path (str): Cloud path of the document to analyze (e.g., "/docs/sample.pdf"). Returns: list[ImageFile]: List of extracted image entries. Each entry contains `path` to the image in cloud storage.
Extract text from a document in GroupDocs Cloud storage. Args: path (str): Cloud storage path to the document (e.g., "/docs/sample.pdf"). Returns: str: Extracted textual content.
Download an extracted image from Cloud storage and return it as an ImageContent. This tool is meant to be used after `parser_extract_images`, which returns cloud paths of extracted images. Args: image_path (str): Cloud path of the extracted image (e.g., "/temp/parsed/img_0.png"). Returns: PIL.Image.Image: The downloaded image. FastMCP serializes it as ImageContent.
Get list of formats supported by GroupDocs.Parser Cloud. Returns: list[str]: A list of supported extensions (e.g., ".pdf", ".docx").
Output schemas are documented in docstrings (e.g., 'list[str]', 'DownloadedFile') but not in formal JSON Schema declarations visible to the MCP client. The FastMCP framework may infer schemas from return type hints, but this is implicit and not verifiable from the source. LLMs relying on the response schema may have incomplete or incorrect expectations.
Error handling is minimal. Most tools lack recovery guidance. Example: parser_extract_barcodes wraps the exception with 'Failed to extract barcodes from {path}', but does not categorize the error (retryable? user-fixable? fatal?), suggest recovery steps, or include the invalid value that triggered the error. File operations silently fail if paths are invalid (no validation before calling the API).
Destructive tools (file_upload, file_upload_local, file_download_local) lack confirmation or dry-run support. An agent could accidentally overwrite critical cloud storage paths or write to invalid local paths without safeguards. No idempotency guarantee documented.
No pagination support. Tools like parser_supported_formats return all formats in a single list without pagination or result limits. If GroupDocs adds thousands of formats, the response could consume excessive tokens. List-returning tools should include limit, offset, and total_count parameters and document the maximum result size.
Parameter naming inconsistency. Some tools use 'path' (parser_extract_text, file_download) while others use 'cloud_path' (file_upload) or 'image_path' (parser_get_image). Similarly, 'local_path' appears in file_upload_local and file_download_local but the parameter name differs in other tools. This forces the LLM to reason about field mappings instead of following a consistent pattern.
file_upload accepts both bytes and base64-encoded string input, but the distinction is undocumented. The description says 'File contents as raw bytes or a base64-encoded string' but does not explain when to use which format, what the LLM should do if it has only one encoding, or what happens if the format is wrong. This ambiguity invites misuse.
No tool composition documentation. It is unclear whether parser_extract_images output (cloud paths) are always compatible with parser_get_image input, or whether there are edge cases. The docstring for parser_get_image says 'This tool is meant to be used after parser_extract_images' but does not guarantee that all extracted image paths are downloadable via parser_get_image, or what happens if a path is not found beyond the FileNotFoundError message.