MCP server for FlexOrch — SDK for machines. Converts unstructured documents (PDF, DOCX, invoices, contracts, payroll, etc.) into structured, LLM-ready datasets with PII masking and quality scoring.
FlexOrch MCP demonstrates strong fundamentals: all 9 tools have explicit descriptions (avg 180 chars), complete input schemas with type definitions, and output models defined in Pydantic. Naming is action-centric (process, status, result, build, search, export, index, chunks, reprocess). However, parameter descriptions vary in quality, some lack constraints/ranges, and a few output fields are undocumented. Tool composition is clean (each tool does one thing), and error handling uses structured Pydantic models with isError/error fields. The critical gap is the absence of per-parameter descriptions in JSON Schema definitions, while descriptions exist in prose, they are not formally attached to schema fields, forcing LLMs to infer parameter semantics. Additionally, tool descriptions mention asynch workflows and polling requirements, but the tools themselves do not enforce or guide this pattern (e.g., no retry helper, no max_retries parameter). Output schemas are documented in Python models but not exposed in tool definition responses, limiting discoverability.
Create a structured dataset from a completed execution (Step 4). Convert extracted fields into a queryable dataset. Returns a job_id — poll job.status until status='completed', then use the dataset_id in dataset.export to download all records.
Retrieve RAG-ready text chunks from an indexed dataset (Pro+ plan required). Returns LangChain/LlamaIndex-ready text chunks with pagination support and quality/PII filtering.
Download a built dataset in text format (Step 5). Returns the full dataset content as text in the requested format (json, jsonl, csv, md, xml, or rag). Binary formats (parquet, hf) must be downloaded via the FlexOrch API directly.
Trigger semantic indexing for a dataset (Pro+ plan required). Indexes chunks for RAG pipelines. Use dataset.chunks to retrieve chunks once indexing is complete. Indexing typically takes 10–60 seconds depending on dataset size.
Full-text and semantic search across indexed datasets. Search existing datasets without processing a new document. Supports multiple search modes and optional filtering by document type, language, and quality grade.
Parameter descriptions not formally attached to JSON Schema fields. Descriptions appear in tool docstrings but not in inputSchema.properties[param].description. LLMs relying on schema introspection will not see constraints like 'max 50 MB', 'enum values', or 'required dependency on execution_id'.
Output schemas (ProcessDocumentResult, JobStatusResult, ExtractionResult, etc.) defined in Python but not exposed in MCP tool response schemas. FastMCP likely auto-derives output schemas from return type hints, but this is framework-specific and non-standard. LLMs cannot introspect what dataset.export will return (filename, content, byte_count fields) without calling it or reading source code.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 80 | 2026-07-28+ | v2 |
Submit a document for processing — this is always the first step (Step 1 of 5). Downloads the file from file_url, then submits it to FlexOrch for automatic classification, structured field extraction, PII detection/masking, and quality scoring. Processing is asynchronous — this tool returns immediately with a job_id. You MUST call job.status(job_id) every 3–5 seconds until status='completed' before calling job.result.
Re-queue an already-uploaded document for reprocessing with updated settings. Useful for re-extracting data with different PII masking or document type hints. Returns a job_id — poll job.status until completion.
Read structured fields extracted from a completed document (Step 3). Use the execution_id from a completed data_process job (job.status response). Returns document type, detected language, quality grade (A–D), PII summary, column list, and extracted field values. If no dataset has been built yet, the response includes a fields_hint guiding you to call dataset.build next. To retrieve all rows as a file, proceed to dataset.build → dataset.export.
Poll a job until it finishes — call this after document.process or dataset.build (Step 2). Call repeatedly every 3–5 seconds until status is 'completed' or 'failed'. For data_process jobs: the completed response includes execution_id — pass it to job.result. For dataset_build jobs: the completed response includes dataset_id — pass it to dataset.export.
dataset.search accepts 'mode' parameter with enum (auto, semantic, structured, hybrid) but description does not explain what each mode does or when to use each. An LLM has no guidance on which mode suits a given query.
dataset.export 'format' parameter accepts 'rag' but description does not explain what 'rag' format returns or how it differs from json/csv. Also notes binary formats (parquet, hf) must be downloaded 'from the API directly', unclear how an agent should do this if the tool does not support them.
Polling pattern (job.status every 3 - 5 seconds) documented in tool descriptions but not enforced by the tools. No max_retries, backoff strategy, or timeout guidance. An agent could poll indefinitely or too aggressively, overloading the API.
document.process 'document_type' parameter lists enum values (invoice, expense_report, etc.) but does not state what happens if the hint is wrong or omitted. Does auto-detection always succeed? What if detection fails, does it error or return a confidence score?
job.result and dataset.chunks return paginated results (row_count, has_more fields) but dataset.chunks accepts explicit pagination (page, page_size), while job.result does not. Inconsistent pagination API across tools forces LLMs to learn two patterns.
Error handling uses isError/error fields in Pydantic models but does not categorize errors as retryable vs. fatal. An LLM seeing error='Dataset not found' cannot determine if retrying will help or if it must ask the user for a different dataset_id.
dataset.index and dataset.chunks require Pro+ plan but tools do not validate plan tier or return a specific error (e.g., 'Indexing requires Pro+ plan') if called on a free account. LLMs will receive a generic error and cannot self-correct.