MCP server for intelligent PDF ingestion and memory persistence. Analyzes PDF structure, extracts text/tables/images, and creates optimized indices using LlamaIndex query engines for semantic search and retrieval.
The server exposes a single tool 'determine_memory_architecture' with a clear description and documented parameters. However, the tool combines multiple responsibilities (PDF parsing, image extraction, table detection, hierarchy analysis, and ingestion strategy selection) into one monolithic operation, violating the single-responsibility principle. The output schema is not formally documented, the code shows a dictionary return with 'structure_envelope' and 'ingestion_plan' keys, but no schema is visible defining the internal structure of these nested objects. Error handling is minimal (only base64 validation and a generic ValueError), lacking recovery guidance for the LLM. The tool name 'determine_memory_architecture' is clear and verb-forward, but the parameter descriptions are adequate but not LLM-optimized, they lack context on when to use the tool or what the tradeoffs are between extract_and_embed_images and look_for_queryable_tables.
Determine the architecture for PDF ingestion. This function kicks off an multi-agent pipeline to determine the best architecture for ingesting a PDF file and storing it. It returns an envelope with the metadata needed for ingestion such as: text_chunks, tables, images, screenshots, OCR text, hierarchy graph All build smartly by the agents based on the PDF content. It also returns a strategy that can be used for ingestion and memory persistence based on Llama-Index Query Engines.
Monolithic tool combines 6+ distinct responsibilities: PDF parsing, image extraction, table detection, OCR processing, hierarchy analysis, and ingestion strategy selection. LLMs cannot reason about which features to disable or why, and cannot retry partial failures.
Output schema not formally documented. Code returns a dict with 'structure_envelope' and 'ingestion_plan', but internal structure (e.g., what fields are in structure_envelope, what type/format is ingestion_plan) is not visible. LLMs cannot plan downstream operations without knowing return fields.
Minimal error handling. Only base64 validation is explicit; no guidance on what to do if the PDF is corrupt, encrypted, too large, or if agent extraction fails. LLM receives raw errors with no recovery path.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 54 | - | v1 |
Tool description (139 chars) is adequate in length but lacks usage context. Does not explain WHEN to use this vs other ingest tools, what the tradeoffs are, or what downstream operations depend on structure_envelope vs ingestion_plan. LLMs cannot disambiguate if multiple ingestion strategies exist.
Parameter 'pdf' is documented as 'The PDF file content as base64 encoded string' but lacks guidance on size limits, supported PDF versions, or expected encoding format. LLMs may pass invalid base64 or oversized files without warning.
Boolean parameters 'extract_and_embed_images' and 'look_for_queryable_tables' have simple descriptions but no guidance on performance/latency tradeoffs or when NOT to enable them. LLMs cannot reason about cost vs. signal.