Local PDF ingestion helpers for LLM agents. Provide the PDF path (relative to the allowed base directory when set) and tune pagination or semantic search parameters as needed.
PDF Toolbox presents a focused domain (PDF reading, chunking, semantic search) with four tools. Tool naming follows verb_noun conventions well (read_pdf, search_pdf, describe_pdf_sections, configure_pdf_defaults). All tools have descriptions and input schemas are formally defined. However, descriptions are inconsistently detailed, error handling guidance is absent, and the configure_pdf_defaults tool conflates configuration management with tool execution (mixing concerns). The server demonstrates intermediate quality: schemas are present and mostly complete, but lack the LLM-optimized depth and error recovery patterns expected of production-grade tools.
Updates the default PDF processing parameters so subsequent calls inherit them.
Generates sequential chunks enriched with offsets and metadata for LLM workflows.
Extracts plain text from a PDF page range while keeping the payload under control.
Runs semantic search over embedded PDF chunks and returns the highest scoring matches.
configure_pdf_defaults violates single-responsibility principle. It mixes agent state mutation (defaults) with tool execution. This conflates concerns: agents should not need to manage tool configuration mid-execution. Should be removed or moved to server initialization only.
Error handling guidance is absent across all tools. Tools do not document what errors are retryable, user-fixable, or fatal. Example: search_pdf with min_score=0.25 default can silently return zero results; tool description does not warn of this or guide recovery. Per pattern:recovery-guide, every tool should tell LLM what to do next on failure.
Output schemas are not documented in tool definitions. Descriptions mention what tools return (e.g., search_pdf returns 'highest scoring matches') but lack structured field documentation. LLMs cannot plan downstream calls without knowing: What fields does search_pdf.results contain? Is chunk_id present? Is offset always included? Per pattern:response-shaper, tool responses must have documented output schemas.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 54 | - | v1 |
describe_pdf_sections mode parameter documentation is minimal. The enum=['chunks','tables'] constraint is present but the description does not explain when to use each mode, what output differences are, or which is the default. LLMs will guess.
search_pdf description does not clarify behavior of empty results. If min_score threshold filters out all matches, does the tool return empty list or an error? The parameter default (0.25) is conservative but undocumented as a starting point for tuning. LLMs may not know how to adjust min_score to get results.
Parameter descriptions lack format/constraint guidance. For example, read_pdf.path accepts 'relative or absolute' paths, but does not specify: What is the base directory? Are symlinks allowed? What happens if path escapes allowed dirs?
Security: read_pdf, search_pdf, describe_pdf_sections accept arbitrary file paths. No validation shown in tool definitions against path traversal (../../../etc/passwd). Code review shows resolve_pdf_path() and set_base_path() exist in services, but tool descriptions do not document path security constraints. Per pattern:tool-gateway, all agent input must be sanitized and constraints documented.
Composition issue: read_pdf returns pages but search_pdf requires pre-indexed PDFs. Tool descriptions do not explain the workflow: agents must understand that search_pdf silently requires prior embedding indexing (not explicit via a tool). Per pattern:tool-chain, tool outputs must contain references and guidance that downstream tools need.