Parse documents (PDF, DOCX, PPTX, HTML) and answer questions by reading only necessary sections. Provides five tools for document parsing, outline retrieval, content reading, semantic search, and markdown export.
DocSlicer MCP server demonstrates solid definition quality with well-named, verb-forward tools and comprehensive descriptions. All 5 tools have clear, actionable descriptions (100-400+ chars) that explain WHAT the tool does, WHEN to use it, and key prerequisites. Input schemas are present with typed parameters and descriptions. However, output schemas are not explicitly documented in the visible code, descriptions mention what is returned (outline, chunks, headings with snippets) but formal JSON Schema output definitions are absent. This forces LLMs to infer response structure. Tool names follow verb_noun pattern (parse, get_outline, read, search, to_markdown). Parameter naming is consistent and descriptive (doc_id, source, headings, query, password). The server includes security-conscious design (password parameter for encrypted docs, file path restrictions via roots). One tool (to_markdown) is correctly marked as WRITE, showing risk awareness. Error handling guidance is embedded in descriptions but not formalized as recovery patterns in code. Overall, this is a well-designed document manipulation API that follows most patterns, but lacks explicit output schema documentation and formal error classification.
Retrieve the heading outline for a parsed document by doc_id. Use this when the outline scrolls out of context. Cheap operation.
Parse a document and return its heading outline. Call this first. Local path or http(s) URL, format detected automatically. A short document returns as `text` — the whole thing, with `is_complete: true`. There is nothing left to fetch: `read` and `search` on it would return only what you are already holding. Answer from it directly. Everything else returns the outline, not the text. Each line carries the tokens `read` would return for that heading.
Read content under specified headings from a parsed document. Provide one or more heading paths (as returned by parse or search). Returns chunk text for each heading, including all subsections.
Search for headings in a parsed document by semantic query. Use this when no heading in the outline looks relevant to your question. Returns headings to read, ranked by relevance, with snippets showing context.
Export a document as markdown and write to disk. Accepts either a source (path/URL) or doc_id. Returns a file path. Nothing enters your context, so document size does not matter.
Output schemas are not explicitly documented. Descriptions state what parse() returns ('outline' with token costs, or full text with is_complete:true) and read/search return ('chunks', 'headings with snippets'), but formal JSON Schema definitions for response structures are absent from visible source. This forces LLMs to infer response field names and types, risking misuse.
Error handling is implicit in descriptions but not formalized. The parse() description mentions 'everything else returns the outline, not the text' and read/search have hints, but there is no explicit error classification (retryable, user-fixable, fatal) or recovery guidance (e.g., 'if search returns no results, try a broader query'). Agents lack guidance on how to recover from failures.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | <=2025-11-25 | v2 |
to_markdown tool accepts 'source' parameter that can be 'path/URL or doc_id', creating type ambiguity. The description attempts to clarify this ('Accepts either a source (path/URL) or doc_id') but the parameter itself is a single untyped string. Best practice would split this into two parameters (source_path, doc_id) with mutual exclusion documented, or validate internally and return clear error messages distinguishing which was attempted.
No pagination guidance in read() and search() return descriptions. If read() returns many chunks under a heading, or search() returns many ranked headings, there is no documented limit, pagination token, or count field. Large result sets could exhaust context windows without the LLM knowing.
password parameter is optional and unvalidated in the visible schema. If a document is encrypted and password is omitted, the error is not documented. Should state: 'Required for encrypted PDFs/Office documents. Returns error if document is encrypted and password is missing or incorrect.'