MCP server for building documentation context, searching indexed documents, and analyzing documentation governance, research provenance, and relationships across markdown, DOCX, HTML, and PDF files.
DocGraph presents well-structured tool definitions with clear, detailed descriptions and explicit JSON Schema input parameters. All 4 tools have descriptions exceeding 100 characters with thoughtful usage guidance. Schemas are present and typed for all parameters. However, output schemas are not formally documented, and error handling lacks recovery guidance. The server demonstrates strong naming clarity and comprehensive parameter documentation with constraints (enums, limits, defaults), placing it in the 'good' category despite lacking some A-tier polish.
PRIMARY TOOL. Build relevant documentation context for a task or topic. Composes governance-aware search + node details + cross-references + bounded source content in one call. For a single known document, use docgraph_node instead. For broad queries, set includeContent=false or reduce maxContentBytes (default 2000, hard cap 6000) to avoid large responses; a 10-node query with default settings can produce 20–50 KB of output.
List all indexed files (.md, .docx, .html, .pdf). Use path filter to narrow scope (bare directory name, e.g. path=docs). For a single known doc, use docgraph_node instead.
Get a single document or heading's full details: metadata, structure, and cross-references. Use 'section' to read the full content of a specific heading section from the source file. For multiple documents, use docgraph_context instead.
Find documents topically similar to a given document using TF-IDF term overlap + shared references + tag overlap (engine=auto/tfidf — the default, always on, no flags). Returns 0 results for a topically unique document (a broad README or changelog commonly has no similar_to edges even when the index is fully built and the engine is working): 0 does NOT mean the engine is off, embeddings are disabled, or the index is broken. Neural similarity is an OPTIONAL add-on layered on top — only if embeddings were stored via docgraph_embeddings action=store (engine=neural) are neural scores added; embeddings being disabled never causes a TF-IDF 0-result. For explicit link tracking use docgraph_graph. Accepts document paths only — heading anchors (doc.md#heading) return empty. The score is a 0-to-1 weighted blend (TF-IDF cosine 50% + shared-reference Jaccard 30% + tag Jaccard 20%); it is NOT a percentage. Each result shows the three signal components that drove its score. No per-vocabulary-term breakdown is available — the engine does not retain individual term contributions, so you cannot identify which specific terms, phrases, or mentions made a score high OR low; any per-term explanation of the TF-IDF component is fabricated. Scores are corpus-relative; 0.4-0.5 can mean near-identical in a corpus with high shared vocabulary.
Output schemas not formally documented. Tools return structured data but lack explicit JSON Schema definitions for response objects. LLMs cannot reliably parse return types or plan downstream tool chains.
Error handling lacks recovery guidance. No tool description indicates what to do if a search fails, a document is not found, or filters return zero results. Error responses are likely bare status codes rather than actionable guidance.
Parameter 'document' in docgraph_node accepts multiple formats (path, heading-qualified name) but only the description hints at format conversion ('strip the trailing #heading:line suffix'). Ambiguous input format invites LLM errors.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 72 | 2026-07-28+ | v2 |
No pagination explicit in docgraph_files. Although a 'limit' param exists, no 'offset'/'next_cursor' mechanism is documented. Unbounded result sets risk context window exhaustion.
docgraph_similar output score (0-1 'weighted blend') is explained in description, but calculation basis is deliberately opaque ('No per-vocabulary-term breakdown is available'). LLMs cannot reason about confidence or decide when scores are trustworthy.