Obsidian RAG MCP demonstrates above-average tool definition quality with consistently complete schemas, good descriptions, and thoughtful parameter validation. All 9 tools are explicitly registered with proper JSON Schema inputs and readable descriptions (85-280 chars). Parameter constraints (min/max values, enum-style filtering) are well-defined. However, the server lacks error recovery guidance in descriptions, does not document output schemas, and offers no tool annotations (readOnlyHint, destructiveHint). The reasoning-layer tools (search_with_reasoning, get_conclusion_trace, explore_connected_conclusions) are sophisticated but their output structures are undocumented, forcing LLMs to infer result types. Naming follows verb_noun convention consistently. Parameter validation occurs at runtime but descriptions lack explicit guidance on what happens on validation failure.
Explore conclusions related to a query or another conclusion. Use this to discover what else the system knows about a topic or find connections between different pieces of information.
Get the reasoning trace for a specific conclusion. Shows the evidence chain: source chunk -> conclusion -> related conclusions. Use this to understand WHY a conclusion was drawn and what evidence supports it.
Get the full content of a specific note by its path. Use this when you need the complete document, not just a snippet.
Find notes related to a given note. Useful for discovering connected information or similar past incidents.
Check the status of the search index. Shows number of files indexed and other statistics.
List recently modified notes in the vault. Useful for finding recent RCAs or documentation updates.
Output schemas are completely undocumented. Tools return results (chunks, notes, conclusions with reasoning traces, metadata) but callers cannot know the structure without reading implementation. LLMs cannot reliably extract fields like 'confidence', 'evidence_chain', or 'related_conclusions' from responses.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). All tools are read-only (as intended), but this is not explicitly signaled in the tool metadata. Agents cannot distinguish safe tools from potentially dangerous ones without parsing descriptions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 49 | - | v1 |
Search for documents with specific tags. Useful when you know the category but want to explore related content. If no query is provided, tags are used as the semantic search query.
Search the Obsidian vault semantically. Returns relevant document chunks based on meaning, not just keyword matching. Use this to find information about specific topics, past incidents, or documentation.
Search the vault with reasoning layer. Returns relevant document chunks PLUS logical conclusions extracted from the content. Use this when you need not just raw text but synthesized insights and patterns. Only available when reasoning is enabled during indexing.
Error handling descriptions are absent. Tools validate inputs (query length, top_k ranges, path format) but descriptions do not explain what the agent should do on validation failure, how to recover, or which errors are retryable.
Tool composition assumes multi-step reasoning. Agents must chain search_vault → get_note → get_related → search_with_reasoning sequentially. No batch tools exist. Descriptions do not hint at likely follow-up calls or dependencies.
Reasoning layer tools (search_with_reasoning, get_conclusion_trace, explore_connected_conclusions) lack output schema documentation. Descriptions mention 'conclusions', 'reasoning traces', 'evidence chains', and 'confidence' but the actual response structure is unknown. LLMs must guess whether conclusions are returned as objects, arrays, or strings.
Parameter 'conclusion_types' in search_with_reasoning is an enum (deductive, inductive, abductive) but no description explains what these reasoning types mean or when to use each. An LLM cannot intelligently filter conclusions without understanding the distinction.
index_status description is vague: 'Check the status of the search index. Shows number of files indexed and other statistics.' The phrase 'other statistics' is undescribed. Output schema is absent, so the agent does not know what fields will be returned.