MCP server for ingesting Obsidian markdown documents into PostgreSQL with vector embeddings and semantic search capabilities
The server demonstrates solid foundational quality with 12 well-defined tools, clear naming conventions, and comprehensive descriptions. Most tools follow verb_noun patterns (query_documents, get_documents_by_tag, list_documents, etc.). Descriptions average 180+ characters and explain WHAT the tool does and WHEN to use it. Parameters are well-described with type information. However, there are notable gaps: output schemas are not explicitly documented in the provided code, error handling guidance is minimal, and some parameter relationships are underdocumented. The server accepts flexible input formats (JSON strings or dicts for filters) which is user-friendly but increases validation complexity. Most tools are READ_ONLY with clear risk classification, but destructive operations (delete_vault, update_vault, ingest) lack explicit confirmation patterns in some cases. Tool composition is strong, tools form clear chains (search → get_document → retrieve details). The code shows high quality (strict ruff linting, mypy type checking, comprehensive testing) but MCP tool definition quality depends heavily on what LLMs see in descriptions and schemas, not internal implementation details.
Delete a vault and all associated data. This operation is irreversible and cascade-deletes all associated documents, tasks, and chunks. Requires explicit confirmation via confirm=True parameter.
Get all tags in the knowledge base with optional filtering and pagination.
Get a single document by exact file_path within a vault, or by UUID document_id. Retrieves full document content, metadata, tags, and obsidian_uri. Either vault_name+file_path or document_id must be provided.
Query documents filtered by frontmatter properties with include/exclude semantics. Returns documents matching the specified property filters.
Query documents filtered by tags with include/exclude semantics. Returns all documents matching the specified tags with flexible matching logic (all or any).
Output schemas not explicitly documented in tool definitions. While parameter input schemas are present and well-typed, the tools do not include documented return schemas (e.g., what fields query_documents returns, what the structure of results looks like). LLMs need to know expected output fields to plan downstream calls.
Destructive operations (delete_vault, update_vault with container_path change) lack explicit pre-flight confirmation. While delete_vault has a 'confirm' parameter requiring True, the description says 'returns an error dict' if confirm=False, but this error handling is not shown in the schema. A safer pattern is MRTR (Multi Round-Trip Request) with result input_required to force explicit user confirmation before irreversible actions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 11 | - | v1 |
Query tasks from Obsidian documents with optional filtering by date ranges and tags.
Get a single vault by name or ID. Retrieves vault details including document count. Either vault_name or vault_id must be provided, with vault_name taking precedence if both are given.
Ingest documents from an Obsidian vault into the database. Processes markdown files, extracts metadata, and stores embeddings for semantic search.
List documents by file_name with optional vault scope. Returns all documents matching the exact file_name. When vault_name is provided, results are scoped to that vault only.
List all vaults with pagination.
Semantic search over document content with optional filters. Large documents are split into chunks for better embedding quality. The search queries all chunks and returns the best matching chunk per document, with document relevance determined by the highest chunk similarity score.
Update a vault's properties. The name field is used for lookup only and cannot be changed. Changing container_path requires force=True as it deletes all documents, tasks, and chunks for the vault.
Filter parameter accepts both JSON strings and dicts (filters as 'object' or 'JSON string'). While flexible, this creates ambiguity: the LLM must decide whether to pass a dict or a JSON-encoded string. The parameter description says 'Can be passed as a JSON string or a dict' but does not clarify which is preferred or how the tool handles both. This risks malformed input and confusion.
Tag-related parameters in get_documents_by_tag and filters in other tools reference a 'match_mode' field and tag format rules (e.g., 'Tags should NOT include the # prefix') but match_mode is not exposed as an explicit parameter in all tools where it applies. The description mentions it but the schema does not declare it as a selectable enum (all/any), leaving the LLM unable to control matching logic.
Error handling and recovery guidance is minimal. Tools accept optional output_file parameters to write results to local or S3 storage, but there is no documented error handling (what if S3 write fails? What does the agent do?). Similarly, ingest tool has no guidance on what happens if a file is unreadable or embedding fails.
Large result limits (max 10000 items) without guidance on pagination. While limit/offset are supported, the tool descriptions do not advise when pagination is necessary or what the practical limits are for LLM context. Returning 10,000 results will exhaust context and degrade reasoning.
Tool names like 'get_documents_by_tag' and 'get_documents_by_property' are similar but distinct. No guidance in descriptions about when to choose one vs the other. An LLM may conflate them or call both redundantly.