Local RAG MCP Server - Easy-to-setup document search with minimal configuration
mcp-local-rag demonstrates strong definition quality with 9 well-structured tools. All tools have explicit names following verb_noun convention (query_, ingest_, delete_, list_, read_, sync_), detailed descriptions (100-300 chars), and complete JSON Schema input definitions with type constraints and parameter descriptions. Tool output schemas are documented in descriptions. Key strengths: consistent naming patterns, comprehensive parameter constraints (enums, min/max bounds), and clear error recovery guidance (e.g., scope prefix matching, visual vs text-only strategies). Weaknesses: no tool annotations (readOnlyHint/destructiveHint) despite clear risk categories (READ_ONLY, WRITE, DESTRUCTIVE); output schema details are described narratively rather than in formal schema blocks; some descriptions could be more concise; error handling is implicit rather than explicit with guided recovery messages.
Delete a previously ingested file or data from the vector database. Use filePath for files ingested via ingest_file, or source for data ingested via ingest_data. Either filePath or source must be provided. Returns deleted (operation succeeded), removedChunks, and existed (whether anything was actually present).
Ingest in-memory content as a string (use ingest_file for files on disk). The source identifier enables re-ingestion to update existing content. Returns { filePath, chunkCount, timestamp, fileTitle }.
Ingest a document file (PDF, DOCX, TXT, MD) into the vector database. Path must be absolute; re-ingesting the same path replaces its existing data. Returns { filePath, chunkCount, timestamp, fileTitle }.
List supported files (PDF, DOCX, TXT, MD) under the configured base directories and whether each is ingested. Returns { baseDirs, files, sources }; sources lists ingested items reported apart from the file scan, chiefly ingest_data content (web pages, clipboard, etc.).
Search ingested documents with hybrid keyword + semantic matching. Use the returned order as the ranking; score may disagree with it. Each has filePath, chunkIndex, text, fileTitle, score (lower is closer), and source (for ingest_data items).
Tool annotations missing: no readOnlyHint, destructiveHint, or idempotentHint declarations despite clear risk categories (READ_ONLY, WRITE, DESTRUCTIVE) marked in source comments.
Output schema details are narrative-only (e.g., 'Returns { filePath, chunkCount, timestamp, fileTitle }') rather than formal JSON Schema blocks in the tool definition. LLMs cannot parse unstructured output descriptions reliably.
Error handling is implicit. Tools do not return structured error messages with actionable recovery guidance (e.g., 'Invalid scope: prefix must be absolute path. Did you mean /docs/api?'). Agents cannot self-correct on failures.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 65 | - | v1 |
Read the chunks immediately before and after a query_documents result, in the same document, for more surrounding context. Pass chunkIndex from the result plus exactly one of filePath (ingest_file) or source (ingest_data). Returns the target chunk (isTarget: true) and its neighbors, ascending by chunkIndex; an out-of-range chunkIndex returns []. Defaults: before=2, after=2 (max 50 each).
Get index status: { documentCount, chunkCount, memoryUsage (MB), uptime (s), ftsIndexEnabled, searchMode }.
Reconcile the index with the files on disk: ingest new and changed files, leave unchanged files alone, and remove index entries for files that are gone. Each changed PDF is re-ingested with the visual profile ("fast" or "quality") already recorded for it, so a PDF indexed with VLM captions keeps them; a PDF with no recorded profile stays text-only. There is no option to change a profile here — use the CLI (mcp-local-rag sync --visual) to set one, or ingest_file to replace the file, where a normal ingest clears the recorded profile. Stored images are unrelated: STORE_IMAGES applies to whatever this run re-ingests and never makes a file changed. Returns { jobId } without waiting for the run to finish; poll sync_status with that jobId for progress and the final outcome. Only one job is kept, and it is lost when the server process exits.
Get the current or latest sync job record: { jobId, state ("running" | "succeeded" | "failed"), total (null until scanning has counted the files on disk), completed (upserted + skipped + empty; pruned is counted separately), summary { upserted, skipped, empty, pruned }, warnings, error (null unless the job failed) }. An unknown jobId means the job was replaced by a newer one or lost with a previous server process.
Descriptions for status and sync_* tools are terse (<80 chars) and lack guidance on WHEN to call them or dependencies (e.g., 'sync_status requires jobId from sync_start'). LLMs may invoke them redundantly or in wrong order.