Hybrid RAG engine — semantic + keyword search, MCP-native, single binary
rag-ferrite is a well-structured RAG server with 13 tools covering search, ingestion, and analytics workflows. Strengths: all tools have descriptions (34 - 392 chars, well within baseline), all tools have clearly named action verbs (query_, ingest_, delete_, list_, check_, benchmark, read_, collection_, chunk_, suggest_, tag_map). Parameters are typed and mostly described. Weaknesses: output schemas are not documented in the source code (tools return `.to_string()` with JSON via `service::*` functions, but the schemas are not visible in the submitted code); several parameters lack detailed format/range constraints; some descriptions could better explain WHEN to use the tool vs. similar ones (e.g., query_documents vs. suggest_collection); error handling is present but recovery guidance could be more specific.
Evaluate retrieval quality against a versioned golden dataset JSON file. Supports the legacy array format and {version, entries}; returns Recall@k, precision, MRR, nDCG, empty-result rate, latency percentiles, and per-query details.
Pre-ingestion quality check: analyze a document before indexing. Returns char count, estimated chunks, language, duplicate detection, and warnings.
Get chunk-level QA report: identify dead chunks (never queried) and cold chunks. Grouped by source with heat scores calculated on-the-fly. Useful for cleaning up noise.
Get collection heat tracking data: which collections are queried most/freshly. Returns heat_score, last_queried_at, and query_count per collection.
Remove a document and all its chunks by source ID.
Index content directly (text, HTML, or markdown) with a source identifier.
Output schemas are not documented in tool definitions. All tools return `.to_string()` serialized JSON, but the response structure (fields, types, nested objects) is not visible in the submitted code. LLMs cannot plan downstream tool calls or extract specific fields without documented schemas.
Parameter format/range constraints are underspecified. 'tags' is described as array with AND logic, but no max length or tag format spec (regex, charset, examples).
Tool descriptions do not explain WHEN to use each tool over similar ones. For example, query_documents vs. suggest_collection both involve searching content, but the description does not clarify: should I call suggest_collection to pick a collection first, then query_documents? Or is suggest_collection a diagnostic tool? This forces LLM guessing.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 62 | <=2025-11-25 | v2 |
Parse and index a document file (PDF, TXT, MD) into the RAG.
List all indexed documents with their metadata.
Search documents using hybrid search (BM25 + vector with RRF fusion). Returns relevant chunks with scores. Tags use AND logic: 1 tag = broad results, 2 tags = precise intersection. Use 1-2 tags max.
Get chunks adjacent to a specific chunk for context expansion.
Get RAG engine status: document count.
Given a query, extract keywords and suggest the best-matching collection based on tag routing. Returns suggested collection, matched keywords, and all candidate collections with scores.
Show the full tag → collection mapping with chunk counts. Useful to understand which tags belong to which collections.
Error handling in code shows validation (e.g., path_not_allowed, 'Provide either file_path or content'), but descriptions do not hint at these errors or recovery steps. LLMs should know: if ingest_file fails with 'path_not_allowed', call check_ingestion with content directly. This guidance is missing.
Destructive operations (delete_file) lack confirmation/dry-run support. The delete_file tool can remove all chunks for a source ID with no undo. Code validates input but description does not warn about irreversibility or suggest a confirmation pattern.
Parameter descriptions lack examples and constraints in narrative form. E.g., 'source_ids' is 'Filter by source IDs', no mention of format (integers? strings? UUIDs?), cardinality (1 - 10 IDs max?), or how they relate to sources returned by list_files. This forces LLM guessing.