Semantic code search and RAG (Retrieval-Augmented Generation) for local repositories with embeddings, reranking, and LLM inference
This is a STDIO-only MCP server with 8 tools that expose ML operations (embeddings, text generation, reranking, model management). Definition quality is mediocre across the board. Most tool descriptions are present but generic, they lack depth, constraints, and guidance on when to use each tool vs. alternatives. Parameter descriptions are sparse or missing entirely. No input schemas are visible in the provided source code, tool definitions appear to be inferred from test/stub references (jsonrpc_tests.rs and python server.py) rather than explicitly registered. The server mixes critical operations (load_model, unload_model, shutdown) with read-only tools without clear permission gating or confirmation patterns. Error handling guidance is absent. The Python backend (srag_ml/server.py) shows ML service architecture but does not reveal full JSON-RPC schema definitions needed to validate parameter typing and constraints.
Generate embeddings for a batch of text inputs using FastEmbed models
Generate text using either local LLM or external API based on configured provider
Load ML models into memory (embedder or LLM)
Query the current status and memory usage of loaded models
Health check endpoint to verify ML service is operational
Rerank a list of documents based on relevance to a query
Gracefully shut down the ML service
Tool definitions inferred from test files and Python source, not explicitly registered in visible Rust code. No explicit JSON-RPC method registration with complete schemas found in source.
No input schemas visible for 'ping', 'model_status', 'shutdown'. These tools show empty input objects {} but no validation constraints, required fields, or type information.
Parameter descriptions are minimal or absent. 'type' parameter in load_model and unload_model lacks constraint details (what exactly are 'embedder' vs 'llm'? are there other types?). 'path' parameter in load_model has no description of expected format, file extension, or precedence rules.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 38 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Unload ML models from memory to free resources
Destructive operations (load_model, unload_model, shutdown) lack confirmation or dry-run patterns. No error recovery guidance provided. 'shutdown' kills the service with no confirmation, agents could invoke this accidentally.
No permission gating or scope declarations. Tools like 'shutdown' and 'load_model' (which modifies service state) have no access control checks visible. No audit trail or permission-gate pattern evident.
Tool descriptions are generic and lack context on WHEN to use each tool vs. alternatives. 'Generate text' is vague, when should an LLM use 'generate' vs querying an external API? The description does not explain the trade-off or when each provider (local vs external) is appropriate.
'embed' tool description mentions 'FastEmbed models' but does not specify which models are available, supported languages, or embedding dimensions. 'rerank' does not document ranking algorithm or scoring scale.
No output schemas documented. Tools like 'generate', 'embed', 'rerank', 'model_status' do not specify return types, field names, or data structures. LLMs cannot plan downstream tool calls or extract data reliably without knowing what fields to expect.
Parameter naming inconsistency. 'generate' uses 'stop' (a string), but typical LLM APIs use 'stop_sequences' (array) or similar. 'embed' uses 'texts' (plural), but 'generate' uses 'prompt' (singular), inconsistent naming across similar tools confuses LLMs.
Error handling guidance absent. No error recovery patterns, retry hints, or actionable error messages visible. If 'load_model' fails due to missing model file, the error response does not suggest calling 'download_model' or checking the models directory.