Exposes RAGScore functionality to AI assistants via Model Context Protocol (MCP). Enables QA dataset generation from documents and RAG system evaluation.
RAGScore MCP exposes 2 tools with clear, domain-specific purposes. Tool naming follows verb-noun convention (generate_, evaluate_). Descriptions are detailed and context-aware, ranging 140 - 280 characters, within the 10 - 1024 baseline and above the median (194 chars). Parameter descriptions are present and mostly specific. However, critical gaps emerge: (1) Output schemas are not documented in the source code, the tool responses are described verbally but return types are not formally declared; (2) Error handling guidance is absent, no indication of retryable vs. fatal errors or recovery paths; (3) Security considerations are mentioned but not enforced in schema (e.g., credentials, API keys in parameters). Parameter schemas are well-structured with types, defaults, and nullability, but lack explicit enum constraints for the `provider` and `purpose` parameters, which should be enumerations. The tools are domain-appropriate and composable (generate QA → evaluate RAG), but lack production-grade robustness in error paths and response guarantees.
Evaluate a RAG API endpoint against a QA dataset. Queries the RAG endpoint with questions and scores the answers using LLM-as-judge (1-5 scale).
Generate QA pairs from documents for RAG evaluation or synthetic data. Scans documents (PDF, TXT, MD) and generates question-answer pairs that can be used to test RAG systems or as training data.
Output schemas not documented. Neither tool declares its return type, structure, or fields. LLMs cannot infer what data is returned, how to chain tools, or what fields downstream tools require.
No error handling guidance. Tool descriptions do not specify what errors can occur, whether they are retryable, or what the LLM should do next (e.g., 'If endpoint is unreachable, verify RAG service is running').
Missing enum constraints for provider and purpose parameters. Both tools accept 'provider' (openai, anthropic, ollama) and 'purpose' (training, faq, compliance, fine-tuning) as free-form strings instead of enumerations. LLMs will hallucinate unsupported values.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No tool annotations for side effects. Neither tool declares whether it is read-only, destructive, or idempotent. LLMs cannot reason about retry safety, generate_qa_dataset writes files, but no hint warns the agent.
Result limits not enforced or documented. evaluate_rag may return hundreds of scored Q&A pairs without pagination. No mention of result caps, offset, or next_cursor in the tool description.
API credentials/keys accepted as optional parameters. Both tools accept provider and model but do not document where API keys come from. If LLMs pass credentials directly, they enter logs and context.