Exposes Cherry Evals as tools for any MCP-compatible AI agent. Agents can search evaluation datasets, create curated collections, and export them — all through the Model Context Protocol.
Cherry Evals MCP server exposes 5 well-named search and discovery tools for evaluation dataset access. All tools are READ_ONLY, reducing security risk. Tool names follow verb-noun conventions (list_, get_, search_), and descriptions are present and reasonably detailed (ranging 115-280 chars, within the 10-1024 baseline). However, input schemas lack proper JSON Schema structure documentation, output schemas are not documented, and several critical quality patterns are missing: no pagination return fields (total count, next_cursor), no error recovery guidance, no tool output structure documentation. The tools appear functional but lack production-grade detail needed for confident agent reasoning and error handling.
Get details about a specific dataset. Args: dataset_id: The ID of the dataset to retrieve.
Search for examples using hybrid keyword + semantic search. Combines keyword matching and vector similarity using Reciprocal Rank Fusion (RRF). Falls back to keyword-only if semantic search is unavailable. Args: query: The search query string. dataset_name: Optional filter to search only within a specific dataset. subject: Optional filter by subject. collection: Qdrant collection to search (default: mmlu_embeddings). limit: Maximum number of results to return (default 20, max 100). keyword_weight: Weight for keyword results (0-1, default 0.4). semantic_weight: Weight for semantic results (0-1, default 0.6).
List all available evaluation datasets. Returns a JSON array of datasets with their names, task types, and example counts.
Search for examples across evaluation datasets by keyword. Searches question and answer text. Returns matching examples with their dataset info, choices, and metadata. Args: query: The search query (matches against question and answer text). dataset_name: Optional filter to search only within a specific dataset. subject: Optional filter by subject (e.g. "math", "history"). limit: Maximum number of results to return (default 20, max 100).
Output schemas are not documented. Tools return search results and datasets, but LLMs cannot see what fields are included, data types, or structure. This forces LLMs to guess at the response format and makes downstream tool composition fragile.
Pagination control is incomplete. search_examples and hybrid_search_examples accept 'limit' parameter but there is no mention of 'offset', 'page', or 'next_cursor' in the response. This violates the paginated-result pattern, tools returning lists must return total count or cursor for safe large-result handling.
No error recovery guidance. Tools lack descriptions of when/why they fail and what the LLM should do next. For example, semantic_search_examples requires embeddings to be pre-generated, if they're missing, the error message should guide the user to generate them first.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Search for examples using semantic (vector) similarity. Embeds the query and finds nearest neighbors in the vector database. Requires embeddings to be generated for the target collection. Args: query: Natural language search query. collection: Qdrant collection to search (default: mmlu_embeddings). subject: Optional filter by subject in payload. limit: Maximum number of results to return (default 20, max 100). score_threshold: Minimum similarity score threshold (0-1).
Parameter constraint information is minimal. semantic_search_examples accepts 'score_threshold' (0-1) but the tool description does not explain what happens if the threshold is set to 0.0 vs 1.0, or what 'score' means semantically (cosine similarity? dot product? normalized?). This ambiguity forces LLMs to reason about the scoring model.
hybrid_search_examples accepts both 'dataset_name' and 'collection' parameters, which are related but not clearly distinguished. The tool description does not explain when to use one vs the other, or if both can be used together. This violates the mutual-exclusivity guidance.
No response field documentation. Tools like 'search_examples' return 'matching examples with their dataset info, choices, and metadata' but there is no schema showing which fields are present. What is 'choices'? Is it an array of strings or objects? Do examples include timestamps? LLMs need explicit field documentation.