Pytest-style framework for evaluating Model Context Protocol (MCP) servers.
MCP Eval is an evaluation framework that hosts 13 sample tools across multiple demo servers (calculator, healthcare, YouTube, generic). Tool definitions are visible with schemas and descriptions, but quality is inconsistent. Most tools have adequate naming conventions (verb_noun pattern: get_*, search_*, summarize_*) and descriptions (50-150 chars), but many descriptions lack actionable depth, parameters are minimally described, and output schemas are not documented. No evidence of error handling guidance, recovery patterns, or security considerations. The framework itself is well-engineered (pytest plugin, OTEL integration, evaluators), but the sample tools it provides are illustrative rather than production-grade. Naming is consistently strong; descriptions and parameter quality drag down the overall score.
Look up drug information from the FDA database.
Gets the current time in the specified timezone. For example, 'America/New_York' or 'Europe/London'.
Get a personalized greeting
Get metadata information about a YouTube video.
Get the full transcript of a YouTube video.
Search for medical literature in PubMed database.
Search within a YouTube video transcript.
Special math tools (special_add, special_subtract, special_multiply, special_divide) use non-standard naming ('special_*' prefix). These violate the verb_noun convention and obscure intent. Should be 'add_numbers', 'subtract_numbers', 'multiply_numbers', 'divide_numbers' to immediately convey what they do and enable LLM parsing.
get_greeting has minimal description ('Get a personalized greeting', 29 chars). This fails to explain when to use it (casual interaction? onboarding? error recovery?) or what makes a 'personalized' greeting. LLMs cannot infer context from vague descriptions.
No tool includes documented output schemas or return type specifications. Tools like fda_drug_lookup and pubmed_search return complex nested results (drug info, adverse events; article metadata, abstracts). Without documented output structure, LLMs cannot plan downstream extraction or pagination. This violates the documented return types baseline (100% of A+ tools have them).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
A special addition tool that adds two numbers and doubles the result
A special division tool that divides two numbers and halves the result
A special multiplication tool that multiplies two numbers and doubles the result
A special subtraction tool that subtracts two numbers and halves the result
Summarizes the provided text to a specified number of words. This is a naive implementation that just truncates the text.
Get a summarized version of a YouTube video transcript.
Parameter descriptions are minimal or missing context. Example: fda_drug_lookup's 'search_type' enum has labels ('general', 'label', 'adverse_events') but descriptions lack guidance on which to use when. pubmed_search's 'date_range' accepts a string like '5' (for 5 years) but does not explain format constraints or valid ranges. LLMs will guess and pass invalid values.
No evidence of error handling, error categorization, or recovery guidance. Tools that call external APIs (FDA, PubMed, YouTube) can fail (network, rate limit, invalid ID). Absence of error handling patterns means LLMs receive raw failures with no actionable next steps.
Tools lack idempotency guarantees. get_greeting, special_add, summarize_text, summarize_transcript are presumably read-only, but no descriptions or annotations declare this. Without idempotent hints, agents cannot safely retry on transient failures.
pubmed_search accepts 'max_results' (1-100 range implied in description) but no minimum/maximum constraints are enforced in schema. Unbounded or weakly constrained numeric parameters allow LLMs to pass invalid values (0, 1000+), causing API errors.
Multi-parameter tools (pubmed_search with query, max_results, date_range) lack dependency or mutual-exclusivity guidance. 'date_range' is optional but its format ('5' for years) is inferred only from the description text. LLMs will guess formats and fail.