A comprehensive testing framework for validating LLM tool calling capabilities with MCP services
testmcpy has three tools with reasonable verb-noun naming (load_test_suite, execute_test_case, list_mcp_tools). All tools have descriptions between 100-200 chars, which is acceptable. Input schemas are present with type declarations for all parameters. However, parameter descriptions are minimal (typically 10-40 chars), falling short of the 72-char baseline for A+ tools. No output schemas are documented in the visible source code. Error handling guidance is absent, tools do not describe what to do if they fail, what errors are retryable, or how to recover. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite execute_test_case being a WRITE operation. The tools are narrowly focused (good composability), but parameter typing lacks depth (no enums, format constraints, or min/max bounds for numeric parameters).
Run a single test case against a subject LLM via the MCP service. Specify the test name and prompt, plus the model and provider to test. Returns pass/fail, score, tool calls, evaluations, cost, and token usage.
Discover all tools available on the MCP service. Returns tool names, descriptions, and input schemas.
Load and parse test cases from a YAML/JSON file or directory. Returns parsed test case definitions with their prompts and evaluators.
No output schemas documented. Tools return complex nested structures (test results, tool metadata) but LLMs have no contract for what fields to expect. This forces LLMs to parse responses blindly and breaks downstream tool composition.
Parameter descriptions are too brief (10-40 chars vs 72-char baseline). 'test_path' = 'Path to test file or directory' lacks guidance on format (absolute? relative? file or directory?), 'prompt' and 'model' lack context (what format? what models are valid?), 'provider' has no enum or list of valid values (claude, openai, ...?).
execute_test_case has a WRITE risk profile but lacks error handling documentation. No guidance on retryability, failure modes (invalid model, provider unavailable, test timeout), or recovery steps. LLMs cannot distinguish between transient and permanent failures.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
No tool annotations present. execute_test_case is destructive (modifies test state, records costs/tokens) but lacks destructiveHint=true. list_mcp_tools and load_test_suite are read-only but lack readOnlyHint=true. This forces clients to infer safety from names alone.
Parameters lack enums, format constraints, and bounds. 'model' accepts any string (claude-3-5-sonnet? gpt-4o? hallucinated model names?). 'provider' has no enum (claude|openai|anthropic|...?). 'test_path' has no format hints (absolute path? glob patterns? directory traversal?).
No documentation of tool return values. What fields does load_test_suite return? What is the structure of parsed test cases? Does execute_test_case return pass/fail booleans, numeric scores, arrays of tool calls? Without schemas, LLMs cannot plan downstream operations or extract cited data.