MCP server that exposes the EvalView testing framework as tools for Claude Code. Provides test creation, baseline snapshots, regression checking, and evaluation of AI agent behavior.
EvalView MCP server defines 3 tools with generally strong descriptions and parameter documentation. All tools have non-empty descriptions (194+ chars, well above the 10-char minimum and aligned with baseline of 194 chars average). All tools have comprehensive input schemas with proper JSON Schema types and detailed parameter descriptions. Naming follows verb_noun conventions (create_test, run_check, run_snapshot). However, some issues prevent a higher score: (1) output schemas are not explicitly documented in the source code excerpts provided, the tool descriptions reference return values but no formal output schema is visible; (2) error handling guidance is mentioned in descriptions but not formalized as part of the tool contract; (3) tool annotations (readOnlyHint, destructiveHint, idempotentHint) are absent, create_test and run_snapshot modify state but lack destructiveHint annotation; run_check is read-only but lacks readOnlyHint. (4) The 'forbidden_tools' parameter in create_test is a novel safety pattern not aligned with standard MCP patterns, though well-intentioned. Strengths: comprehensive parameter descriptions with clear constraints, intelligent auto-detection logic documented in descriptions, support for natural identifiers (test names rather than opaque IDs), and clear composition (3 separate tools for distinct concerns). Parameters are well-constrained with enums and numeric bounds where appropriate.
Create a new EvalView test case YAML file for an agent. Call this when the user asks to add a test, or when you want to capture expected agent behavior. After creating a test, call run_snapshot to establish the baseline. No YAML knowledge required — just describe the test. IMPORTANT: Automatically detect test_path by looking for a 'tests/evalview/' directory in the current project. If found, use it. Otherwise use 'tests'.
Check for regressions against the golden baseline. Returns a diff summary for each test: PASSED, OUTPUT_CHANGED, TOOLS_CHANGED, or REGRESSION. REGRESSION means the score dropped significantly — treat this as a blocking failure. TOOLS_CHANGED / OUTPUT_CHANGED are warnings: the agent's behavior shifted but may be intentional. Also returns observability signals: behavioral anomalies (tool loops, stalls), trust scores (benchmark gaming detection), and coherence issues (multi-turn context loss). Use this after any code change (prompt, model, tools) to confirm nothing broke. If you see a regression, show the diff to the user and offer to fix it before moving on. Use heal=true to auto-retry flaky failures and distinguish non-determinism from real drift. IMPORTANT: Automatically detect test_path by looking for a 'tests/evalview/' directory in the current project. If it exists, pass it as test_path. If the project has a custom test location, use that instead.
Run tests and save passing results as the new golden baseline. Use this to establish or update the expected behavior after an intentional change. Future `run_check` calls will compare against this snapshot. Call this: (1) after creating a new test with create_test, (2) after confirming a behavioral change is intentional, (3) before making large refactors so you have a clean rollback point. Only passing tests are saved — failing tests are skipped with a warning. IMPORTANT: Automatically detect test_path by looking for a 'tests/evalview/' directory in the current project. If it exists, pass it as test_path.
Output schemas not formally documented. Tool descriptions reference return values (e.g., run_check returns 'PASSED, OUTPUT_CHANGED, TOOLS_CHANGED, REGRESSION') and observability signals, but no explicit JSON Schema for response structure is visible in the source code. LLMs cannot plan downstream processing without knowing response field types.
Tool annotations missing. create_test (WRITE, modifies files) and run_snapshot (WRITE, destructive reset option) lack destructiveHint annotation. run_check (READ_ONLY) lacks readOnlyHint. These annotations inform MCP clients about risk profile and help with rate limiting, retry logic, and UI warnings.
Error handling guidance not formalized in tool contracts. Descriptions mention error classification (retryable vs fatal) and root-cause analysis, but no explicit error response format or error codes are defined. LLMs cannot reliably distinguish permanent failures from transient ones without a structured error contract.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Parameter 'budget' (run_check) is numeric but lacks explicit bounds documentation. Description says 'Maximum total budget in dollars (e.g. 0.50)' but does not specify minimum or maximum range. LLMs may pass nonsensical values (negative, zero, extremely large).
Safety-critical parameter 'reset' (run_snapshot) lacks explicit confirmation or dry-run guard in schema. While 'preview' param provides dry-run, 'reset=true' directly deletes all baselines without requiring explicit confirmation. Agents may invoke this unintentionally.