Typed tools for corpus, strategy, and benchmark work on retrieval architectures. Compare retrieval strategies on a corpus with reproducible evidence.
KB Arena MCP Server has 6 tools with generally clear names and reasonable descriptions, but significant gaps in schema completeness, parameter documentation, and error handling guidance. All tools are read-only or stateful with clear risk classifications, which is good. However, input schemas lack detailed parameter descriptions and type constraints. Description quality is uneven, some tools (list_corpora, start_benchmark) provide helpful context about behavior and next steps, while others (job_status, read_manifest) are terse. No enum constraints are visible for multi-value parameters like 'strategy', 'split', or 'tier'. Output schemas are not documented. Error handling descriptions are minimal, tools indicate when they raise but do not guide recovery or explain error categories. Security model is reasonable (no credentials in params) but lacks explicit permission declarations or audit trail documentation.
Poll the status and outcome of a background benchmark job.
List every corpus under the configured datasets root, with pipeline status. Raises if the configured root itself is missing, so a misconfigured KB_ARENA_DATASETS_PATH reads as an error, not as "no corpora exist".
List every built-in strategy and its runtime status. Reads `kb_arena.strategies.catalog.STRATEGY_CATALOG` fresh on every call, so a strategy added to the catalog shows up here without a code change to this tool.
Read a benchmark run's manifest file to inspect its configuration, strategies, corpus, and benchmark parameters. Use this before citing a run's numbers to verify what was actually measured.
Start a benchmark run in the background and return a job id to poll. A benchmark can run for minutes, so this schedules the run and returns right away. Poll `job_status` with the returned job id for the outcome.
Check whether a corpus exists under the configured root and is buildable. A corpus that does not exist is a normal, reported outcome (`valid: false` with a reason). An unsafe corpus name (path traversal) raises instead, because that is not a validation result, it is a rejected request.
Parameter descriptions missing or incomplete. 'start_benchmark' has 5 params with default values but descriptions lack guidance on valid ranges, enum constraints, or interdependencies. E.g., 'tier' is described as 'Benchmark tier/complexity level' with default 0, but no min/max or valid values stated. 'strategy' accepts comma-separated list but no examples or constraint documentation. 'split' parameter has empty default and no description of valid values.
Output schemas not documented. No tool explicitly documents what fields are returned or their types. LLMs cannot plan downstream operations or extract needed data (e.g., job_id from start_benchmark, status from job_status) without seeing output structure in code or README.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 61 | 2026-07-28+ | v2 |
Error handling lacks recovery guidance. Descriptions state when tools 'raise' or return errors (e.g., 'validate_corpus: raises if unsafe corpus name') but do not categorize errors as retryable/user-fixable/fatal or suggest next steps. E.g., 'list_corpora raises if KB_ARENA_DATASETS_PATH missing' but no guidance to LLM on how to fix or recover.
Terse descriptions for job_status and read_manifest. 'job_status' is described as 'Poll the status and outcome of a background benchmark job' (55 chars) and 'read_manifest' as 'Read a benchmark run's manifest file...' (44 chars). These fall below the 10 - 1024 char baseline, and lack guidance on when to call, what to expect, or how the tool fits into the workflow.
start_benchmark uses free-form string parameters where enums would help. 'strategy' parameter accepts comma-separated list with no enum constraint or documented valid values. LLM has no way to know which strategies are valid without calling list_strategies first. 'split' parameter has empty default and no enum, unclear what valid values are.
No pagination or result limits documented. 'list_corpora' and 'list_strategies' may return many items but no limit, offset, or pagination mechanism visible in schemas. For large deployments, these tools could overwhelm context windows. No documentation of typical list sizes or truncation behavior.
No idempotency guarantees for start_benchmark. Description states tool 'schedules the run and returns right away' but does not confirm whether calling twice with identical parameters returns the same job_id (idempotent) or spawns duplicate runs. Agents may retry on ambiguous failures and spawn unintended duplicate benchmarks.
No scope declarations or permission gates. Tools do not declare what permissions they require (e.g., 'read:corpus', 'write:benchmark'). 'start_benchmark' is a write operation but no description of required roles or audit trail. Cannot configure least-privilege agent access.