MCP server for the Wikipedia racing benchmark. Exposes the benchmark as MCP tools: start_episode, get_page, click_link, search_page, score_episode.
This MCP server exposes a Wikipedia racing benchmark with 6 tools. While tool naming follows verb-noun convention and schemas are present with type definitions, the server has significant gaps in description quality, parameter documentation, and error handling. Most tools have minimal descriptions (under 50 chars), and critical parameters lack context. Output schemas are not documented. The benchmark is read-only by design, which simplifies some concerns but does not excuse incomplete documentation for agent planning.
Attempt to click a link on the current page.
Return current page observation.
Return the full text of a specific section.
Score a completed episode.
Search for a query within the current page text.
Start a new episode. Seed indexes into the pre-sampled catalog.
Missing output schemas: No tool documents what fields are returned or their types. Agents cannot plan downstream operations or extract required data without reverse-engineering responses.
Descriptions too brief and lack LLM guidance: 5 of 6 tools have descriptions under 50 chars. Descriptions do not explain WHEN to use the tool, WHAT it returns, or HOW it affects state. Example: 'Return current page observation' does not say what fields the observation contains.
Parameter descriptions lack constraint guidance: 'seed', 'title', 'query', and 'section' parameters have descriptions but do not specify valid ranges, formats, or how to determine valid values. Example: 'title' parameter does not explain how to learn valid link titles from page_observation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 46 | - | v1 |
No error handling guidance: Tool descriptions do not indicate what can fail, how failures are reported, or what the agent should do next. Example: What happens if click_link is called with an invalid link title? What does the error response look like?
Episode lifecycle unclear: No tool or documentation explains the episode state machine or validity window. get_page, click_link, get_section, and search_page all require episode_id but do not document whether an episode expires, what triggers completion, or whether operations are replay-safe (idempotent).
Naming of 'seed' parameter is unclear: Parameter is described as 'Optional seed index into the pre-sampled episodes catalog' but does not explain what happens if omitted, what the valid range is, or whether a random episode is selected. 'seed' typically refers to a random number seed, not an index.