MCP server for mechafil-server API endpoints providing Filecoin economic simulation and historical data analysis
The server exposes 5 tools (with one duplicate: get_historical_data appears twice). All tools have descriptions and input schemas are visible in the Pydantic model definitions. However, several gaps reduce quality: (1) Descriptions lack explicit WHEN/WHY context for agent selection, they explain WHAT the tool does but not when to prefer it over alternatives. (2) Parameters have type hints but descriptions are often generic or under-optimized for LLM reasoning. (3) Output schemas are not formally documented in code, agents cannot predict response structure without calling the tool. (4) Error handling is absent from tool definitions, no guidance on recovery or retry semantics. (5) One duplicate tool (get_historical_data) indicates maintenance debt. The simulate() tool has the most detailed description, but fetch_context() and provide_plot() lack clarity on expected inputs/outputs.
Return the authoritative system prompt text with dynamic documentation inserts. Call once at startup (per session) before any other tool.
Fetch historical network metrics from `/history` endpoint. - Omit `fields` to retrieve all available fields. - Use `fields=['field_name']` to return only specific series. - Always called after `fetch_context()`.
Fetch historical network metrics from `/history` endpoint. - Omit `fields` to retrieve all available fields. - Use `fields=['field_name']` to return only specific series. - Always called after `fetch_context()`.
Build a strict chart specification for the UI to render. Accepts series name(s) or series descriptors; returns a specification for the UI to render as a chart.
Run a MechaFil simulation via `/simulate`. - Always align `forecast_length_days` with the user's horizon. - Use `requested_metrics` (list) to return one or more metrics in a single run (defaults to ['1y_sector_roi']). - Output is Monday-sampled. The response includes an `Explanation` reflecting the actual inputs used after defaults are applied.
Duplicate tool definition: get_historical_data appears twice in the tool list with identical schema and description. This creates ambiguity and suggests the tool roster was not properly deduplicated.
Output schemas are not documented in tool definitions. Agents have no way to know what fields fetch_context, get_historical_data, simulate, and provide_plot return without calling them or reading external API docs. This violates pattern:tool and forces LLMs to guess at response structure.
Tool descriptions lack explicit WHEN/WHY guidance. The simulate() description explains input parameters but does not articulate when an agent should call simulate vs get_historical_data, or what business question each answers. This increases LLM confusion when multiple tools exist.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 23 | 2024-11-05+ | v1 |
provide_plot() has a vague description ('Build a strict chart specification...') that does not explain what 'strict chart specification' means, what UI framework is expected, or what format the response uses. This is under 100 chars and lacks clarity for LLM selection.
The 'series' parameter in provide_plot accepts type=['string','object','array'] but the description does not clarify the structure of each type. An LLM does not know whether to pass a string series name, a dict descriptor, or an array of either. This is under-specified.
No error handling documented in any tool definition. Tools do not declare what errors they may raise, whether they are retryable, or what guidance to offer an LLM on recovery. This violates pattern:recovery-guide.
The simulate() tool accepts 'requested_metrics' as a list but the description says 'Use exact API metric identifiers; if unsure, ask the user to choose from a short list.' This is poor guidance, the tool should either validate metrics server-side and return a clear error, or provide a discovery tool (e.g., list_available_metrics) that agents can call first.
Parameter descriptions for simulate() use lists and defaults extensively but lack explicit constraints. E.g., 'list (len = forecast_length_days)' is informal, the schema should enforce this as a minItems/maxItems constraint, not rely on LLM compliance. Current approach risks malformed input.