AI-driven information aggregation system for tracking academic and social trends via MCP protocol
Horizon MCP has 11 tools with complete JSON schemas and descriptions visible in src/mcp/server.py. However, the server exhibits several definition quality issues: (1) Naming is generally verb-noun style but lacks clarity distinction between similar operations (e.g., hz_score_items, hz_filter_items, hz_enrich_items follow a pipeline pattern that could be better decomposed); (2) Descriptions are present but brief (ranging 40-100 chars), with several falling below the 50-200 char sweet spot for LLM-optimal guidance; (3) Parameter schemas are well-typed with descriptions but lack constraint documentation (enums, ranges, patterns) that would prevent invalid inputs; (4) No output schemas are documented, tools return structured data but the response format is not formally specified in tool definitions; (5) Error handling is generic (HorizonMcpError base class with code/message/details) but lacks recovery guidance specific to each tool; (6) Parameter relationships are undocumented (e.g., hz_run_pipeline combines 5 operations, unclear which params apply to which stages). The server is functional and follows basic patterns but falls short of production-grade quality for autonomous agent use.
Enrich filtered items into the enriched stage.
Fetch and deduplicate content into the raw stage.
Filter scored items into the filtered stage.
Generate a markdown summary from a stage.
Get server metrics and performance data.
Read run metadata.
Read items from a run stage.
hz_run_pipeline violates single-responsibility principle by combining fetch → score → filter → enrich → summarize into one tool. LLM cannot control individual stages or recover from mid-pipeline failures. Should decompose into composable tools or add granular control parameters (skip_score, skip_filter, etc.) with explicit documentation.
Output schemas are not documented in tool definitions. hz_fetch_items, hz_score_items, hz_filter_items return structured data but the response schema (field names, types, structure) is not specified. LLMs cannot plan downstream calls or extract fields without trial-and-error.
Parameter relationships are undocumented. hz_run_pipeline accepts both 'enrich' (boolean) and 'source_stage' defaults throughout pipeline, but it's unclear which parameters apply to which stage. 'threshold' applies only to filter stage; 'language' only to summarize. This forces trial-and-error.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 46 | - | v1 |
List recent runs and stage states.
Run fetch -> score -> filter -> enrich -> summarize in one call.
Score a stage into the scored stage.
Validate Horizon config and required environment variables.
Descriptions are too brief and lack LLM-optimized guidance. hz_score_items ('Score a stage into the scored stage') does not explain WHEN to use it, WHAT scoring means, or WHAT to expect in output. Baseline is 50-200 chars optimal; many tools are 40-75 chars (too terse).
No enum or constraint documentation for string parameters. 'language' in hz_generate_summary defaults to 'zh' but no enum of valid languages is specified. 'stage' in hz_get_run_stage has no list of valid stage names (raw, scored, filtered, enriched, etc.). Forces agent trial-and-error.
No error recovery guidance. Generic HorizonMcpError class returns code + message but does not guide LLM on next steps. E.g., if 'run_id not found', should suggest hz_list_runs; if 'stage does not exist', should list available stages. Errors are not actionable.
No confirmation/dry-run for destructive operations. hz_score_items, hz_filter_items, hz_enrich_items, and hz_run_pipeline modify state (WRITE risk) with no dry-run or confirmation step. Agents can trigger data loss without safeguards.
Naming confusion among pipeline stages. hz_fetch_items, hz_score_items, hz_filter_items, hz_enrich_items, hz_generate_summary all operate on implicit 'stages' with 'source_stage' parameters (raw, scored, filtered, enriched). The stage lifecycle is not explicitly documented, and similar names (score vs scored) risk LLM confusion.