PolicyAgents exhibits pervasive structural issues that make it unsuitable for production use. The four tools lack proper input schemas visible in the provided source code, descriptions are present but generic and undifferentiated, and there is no evidence of output schema documentation. The server appears to be a Streamlit frontend wrapping LangGraph agents rather than a properly structured MCP server with formal tool contracts. Tool definitions are inferred from brief Chinese descriptions rather than explicit schema registrations. The codebase shows Streamlit UI patterns (st.set_page_config, st.markdown, session state) and LangGraph orchestration, but no evidence of MCP tool registration with JSON Schema input schemas or documented output types. All four tools accept only a single string parameter ('policy_text') with identical descriptions in Chinese, suggesting copy-paste tool creation rather than careful design. No error handling guidance, no parameter validation rules, and no output schema documentation are visible.
政策解读专家。详细分析新政策弥补了既有政策体系中的哪些空缺、盲点或执行难点,说明对应条款/机制如何填补,指出与既有政策的衔接关系与潜在冲突,给出证据化、可引用的表述。必要时可检索外部资料。
政策意图分析助手。对政策进行意图与要点归纳,包括:1) 政策目标与价值取向 2) 关键举措与制度设计 3) 目标对象与适用范围 4) 与既有政策的关系与边界 5) 可能需要重点解读的术语或口径
政策风险与约束评估专家。客观识别可能的消极影响与风险,包括但不限于:执行成本、制度摩擦、区域与群体不均衡、市场扭曲、道德风险、合规与监管压力、外部性等,提出风险缓释建议。
政策影响评估专家。系统分析政策在宏观与微观层面的积极影响,包括但不限于:经济增长、产业结构优化、创新驱动、民生改善、治理效能提升等,给出受益主体与影响机制。
No input schemas visible in source code. Tools define only a bare 'policy_text' string parameter with no JSON Schema type annotations, constraints, length limits, or format specifications.
Output schemas not documented. No evidence in the codebase that any tool documents its return type, field structure, or what data the LLM should expect after invocation. This prevents agents from planning follow-up calls or extracting structured results.
Non-specific, uninformative descriptions. All four tools have nearly identical descriptions in Chinese (例: '政策意图分析助手' = 'policy intention analysis assistant') that provide no actionable context about WHEN to call each tool, WHY to prefer one over another, or WHAT the output contains. Descriptions are domain jargon, not LLM-optimized guidance. Baseline for A+ tools: 50 - 200 chars of specific, action-oriented text. These descriptions are generic and repetitive.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 25 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 16 | - | v1 |
Tool naming does not follow verb_noun convention. Names like 'intent_analyst', 'gaps_analyst', 'positive_analyst', 'negative_analyst' are noun-based roles, not action-oriented verbs. Better names: 'analyze_policy_intent', 'analyze_policy_gaps', 'assess_policy_positive_impacts', 'assess_policy_risks'. LLMs infer tool purpose from the name before reading descriptions, these names are ambiguous about whether they analyze, summarize, evaluate, or retrieve.
No error handling or recovery guidance. Source code shows no evidence of tools returning structured errors, categorizing failures as retryable vs. fatal, or suggesting next steps when called with invalid input. Raw exceptions or missing data lead to silent failures.
Parameter lacks constraints and validation rules. The single 'policy_text' parameter accepts any string with no documented length limit, format requirement, or encoding. If a user passes a 10MB policy file or binary data, the tool may hang or crash without guidance on valid input ranges.
Likely tool inference rather than explicit registration. The source code provided is a Streamlit UI (app.py) that imports from main.py and calls run_analysis(). No explicit MCP tool registration, JSON Schema definitions, or Tool objects are visible in the provided code. Tool definitions are inferred from brief parameter descriptions, not formally defined.