A demonstration MCP server with compliance warning assessment, database integration, and multiple tool implementations
This server has 19 tools with significant inconsistency in quality. Strengths: all tools have at least minimal descriptions and input schemas are present for most tools. Weaknesses: descriptions are often terse and lack actionable context for LLM selection; parameter descriptions are sparse or missing; error handling guidance is absent; output schemas are not documented; tool naming is inconsistent (greet_user vs get_greeting; seed_demo_kb uses Chinese description); no security analysis visible in tool definitions. The compliance_warning tools (tools 8-17) show more detailed schemas but still lack error recovery guidance and clear output documentation. Simple arithmetic tools (add, subtract) are well-named but have minimal descriptions. Database tools (18-19) are severely underspecified.
Add two numbers
【第一步】执行合规风险初步筛查。 这是评估流程的入口。它会: 1. 解析业务数据 2. 运行规则引擎检测硬性风险(Signals) 3. 检索相似的历史案例和制度条款(Hits) 注意:此工具返回的是原始证据(Evidence),不包含最终评分。 获得输出后,你通常需要继续调用 `calculate_risk_score` 来计算量化风险。
一键评估工具:使用内置的示例数据运行一次完整的风险评估流程。
【第二步】计算量化风险评分。 必须使用 `assess_compliance_risk` 的输出作为此工具的输入。 它会根据信号严重度和案例相似度,应用数学模型计算客观的风险概率(Probability)和等级(Level)。
获取指定业务系统的示例输入数据(Payload)。支持:decision(议事), procurement(招标), analytics(分析)。
Get all modules from app_module table.
资源获取:根据 ID 获取特定历史案例的详细 JSON 内容。
Seven tools (add, subtract, get_settings, greet_user, seed_demo_kb, query_db, get_app_modules) have descriptions under 30 characters, providing insufficient context for LLM tool selection.
Tool naming is inconsistent. Both 'get_greeting' and 'greet_user' exist for similar purposes (one returns a greeting, the other generates one). Conflicting names force LLMs to reason about subtle differences.
No output schemas are documented for any tool. Rubric states: 'Document the output schema. LLMs need to know what fields to expect.' For example, assess_compliance_risk returns 'Evidence' but the structure is not specified in tool metadata. Similarly, get_app_modules returns results but the schema is completely undocumented.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Get a personalized greeting
资源获取:根据 ID 获取特定制度条款的详细 JSON 内容。
Get application settings.
Get weather for a city.
Generate a greeting prompt
向知识库动态录入一个历史合规案例(审计结果或否决记录)。
向知识库动态录入一条新的制度条款。
Tool that uses initialized resources.
Read a document by name.
获取指定业务系统的输入字段说明,帮助了解需要提供哪些合规审查要素。
初始化知识库,填充演示用的制度条款和历史案例数据。
Subtract two numbers
query_db tool (file: database.py) has zero meaningful definition. Description is 'Tool that uses initialized resources' (39 chars, vague). Input schema is empty ({}). No explanation of what it queries, what database, or what it returns. Violates pattern:tool completely.
get_app_modules has similarly poor definition. Description is 'Get all modules from app_module table.' (40 chars, generic). Input schema is empty. No explanation of output structure, pagination, or filtering options. If results are large, no limit or pagination guidance.
No error handling guidance in any tool description. Rubric pattern:recovery-guide states: 'Error responses must tell the LLM what to do next.' For example, ingest_policy could fail if doc_id is a duplicate or if scope is invalid, but no error guidance is documented. No tool explains retry behavior or when to escalate.
Parameter constraints are under-specified. Example: get_weather accepts 'unit' with default 'celsius' but nowhere is it documented that only 'celsius' and 'fahrenheit' are valid. The schema shows no enum. calculate_risk_score accepts 'signals' as array but does not specify the structure of signal objects.
Destructive operations lack confirmation or dry-run support. seed_demo_kb and ingest_policy/ingest_case are WRITE operations but do not offer a dry-run or confirmation step. Agents make mistakes, a confirm_before_execute pattern prevents catastrophic errors.'
Tool composition issue: assess_compliance_risk and calculate_risk_score are documented as a two-step flow ('【第一步】' and '【第二步】'), but there is no clear way for the LLM to know if it should always call both or if calculate_risk_score is optional. The relationship and required ordering is not explicit in tool-level metadata.