This server has moderate definition quality with significant gaps. Tool naming is generally action-verb based and clear (health_check, get_mcp_metrics, run_tests, list_testcases, etc.), which is positive. However, descriptions are often translated from Chinese and lack the depth needed for LLM planning. Most critically, parameter descriptions are sparse or missing entirely, and input schemas lack proper type annotations and constraints. Only 2 of 10 tools have reasonably documented schemas; the rest rely on vague descriptions like 'object or null' or 'object or dict or string' that provide zero validation guidance. Error handling exists but is generic, responses include error_code and error_message fields but lack actionable recovery guidance. The server also lacks proper enum constraints for multi-value parameters (e.g., test_type accepts 'all|integration|unit' as free-form strings, not enums). Output schemas are partially documented via Pydantic models (HealthResponse, McpMetricsResponse, RunResponse) visible in the code, but these are not exposed in tool registration metadata, so LLMs cannot infer output structure from tool definitions alone.
删除指定 YAML 测试用例文件及其对应的 pytest 脚本。
返回最近 MCP 工具调用的成功率、耗时与错误分布。
获取指定 run_id 的测试结果历史记录。
获取指定 YAML 测试用例内容,同时返回校验结果(整合读取 + 校验)。
返回服务版本与基础路径信息。
列出指定目录下的 YAML 测试用例文件,支持集成测试 (testcase) 和单元测试 (unittest)。
执行单个或批量 YAML 测试用例,支持集成测试 testcase 和单元测试 unittest。
校验 YAML 测试用例结构是否规范,返回详细的错误列表。
Parameter type annotations are vague or missing. Examples: 'object or null', 'object or dict or string', 'string or null' without specifying required object structure, valid enum values, or format constraints. This violates the constraint that parameters must have strict type definitions.
No enum constraints for parameters that accept a fixed set of values. 'test_type' accepts 'all|integration|unit' and 'mode' accepts 'summary|full', but these are documented in descriptions as free-form strings, not as enum schemas. LLMs cannot reliably select valid options without explicit enum declarations.
Parameter descriptions are minimal or missing. Examples: 'Optional, specify project root directory' (for workspace), 'Specification for Python interpreter path' (for python_path), 'Maximum record count, default 500' (for limit). None explain WHY these parameters are needed or what they control in the context of the operation.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
写入 YAML 测试用例并生成 pytest 脚本,或仅重新生成已存在 YAML 对应的 pytest 脚本。
根据输入的单元测试结构写入 YAML 文件,并生成对应的 pytest 单元测试脚本。
Output schemas are not documented in tool definitions. While Pydantic models (HealthResponse, McpMetricsResponse, etc.) exist in the code, tool registration via FastMCP does not expose output schema metadata. LLMs cannot plan downstream tool calls because they don't know what fields to expect from responses.
Error handling is generic and lacks recovery guidance. Error responses include 'error_code' and 'error_message' but do not suggest next steps. Example: when 'workspace' parameter is missing, the tool should return 'Workspace is required. Specify an absolute path to your project root.' instead of a bare error code.
write_testcase and write_unittest accept 'testcase' and 'unittest' parameters described as 'object or null' with nested structure hints in descriptions only. No explicit JSON Schema for nested fields (name, description, requests, assertions, etc.). Nested objects without schema definitions force LLMs to guess the structure.
Descriptions are translated from Chinese and sometimes lack clarity for English-speaking LLMs. Example: 'Write a YAML test case and generate pytest script, or only re-generate the pytest script for an existing YAML' is ambiguous, does 'existing' mean the YAML or the pytest script? How does the tool decide which action to take?
Destructive tool (delete_testcase) has no confirmation or dry-run mechanism. An LLM could accidentally delete test files. The description states 'Delete the specified YAML test case file' but does not warn that this is irreversible or offer a safety pattern.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in FastMCP registration. While risk markers are documented in the provided tool list (READ_ONLY, WRITE, DESTRUCTIVE), they are not declared as tool properties in the MCP schema, so LLMs cannot reason about safe retry behavior or side effects.