Stable since v1.0.0 — MCP server + cross-host agent skill (Claude Code / Codex / OpenClaw / Hermes) for QA testing across pytest / Jest / Cypress / Go / Maestro / Schemathesis / Newman, plus v1.1 Edge AI Runner (RTSP + YOLO). 22 frozen tools, bilingual EN/zh-TW QA knowledge, AI Visual Challenge Solver (reCAPTCHA / hCaptcha), OWASP API Top 10 (2023) scanner, universal qa_plan + verify bookend across 5 core tools (AI 測試大師)
The mk-qa-master server defines 6 tools across multiple test runners (pytest, Jest, Cypress, Go, Maestro). Tool names follow clear verb_noun patterns (get_*, list_*, run_*), which is excellent. Descriptions are substantial and include usage guidance, prerequisites, and edge cases, well above the 194-char baseline. However, there are critical gaps: (1) Input schemas are present but lack type definitions for most parameters; (2) Output schemas are documented in descriptions but not formally declared in the tool registration; (3) Parameter descriptions are dense and often bilingual (English/zh-TW), mixing pattern guidance with implementation details. The three read-only tools (get_runner_info, list_tests, get_test_report, get_failure_details) are well-scoped. The write tools (run_tests, run_failed) include risk annotations and security guardrails (filter validation, timeout handling) but lack explicit error categorization (retryable vs. fatal). The tool composition is generally sound, each tool has a single responsibility, but the schema completeness prevents a higher score.
Extract full root-cause-analysis materials for every failed test in the most recent run. Behavior: - Reads report.json, filters tests where outcome == 「failed」 - pytest: parses Playwright trace.zip → extracts real API call sequence (Frame.*, Page.*, Locator.*, ElementHandle.* events) as steps[] - Maestro: parses flow YAML for `takeScreenshot:` directives → resolves <name>.png at PROJECT_ROOT root - Best-effort resolves screenshot / trace.zip / video / recording paths from --output / --debug-output artifact directories Returns: list[{nodeid, title, message, duration, steps[], screenshot, trace, video}] When to use: - run_tests just reported failed > 0 → drill into each case - User asks 「why did it fail / show me the trace / what broke」 - Filing a JIRA bug → use the artifact paths to attach screenshot+trace - Comparing failure signatures across runs (pair with get_test_history) When NOT to use: - Want the summary count only → use get_test_report (lighter) - No tests have been run yet → returns [{error: 「找不到報告」}] - Want details for PASSING tests too → not supported here; the HTML reporter renders those via a different path Edge cases: - test_id substring matches nothing → empty list, no error - screenshot/trace/video missing on disk → those fields are null but the entry stays
回傳目前由 QA_RUNNER 環境變數選定的測試 runner(pytest / jest / cypress / go / maestro 五選一)加上 server 編譯時內建的全部 runner 清單。建議每個 session 第一個呼叫——AI 用它判斷後續該產 Playwright .py 還是 Maestro .yaml、要不要 headed browser,避免後面拿錯模板。也用來確認專案環境設定正確:QA_PROJECT_ROOT 指對地方、QA_RUNNER 沒拼錯。回傳 shape:{active: 'pytest', available: ['pytest', 'jest', ...]}。
讀上一次 run_tests 留下的 report.json,回傳一個輕量摘要:total / passed / failed / skipped / flaky_in_run(auto-retry 救回的數量)/ duration(秒)。比再跑一次 suite 便宜得多——適合在連續操作中間反覆查狀態。未跑過時回 {error: 找不到報告,請先執行 run_tests}。拿到摘要後若 failed > 0,接 get_failure_details 拿錯誤細節。 v1.3.0+: Edge AI runner attaches an optional `edge_metrics` block to each test entry ({p95_latency_ms, fps, iou_per_frame, labels_covered}). `get_optimization_plan` reads these to surface 4 Edge-specific flake signals (latency_p95_exceeded_sla, fps_variance_across_runs, iou_jitter_per_tc, coverage_gap_per_label) alongside the standard flake/broken/slow_regression categories. Non-edge runs have no `edge_metrics` field and see no signal changes.
Input schemas present but lack type definitions for most parameters. Rubric states: 'If parameters lack type definitions: schema score CANNOT exceed 30.' The run_tests tool accepts 'filter', 'headed', 'browser', 'plan_id' parameters, but the visible schema in server.py does not show explicit 'type' fields for each parameter (e.g., is 'plan_id' a string? is 'headed' boolean?). This forces LLMs to guess parameter types.
Output schemas documented only in prose descriptions, not formally in tool registration. The run_tests description promises returns like {exit_code, raw_exit_code, stdout_tail, stderr_tail, retry_enabled, flaky_in_run, plan_verification}; get_failure_details promises {nodeid, title, message, duration, steps[], screenshot, trace, video}. However, no formal outputSchema field is visible in the Tool definition. LLMs cannot plan downstream calls without knowing response structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 55 | 2026-07-28+ | v2 |
用 runner 的原生 collection 機制列出受測專案內所有可執行測試:pytest 走 `pytest --collect-only`、Jest 走 `npx jest --listTests`、Cypress 走 `cypress/e2e/*.cy.*` glob、Go 走 `go test -list .*`、Maestro 走 `*.yaml` 遞迴掃。回傳一份逐行 nodeid / 檔名清單。用法:run_tests 前確認 collection 沒漏、generate_test 前避免跟既有 case 重複。
只重跑上次失敗的測試——比跑整套套件快很多,適合修完一個 bug 後驗證迭代。pytest 走 `--lf`(last-failed)、Jest 走 `--onlyFailures`、Cypress 解析上次 report.json 的 failures[] 反查 spec 重跑、Go 撈失敗的 Test 名組成 regex 餵 -run、Maestro 反查 nodeid 對應 .yaml 重跑。需要先有過一次 run_tests(不然 report.json 不存在)。回傳 shape 跟 run_tests 一樣,接 get_test_report / get_failure_details 同樣方式檢視。
Execute the test suite under the active QA_RUNNER and produce a structured report. The single most-called tool — invoke whenever a user says 「跑/run/test/check/驗證/執行」, after generate_test (verify new test), or after a fix (confirm bug gone). Behavior: - Invokes the runner's native CLI under QA_PROJECT_ROOT — pytest with --screenshot=on / --tracing=on / --video=retain-on-failure, or `npx jest --json`, `npx cypress run --reporter json`, `go test -json`, `maestro test --format junit` - Optional `filter` narrows the scope: pytest -k expr, jest -t pattern, cypress --spec glob, go -run regex, maestro flow-name substring - Writes report.json (pytest-json-report shape, runner-agnostic) + JUnit XML - Snapshots the run into history/ and auto-triggers optimizer.write_plan() → optimization-plan.md is refreshed - Maestro: auto-retries flows that failed on first attempt (MAESTRO_RETRY=true), surfaces flaky_in_run count Returns: {exit_code, raw_exit_code, stdout_tail, stderr_tail, retry_enabled, flaky_in_run, ...} When to use: - After writing a new test → verify it actually passes - Smoke before a release - Whenever the user prompt contains a run/test verb When NOT to use: - Inspecting last results without re-running → use get_test_report (cheaper) - Re-running only failed cases → use run_failed (way faster) - Enumerating which tests exist → use list_tests Edge cases: - No tests match `filter` → exit_code != 0 with 「no tests ran」 in stderr_tail - QA_TIMEOUT_SECONDS exceeded → exit_code 124 + `[TIMEOUT…]` tag in stderr_tail - `filter` starting with `-` or containing `..` → blocked by security guardrail, returns {error: …} Plan bookend (v0.10.0): pass `plan_id` from a prior qa_plan call and the response auto-attaches `plan_verification` — the critical points are checked against the just-written report.json via the same flow run_api_security_scan uses. Omit `plan_id` to keep the legacy shape (no plan_verification key). When verify_plan fails (unknown / expired plan_id), the run still succeeds; the error envelope is surfaced *under* `plan_verification`.
Error handling descriptions are informative but lack categorization (retryable vs. user-fixable vs. fatal). The run_tests tool documents edge cases ('No tests match filter → exit_code != 0 with 「no tests ran」', 'QA_TIMEOUT_SECONDS exceeded → exit_code 124'), but does not explicitly tell the LLM which errors should trigger a retry, which should prompt user correction, and which are unrecoverable. Rubric: 'Categorize errors as retryable, user-fixable, or fatal.'
Tool annotations (readOnlyHint, destructiveHint) mentioned in rubric as current pattern are not visible in the source. The Risk fields in the server definition ('Risk: READ_ONLY' for get_runner_info, list_tests, etc.; 'Risk: WRITE' for run_tests, run_failed) suggest awareness of side effects, but these are not declared in the Tool object's inputSchema or as explicit MCP annotations. This is a gap between intended semantics and protocol declaration.
Descriptions are dense, bilingual (English + zh-TW), and often mix usage guidance with implementation details. Example: 'pytest 走 -k 表達式(支援 and/or/not)' mixes the English action ('use -k expression') with Chinese and technical jargon. The Rubric recommends 10 - 1024 chars; these descriptions often exceed 500 chars. While comprehensive, they risk overwhelming LLMs and diluting the core intent.
Parameter 'plan_id' in run_tests is optional but its semantics depend on context (v0.10.0+ feature). Description says 'Omit plan_id to keep the legacy shape (no plan_verification key)', this is a dependency that should be documented as a note but is buried in prose. Rubric: 'When one parameter's valid values depend on another... document this in both parameter descriptions.'