AI-powered app test automation MCP server - scenario generation, execution, detection, and auto-fixing
test-genie-mcp has 21 tools with visible schemas and descriptions. Naming follows verb_noun convention consistently (analyze_, generate_, create_, run_, detect_, suggest_, apply_, etc.). However, several critical issues reduce overall quality: (1) Many descriptions lack specificity about prerequisites, return structures, or error conditions. (2) Parameter descriptions are present but often generic (e.g., 'Path to the project root' appears 20+ times without explaining why projectPath is required for each distinct tool). (3) No documented output schemas, tools return JSON but LLMs cannot plan what fields to expect. (4) Several tools combine multiple concerns (run_full_automation does analyze→scenarios→test→detect→fix→report in one call; run_iterative_fix_loop repeats test→detect→fix cycle). (5) STDIO transport is a hard ceiling at 50 per spec rules, but definition quality alone is mediocre-to-fair. Positive aspects: all tools have names, descriptions, and input schemas with types; enums are used appropriately for constrained inputs (platform, testType, coverage, etc.); parameter descriptions exist for all inputs. The server is functional but would benefit from output schema documentation, tighter tool composition, and richer error handling.
[mode: real] Static analysis of the project: screens, components, APIs, state. Auto-detects platform when not provided.
[mode: real] Profile and analyze app performance: CPU, memory, rendering, network.
[mode: real] AST-based analysis: extract components, dependencies, code patterns.
[mode: real] Detect potential race conditions and concurrency issues.
[mode: real] Security audit: hardcoded credentials, injection vulnerabilities, insecure patterns.
[mode: real] Apply a confirmed fix to the codebase.
[mode: real] Interactive confirmation of a fix suggestion.
No documented output schemas, tools return JSON but descriptions do not explain what fields consumers should expect. This forces LLMs to infer structure, increasing hallucination risk and preventing multi-tool chaining based on known field names.
run_full_automation and run_iterative_fix_loop combine multiple concerns in a single tool. run_full_automation does 6 sequential steps (analyze → scenarios → test → detect → fix → report); run_iterative_fix_loop repeats a 3-step cycle. These violate single-responsibility and prevent agents from composing workflows. Splitting into single-purpose tools (e.g., run_tests, detect_issues, apply_fixes) would enable more flexible agent reasoning.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
[mode: real] Build a test plan from stored scenarios with filtering / scheduling.
[mode: real] Detect race conditions, null refs, state inconsistencies.
[mode: real] Detect memory leaks, retain cycles, unclosed resources.
[mode: real] Comprehensive health check: test status, issues, performance, security, recommendations.
[mode: real] Generate CI/CD pipeline configuration (GitHub Actions, GitLab CI, etc.).
[mode: real] Aggregate test results, issues, fixes, and metrics into a report.
[mode: real] Generate test scenarios from analyzed app structure.
[mode: real] Rollback a previously applied fix.
[mode: real] End-to-end automation: analyze → scenarios → test → detect → fix → report.
[mode: real] Repeat: run test → detect issues → suggest/apply fixes → test again.
[mode: hybrid] Run a stored scenario; real subprocess where possible, falls back to simulated.
[mode: simulated] Random / sequential user-behavior simulation to find issues.
[mode: hybrid] Concurrency / load test against an endpoint or UI surface.
[mode: real] Generate rule-based fix suggestions for detected issues.
Descriptions lack actionable context. Most tools repeat generic phrasing ('Path to the project root', '[mode: real]') without explaining when to call this tool vs. related ones, what data it requires as input (e.g., must analyze_app_structure run first?), or how errors should be handled. Descriptions should answer: What does it do? When use it instead of similar tools? What fields does it return?
No error guidance in descriptions. Tools like apply_fix (IRREVERSIBLE) and run_full_automation lack error handling documentation. What should an LLM do if a fix fails to apply? If automation detects too many issues? Irreversible operations must document recovery paths.
Confirmation-only tool (confirm_fix) expects manual LLM reasoning to choose fixId, but no tool returns fix IDs that can be chained to this tool. The workflow (suggest_fixes → confirm_fix → apply_fix) requires the LLM to extract a fixId from suggest_fixes output, but the output schema is undocumented.
Platform parameter repeated across 12 tools with same enum but no guidance on auto-detection behavior. Description says 'Auto-detects platform when not provided' for analyze_app_structure, but unclear if other tools also auto-detect or require explicit platform. Should document detection precedence (package.json? file structure? other signals?).
Parameter descriptions use generic templating. Every tool with projectPath repeats 'Path to the project root' verbatim. The description should explain why the path is needed for THIS tool and what it must contain (e.g., 'Absolute or relative path to the app's root, must contain package.json or AndroidManifest.xml').
No output pagination for tools returning lists. generate_scenarios and generate_report likely produce large result sets, but descriptions do not mention pagination, limits, or how to cap results. Without pagination docs, LLMs may request all scenarios (unbounded) and blow context window.
Batch operations missing. Agents may need to run multiple scenarios, apply multiple fixes, or generate multiple plans. Separate tools for batching (run_scenario_tests, apply_fixes) would reduce token waste vs. calling single-item tools in a loop.
Risk markers in tool descriptions ([mode: real], [mode: hybrid], [mode: simulated]) are not standardized schema fields and may confuse LLMs that do not parse markdown annotations. Should use structured tool annotations (readOnlyHint, destructiveHint, idempotentHint) in a capabilities extension or response wrapper.