AI-powered E2E integration tester via MCP — analyzes microservices, tests APIs, crawls frontends with Playwright
SecSee has 7 tools with reasonable names (verb_noun pattern) and all tools have descriptions. However, there are critical gaps in parameter descriptions, schema completeness, and error handling guidance. The server follows a logical analysis→test workflow but lacks the polish required for production use. Parameter descriptions are often minimal (e.g., 'Absolute path to the project root directory' for projectRoot across all tools), and output schemas are not formally documented. The 'analysis' parameter in update_working_doc uses z.any() instead of properly typed nested schemas, violating the schema documentation requirement. Error handling provides no recovery guidance. Tool compositions rely on an implicit working.json format that users must understand but is never formally specified in the tool outputs.
Scan the project to discover services and collect source files. Returns the full project structure with file contents so the LLM can analyze endpoints, DB interactions, auth schemes, etc. After reviewing the output, call update_working_doc with the structured analysis.
Run docker compose up -d --wait to start all services
Run docker compose down to stop and remove all containers
Run all API integration tests based on working.json analysis
Launch Playwright and crawl the frontend, testing all pages and forms
Test a specific API endpoint by path and method
update_working_doc uses z.any() for critical 'analysis' parameter structure instead of proper typed schema. The nested object properties (services, endpoints, dbInteractions) are documented only in comments within z.any().describe() calls. This violates schema documentation requirements and prevents LLMs from understanding what structure to provide.
Parameter descriptions across all tools are minimal and provide no guidance on constraints, expected formats, or dependencies. E.g., 'Absolute path to the project root directory' repeated on projectRoot param in 6 tools, but no guidance on whether relative paths work, symlinks are followed, or permissions required.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
Accept structured project analysis from the LLM and write working.md + working.json.
No output schemas documented for any tool. The code shows return objects (text content wrapped in { type: 'text', text: '...' }) but tool definitions do not specify what structure LLMs should expect. The working.json format returned by analyze_project and implied by update_working_doc is never formally specified in tool outputs.
Error handling provides no recovery guidance. spin_up_services and tear_down_services return raw stdout/stderr on failure with no guidance on root causes or next steps. 'Failed to start services' followed by docker output gives the LLM no hint whether to retry, check configuration, or ask user.
tear_down_services has destructive capability (removeVolumes parameter) but no confirmation/dry-run pattern. An LLM could call tear_down_services with removeVolumes=true and lose all database volumes without warning. No destructiveHint tool annotation present.
Tool composition assumes implicit dependencies and state. analyze_project output tells user to 'call update_working_doc' but doesn't explain why or what happens if skipped. test_apis requires working.json to exist (readWorkingJson called inside) but no tool documents this dependency or provides recovery if file missing.
test_single_endpoint method parameter uses free-form string (no enum constraint) instead of limiting to HTTP verbs. Description says '(GET, POST, PUT, DELETE, PATCH)' in text but no enum validation. LLM could pass 'RETRIEVE' or 'FETCH' and tool would fail.
analyze_project returns a markdown summary with embedded source code (potentially megabytes) instead of structured JSON. While helpful for human review, this wastes tokens and makes it hard for LLMs to extract actionable data programmatically. No pagination or limit on file count.