The server provides 15 well-documented tools with mostly complete schemas and descriptions. Tool naming follows verb_noun patterns consistently (e.g., wopee_create_blank_suite, wopee_fetch_analysis_suites). Most descriptions are substantive (100-300+ chars), explaining what, when, and why. However, several issues limit the score: (1) Tool descriptions sometimes redundantly state what other tools do (e.g., wopee_fetch_artifact and wopee_update_artifact have identical descriptions despite being different operations); (2) Some parameter enums lack descriptions (e.g., cookieConsent values ACCEPT_ALL/DECLINE_ALL/IGNORE are not explained); (3) No explicit error handling guidance documented in descriptions; (4) Output schemas are not formally documented in the tool definitions; (5) wopee_dispatch_agent and wopee_dispatch_analysis lack clarity on rate limiting and retry behavior despite mentioning it in description. The tool set is functionally cohesive and covers a reasonable domain (test suite management, artifact generation, test execution, chat integration, GitHub integration), but definition polish is below A-grade standards.
Create a new empty analysis suite in the current project. Use this as the first step when you want to manually build a test suite — the returned suite UUID is needed by wopee_generate_artifact, wopee_fetch_artifact, wopee_update_artifact, and wopee_dispatch_agent. If you want to auto-analyze a web app instead, use wopee_dispatch_analysis which creates and populates a suite in one step. Takes no input parameters; uses WOPEE_PROJECT_UUID from environment. Not idempotent: each call creates a new suite. Returns the suite object with UUID, name, type, and status. Fails if WOPEE_PROJECT_UUID is not configured.
Create a new GitHub issue in the project's connected repository. Use this to report bugs found during testing, track test failures, or create action items from chat discussions. The issue will be created in the GitHub repository linked to the current project. Requires the project to have GitHub integration configured and WOPEE_PROJECT_UUID to be set.
Dispatch an autonomous AI agent to execute specific test cases. The agent opens a real browser, navigates the app, follows test steps, and reports results. Tests run ASYNCHRONOUSLY (1-3 minutes). This tool returns tracking info confirming dispatch — NOT final results. Do NOT interpret the response as pass/fail. Results arrive later via chat notifications. Prerequisite: test cases must exist in the suite (generate with wopee_generate_artifact type USER_STORIES_WITH_TEST_CASES). Use wopee_fetch_recent_executions or wopee_fetch_executed_test_cases to check status later.
Create a new analysis suite AND dispatch an AI crawling agent in one step. The agent opens a real browser, navigates from the starting URL, discovers pages, and maps the application structure. Use this when you want to auto-analyze a web app — it combines suite creation and crawling. Use wopee_create_blank_suite instead if you want to manually populate the suite. Optionally accepts starting URL, login credentials, cookie preferences (ACCEPT_ALL, DECLINE_ALL, IGNORE), custom variables, and free-text instructions to guide the crawl. Not idempotent: each call creates a new suite and starts a new crawl. Side effects: creates a suite and execution records on the platform. Rate limit: 10 seconds between dispatches per project; concurrent calls auto-retry with exponential backoff. Returns the created suite object on success.
wopee_update_artifact has a copy-pasted description from wopee_fetch_artifact. The description reads 'Retrieve a specific test artifact...' but the tool is meant to UPDATE. This creates ambiguity about the tool's actual function and violates the pattern that descriptions must explain WHAT the tool does.
Parameter enums lack descriptive text. The cookieConsent enum (ACCEPT_ALL, DECLINE_ALL, IGNORE) in wopee_dispatch_analysis is not explained. LLMs need to understand what each option does (e.g., 'ACCEPT_ALL: Accept all cookies automatically; DECLINE_ALL: Reject non-essential cookies; IGNORE: Do not interact with cookie dialogs'). Currently, the LLM must guess the intent.
No documented error handling or recovery guidance. Tools like wopee_dispatch_agent and wopee_dispatch_analysis mention rate limiting and retries in prose but do not explain what error the LLM should expect, whether it is retryable, or what to do next. This violates the pattern that error responses must tell the LLM what to do next.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 57 | 2026-07-28+ | v2 |
List all analysis suites in the current project. Returns an array of suite objects with UUIDs, names, types (ANALYSIS, AGENT, etc.), and running statuses (IDLE, IN_PROGRESS, FINISHED). Use this to discover existing suites before calling other tools — you need a suite UUID for wopee_generate_artifact, wopee_fetch_artifact, wopee_update_artifact, and wopee_dispatch_agent. Read-only: does not create or modify anything. Takes no input; uses WOPEE_PROJECT_UUID from environment. Returns an empty array if no suites exist. Fails if WOPEE_PROJECT_UUID is not configured.
Retrieve a specific test artifact from a suite. Returns the artifact content as text. Use this to review what wopee_generate_artifact created, or to retrieve existing artifacts before editing with wopee_update_artifact. Does NOT modify any data — this is a read-only operation. If the requested artifact type has not been generated yet for this suite, returns an empty result. For PLAYWRIGHT_CODE, you must provide the test case identifier (e.g. 'US004:TC006'); omitting it returns an error. For all other types, the identifier parameter is ignored.
Retrieve results of test cases executed by the autonomous agent. Returns each test case with its execution status (IN_PROGRESS, FINISHED, FAILED), agent report (natural language findings), and code report (technical details). Read-only: does not trigger any execution. Use this after wopee_dispatch_agent to check results — if status is IN_PROGRESS, wait and call again. Requires suite UUID. Optionally accepts an analysis identifier (e.g. A068, found in suite data) to filter to a specific analysis run. Returns an empty array if no test cases have been executed in this suite. Do NOT use this to fetch test artifacts like user stories or code — use wopee_fetch_artifact for that.
Fetch the most recent test case executions for the current project (up to 20, newest first). Use this to check the status of recently dispatched tests without needing to remember specific suite UUIDs. Returns each run's verdict (PASSED, FAILED, or INCOMPLETE — the run never established a result, e.g. an infrastructure error, and says nothing about the application) or, for a run with no verdict yet, its execution status (IN_QUEUE, IN_PROGRESS, FINISHED, FAILED, STOPPED), plus agent reports. Takes no input; uses WOPEE_PROJECT_UUID from environment. Prefer this tool when the user asks 'what's the status?' or 'how did the tests go?' and you don't have the specific suite UUID handy.
The authoritative tool for how many tests exist and their latest status. Returns, per analysis, the FULL list of authored test cases joined with their latest execution status — including never-run ones as NOT_RUN. Use this for questions like 'how many tests do I have', 'list the scenarios/test cases in A001', or 'show executed and not-run tests in one table'. Terminology: a 'scenario' is a test case; test cases are grouped under user stories (US001) and identified as US001:TC001. Reusable blocks (user story R001) are counted separately (reusableBlockCount) and are building blocks, not runnable, so they never carry an execution status. Regular tests are all non-R001 test cases. Read-only. Takes an optional analysisIdentifier (e.g. A001) to scope to one analysis; omit to cover every analysis in the project. Prefer this over wopee_fetch_recent_executions / wopee_fetch_executed_test_cases when the user asks about totals or the complete list — those return only test cases that have already run.
Read the run-time variables (additionalVariables) that drive analysis/agent runs, at either level. level: PROJECT returns the project-level variables (uses WOPEE_PROJECT_UUID from the environment); level: ANALYSIS returns a specific analysis suite's variables and requires suiteUuid. Read-only. Returns a JSON string array of { key, value, sourceType } entries, or [] when none are set. Use wopee_fetch_analysis_suites to discover suite UUIDs.
Generate AI-powered test artifacts for a suite using the Wopee.io AI engine. Each call creates one artifact type — call multiple times for different types. Generation order matters: APP_CONTEXT must be generated before user stories, and user stories before test cases. If called out of order, the AI may produce lower quality results. On success, returns confirmation that generation started. Use wopee_fetch_artifact to retrieve the generated content once ready. Do NOT use this to update existing artifacts — use wopee_update_artifact instead. Generating the same type again overwrites the previous version.
Read recent messages from the current project's chat room. Returns the last N messages in chronological order, including sender info and timestamps. Use this to understand the conversation context or review what has been discussed. Requires WOPEE_PROJECT_UUID to be configured.
Send a message to the project's chat room. Use this to notify users, share findings, or log observations. The message will appear in the chat history visible to all project members.
Retrieve a specific test artifact from a suite. Returns the artifact content as text. Use this to review what wopee_generate_artifact created, or to retrieve existing artifacts before editing with wopee_update_artifact. Does NOT modify any data — this is a read-only operation. If the requested artifact type has not been generated yet for this suite, returns an empty result. For PLAYWRIGHT_CODE, you must provide the test case identifier (e.g. 'US004:TC006'); omitting it returns an error. For all other types, the identifier parameter is ignored.
Update the run-time variables (additionalVariables) that drive analysis/agent runs. Can update at either PROJECT or ANALYSIS level. Accepts a JSON object of key-value pairs to set or merge with existing variables.
Output schemas are not formally documented in tool definitions. While input schemas are declared with types and descriptions, the response structure for each tool (e.g., what fields are returned by wopee_fetch_analysis_suites, wopee_fetch_test_inventory) is not formally specified. LLMs have to infer structure from the description alone, risking misuse of response fields.
Conflicting descriptions for discovery vs. action. wopee_fetch_artifact and wopee_update_artifact have nearly identical descriptions, making it unclear which tool the LLM should call when. The description for wopee_update_artifact is verbatim the fetch_artifact description, suggesting a documentation error.
Missing parameter descriptions in some tools. wopee_update_artifact's 'type' parameter lacks a description explaining what types are valid (APP_CONTEXT, USER_STORIES, USER_STORIES_WITH_TEST_CASES, PLAYWRIGHT_CODE). The enum is implied from wopee_generate_artifact but not explicitly stated here.
Tool descriptions are overly verbose with mixed concerns. Some descriptions (e.g., wopee_dispatch_agent, wopee_fetch_test_inventory) intermix explanation of the tool's purpose with recommendations about when NOT to use other tools. While helpful, this violates the principle of single responsibility, the tool should describe itself, not other tools' domains.