Basin MCP defines a single tool 'check-generated-code' with a verbose, prescriptive description (2,500+ chars) that reads more like implementation documentation than an LLM-optimized tool guide. The input schema is technically present but extremely complex (15+ parameters, many undescribed or poorly described), lacks clear parameter relationships, and violates single-responsibility design. The schema includes nested objects with inconsistent structure (formSubmissions.forms vs formSubmissions, elements.endpoints vs elements). No documented output schema. The tool name itself is weak, 'check-generated-code' is passive and ambiguous; 'verify-code-quality' or 'test-code' would be clearer. Parameter descriptions are verbose repetitions of the main tool description rather than tight, parameter-specific guidance. Critical parameters like 'credentials' expose sensitive data as a parameter object (anti-pattern per pattern:secret-injection). No error handling guidance, the tool simply documents what Basin depends on (a running web service) without explaining recovery paths.
Code Quality / Reliability Check Tool This tool must be invoked after every code generation step. This tool ensures code quality and reliability by testing the code for detecting bugs, hallucinations, and unintended changes. Again, it's vital that this tool is called every time calling agent generates a code. When to use: Always run this tool after each code generation to check the generated code to detect errors or unintended changes. This tool returns errors in the current code and some guides on how to fix them, use these guides to rewrite the code to fix the errors this tool discovers. How to use: Provide: 1) Original user prompt, 2) test type, enter 'general' for user-specified thorough tests, or 'sanity' for quick automated sanity checks that runs at the end of each code generation event (choose 'sanity' for all other cases), 3) List of affected files, 4) Generated code, 5) thread id / context id for this conversation, used to keep track of before and after changes, 6) URL of the running local server, 7) URLs of the new and changed endpoints, 8) credentials to use if testing requires authentication, 9) List of features added, changed, or removed per endpoint, 10) List of features that have submissions, AND name and/or id of the submit button element / field for those features (must include the element id or name), 11) Id and names of each of the fields, elements, and buttons per each endpoint, both modified and unmodified, 12) list of external API calls and 3rd party services the feature uses, 13) Structure of files in the project, 14) Lengthy and very well detailed description of related features and endpoints. It's imperative that you provide the list of fields, elements, and buttons for each endpoint, both modified and unmodified, if they are available. Fields, elements, and buttons has to be uniquely identifiable by id and name. for 5), thread id, it should be unique id of the LLM chat thread, if that is not available, leave it empty. Also, make sure to specify if features trigger a navigation to another endpoint or if they don't navigate to any new endpoint, and provide the new endpoint url for each feature that navigates to another endpoint if known. Things to note: - Basin depends on a running web service to test for validity of the features, run a server, or ask the user to run one prior to using this tool, and provide the URL to that server to this tool.
Tool name is passive and vague. 'check-generated-code' does not start with a clear action verb and conflates validation, testing, and reporting. LLMs cannot infer whether this tool is for linting, execution, security analysis, or something else. Should use verb_noun format: 'verify_code_quality', 'test_generated_code', or 'validate_code_changes'.
Tool description is 2,500+ characters, far exceeding the 1,024-char baseline upper limit for tool descriptions. It reads as implementation documentation (prescriptive step-by-step 'Provide: 1) ... 2) ... 14)...') rather than LLM-optimized guidance. LLMs waste tokens parsing this verbose text and struggle to extract the core action and preconditions. Compress to: 'Tests generated code for bugs, hallucinations, and unintended changes. Requires: original prompt, test type (sanity/general), affected files, generated code, running server URL, endpoint URLs, features affected, form submission details, element IDs, and external API usage. Returns: detected errors and remediation guidance. Prerequisites: Basin server must be running.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 33 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 31 | - | v1 |
Credentials parameter exposes sensitive data as a tool input object (properties: username, password). Per pattern:secret-injection, credentials must NEVER be tool parameters, they enter agent logs, prompt history, and traces. Move credentials to server-side secret injection via environment variables or a vault. If authentication is needed, accept a 'credentials_label' or 'secret_name' that the server resolves securely, not the actual username/password.
Input schema is malformed and inconsistent. 'formSubmissions' declares 'forms' as a schema object but the property is not 'forms', it should be 'items'. 'elements' declares 'endpoints' as a schema object but should declare 'items'. The schema structure violates JSON Schema conventions and is unparseable by strict validators. Each complex parameter (formSubmissions, elements) needs proper array-of-objects structure with clear property names.
15 input parameters with inconsistent descriptions. Many parameters ('happyPathEndpoints', 'sadPathEndpoints', 'externalApiCalls', 'fileStructure', 'formSubmissions', 'elements') lack clear descriptions of expected structure and content. 'fileStructure' is simply described as 'Structure of files in the project', does it expect a nested object? A flat list? A JSON tree? A file manifest string? LLMs cannot infer. Each parameter description should explain: (a) data type/format, (b) how to populate it, (c) why it matters. Baseline: average params per tool is 4; this tool has 15, consider splitting into multiple focused tools or leveraging a configuration file.
No documented output schema. LLMs need to know what fields to expect from this tool so they can parse errors, extract remediation guidance, and plan next steps. Document: returns { errors: [ { code, message, remediation } ], passed: boolean, summary: string, ... }. Without this, agents cannot reason about downstream actions.
No error handling guidance. Tool description says 'Basin depends on a running web service' but does not explain what the LLM should do if the service is down, unreachable, or returns a timeout. Per pattern:recovery-guide, error responses must tell the agent what to do next: 'Service unreachable. Check that the Basin server is running at [URL]. If not, ask the user to start it and retry.'
No distinction between retryable and fatal errors. When serverUrl is unreachable or invalid, is the error retryable? When generated code has bugs, should the agent retry or ask the user to fix the code? LLMs need error categories (retryable, user-fixable, fatal) to know whether to retry, escalate, or ask for clarification.
Tool violates single-responsibility principle (pattern:tool). It accepts 15 parameters covering code quality, testing, form validation, element inspection, and endpoint navigation. This combines linting, testing, API validation, and UI inspection into one mega-tool. Consider splitting: (1) verify_code_syntax (input: code, language) → errors; (2) test_endpoints (input: serverUrl, endpoints, testType) → results; (3) validate_forms (input: formSubmissions, elements) → validation_results. Each focused tool is easier for LLMs to reason about and compose.
Parameter 'testType' is an enum (sanity, general) but the description does not explain when to use each. Add: 'sanity: quick automated checks after each generation (~1-5 sec). general: thorough, user-specified tests including integration scenarios (may take 10-60 sec). Default to sanity for speed unless the user explicitly requests thorough testing.'