Two tools are defined with explicit Zod schemas and descriptions in src/initTools.ts. However, significant quality gaps prevent a higher score: (1) Both tools expose database credentials as optional parameters despite warnings in descriptions, a critical security anti-pattern; (2) Descriptions are lengthy (500+ chars) and mix usage guidance with restrictions, violating the concise 10-1024 char baseline; (3) Parameter descriptions in schemas are minimal (often one phrase); (4) No output schema is documented, responses are inferred to be text/markdown but not formally specified; (5) Error handling is minimal (single catch block returning generic error text); (6) Tool names are sufficiently specific but lack clear verb_noun convention ('query-database' vs. 'query_database'); (7) Turbo variant includes a large table of heuristics in the description, which is documentation-as-description anti-pattern. The Zod schema presence and explicit registration prevent this from falling below 40, but quality is below median.
# Data Analyst Agent: Reasoning Analysis Mode This tool translates natural language into multi-step SQL analysis plans and executes them against databases. Use this for complex analytical questions requiring more reasoning. ## Capabilities - Performs multi-step analyses with each step building on previous results - Analyzes data across multiple tables with complex relationships - Handles complex queries requiring careful reasoning and planning - Produces comprehensive markdown reports with insights ## When to Use - For complex analytical questions requiring deep reasoning - When accuracy and comprehensiveness is more important than speed - For queries involving multiple tables or complex relationships - When detailed insights and explanations are needed ## Effective Prompts - Be specific about metrics, time periods, and entities of interest - Include relevant business context for interpretation - Specify desired output format (tables, charts, insights) - For complex analyses, break down into logical components ## Restrictions: - Don't sent database credentials in the payload, it's handled by the server. - Don't sent API keys in the payload, it's handled by the server.
### **When to Use (Natural-Language Heuristics)** Because the model sees only the *user's question* and minimal schema hints, Turbo Mode should activate automatically **whenever the request exhibits every one of these surface-level cues**. Each cue corresponds to a first-principles driver of SQL complexity that the model *can* infer without deep schema knowledge: | Signal in the User's Question | Why It Indicates Turbo Is Safe | | ------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | | **Single Factual Verb** — verbs like "count," "list," "show," "sum," "average," or "max/min," used **once**. | One aggregate or projection keeps the SQL to a single `SELECT`. | | **At Most One Qualifier Clause** — a lone filter such as a date range, status, or simple equality ("where status = 'active'"). | Few filters avoid nested logic or subqueries. | | **No Comparative Language** — absent words like "versus," "compare," "trend," "change over time," "prior year," "by each," etc. | Comparisons imply multiple groupings, time windows, or self-joins. | | **No Multi-Dimensional Grouping Phrases** — avoids "by region and product," "per user per month," "split across categories." | Multiple dimensions require complex `GROUP BY` and often joins. | | **Mentions One Table-Like Concept** — either explicitly ("in `orders`") or implicitly ("orders today," "users last week"). | Referencing several entities hints at join logic the model can't verify quickly. | | **Requests Raw IDs or a Small Top-N List** — e.g., "give me the top 5 order IDs." | The result set will be tiny, so execution latency is dominated by query planning—not data transfer. | | **No Need for Explanation or Visualization** — the user asks only for the numbers or rows, not "explain why" or "graph this." | Generating narrative or charts costs tokens and time; Turbo avoids it. | > **Quick mental check**: *Could you answer this with a single short sentence and a single‐line SQL query template?* > If yes, Turbo Mode is appropriate. --- ### Limitations * Unsuitable for multi-step or exploratory workflows * May miss domain nuances captured in the standard reasoning path * Provides limited explanation and simplistic visuals --- ### Effective Prompts * "How many active users signed up last week?" * "List the five most expensive orders." * "Show the total revenue for March 2025." * "What is the average session length today?"
Credentials exposed as tool parameters. Both 'query-database' and 'query-database-turbo' accept database user, password, and connection strings as optional input parameters, despite server-side environment variable support. This violates pattern:secret-injection and creates audit/compliance risk, secrets in tool parameters are logged and could appear in agent traces.
No output schema documented. Tools return { content: [{ type: 'text', text: <markdown> }] } but this structure is implicit in code, not declared in a schema property. LLMs cannot infer what fields to expect in the response.
Descriptions exceed 1024 character baseline. query-database description is ~520 chars; query-database-turbo is ~1800+ chars (includes full decision table). This violates pattern:tool-description baseline (10-1024 chars) and wastes tokens in LLM context. Move detailed guidance to docs and keep descriptions concise.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 22 | - | v1 |
Parameter descriptions are minimal or missing context. E.g., 'prompt' param is described only as 'Natural language question or analysis request' with no guidance on specificity, format, or constraints. 'databaseConfig' is a complex nested object with only one-sentence descriptions per field.
Minimal error handling and recovery guidance. Single catch block returns generic 'Error: <message>' text. No categorization (retryable vs. fatal), no recovery hints. Pattern:recovery-guide requires 'User not found. Try search_users() with...' style actionable guidance.
Ambiguous parameter semantics. databaseConfig is optional; when omitted, the server uses environment variables. When provided, should it override or merge? The code shows conditional logic ('DONT_USE_DB_ENVS === "true"') but the parameter description does not explain this behavior. Undocumented dependencies per pattern:tool-description violate mxe:param-relationships.
Tool naming convention inconsistent with verb_noun baseline. 'query-database' uses hyphens; Zod schema and codebase use camelCase (databaseConfig). While functional, this violates the mxe:chat-data-model principle of matching user intent language.