NO OUTPUT SCHEMAS DOCUMENTED: Not a single tool specifies what it returns. LLMs cannot infer response structure, preventing downstream tool chaining and forcing agents to blind-guess field names. Critical for composition patterns.
Add output schemas to ALL tools. Specify what fields are returned, their types, and their meaning. Example for get_statistics: returns {document_count: int, chunk_count: int, embedding_coverage: float, total_embeddings: int}. Critical for tool chaining.
Rewrite ALL tool descriptions to 50 - 200 characters, following pattern:tool-description. Template: '[WHAT it does] + [WHEN to call it instead of similar tools] + [key output]'. Example: 'Batch insert 1-1000 documents into the database. Call this instead of insert_document when adding multiple documents for better performance. Returns {inserted_count, failed_count, errors}.'
Add enum constraints for parameters with restricted values. Example: train's 'algorithm' parameter should be {type: string, enum: [random_forest, linear_regression, gradient_boosting, ...], description: 'ML algorithm to train...'}
Document all parameter formats and constraints in descriptions. Example for 'chunk_size': 'Chunk size in characters (1 - 100000, default 1000). Larger chunks reduce API calls but may lose semantic boundaries.'
Remove raw SQL exposure. Replace execute/execute_one with domain-specific tools like search_documents(query, limit), get_embedding_stats(), or execute_sql_read_only(pre-validated_query_name). Sanitize all LLM-provided input.
Replace 'conn' parameter (database connection) with server-side secret injection. Store DB credentials in environment or vault; never expose to tools or agents.
Add mutual exclusivity documentation. Example for insert_document: 'Provide either (filepath) OR (filename, title, content), not both. If filepath is set, filename/title/content are ignored.'
Score history
Overall score trend
↓ 9 points across a rubric change (v1 → v2)
27/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-21
F
27
2026-07-28+
v2
2026-03-09
F
36
-
v1
GENERIC, UNHELPFUL DESCRIPTIONS: 23 of 26 tools have descriptions under 50 characters or are single-phrase non-actionable (e.g., 'Get overall statistics', 'Execute SQL query and return results'). No guidance on WHEN to call the tool, WHY it exists, or what problem it solves. Violates pattern:tool-description critical rule that descriptions must enable LLM selection.
RAW SQL EXPOSURE + CREDENTIAL RISKS: Tools 'execute', 'execute_one', 'aexecute' accept raw SQL query strings as parameters. This creates SQL injection vulnerability if LLM-provided input is not sanitized. Additionally, create_schema() and other db_operations expose 'conn' (database connection object) as a parameter, credentials must never be tool parameters, only server-side injected. Violates pattern:secret-injection and pattern:tool-gateway.
NO MUTUAL EXCLUSIVITY DOCUMENTATION: Parameters like insert_document's (filepath + filename + title) have unclear relationships. Can you provide all three or are they alternatives? No docs state this. train() tool has both 'features' and 'algorithm' but no guidance on how they interact. Violates review:param-relationships.
MISSING ERROR HANDLING GUIDANCE: No tool description explains error recovery (e.g., 'If chunk size is too large, try reducing it and retry'). No indication of which errors are retryable vs. fatal. Violates pattern:recovery-guide and pattern:error-classification.
AMBIGUOUS / VAGUE TOOL NAMES: 'execute' and 'execute_one' are generic verbs with no object. 'process_file' does not say what kind of processing (format conversion? parsing? embedding generation?). 'train' alone is vague, does it train a model, an agent, or calibrate parameters? Violates pattern:tool naming critical rule ('start with action verb that reflects the action').
NO PAGINATION / RESULT LIMITS DOCUMENTED: Tools like 'get_statistics', 'find_documentation_files', 'execute' (when retrieving rows) do not document max result count or pagination. Large result sets can exhaust context windows. Violates pattern:paginated-result.
MISSING IDEMPOTENCY DECLARATIONS: insert_document, batch_insert_documents, and train are WRITE operations that may have side effects on retry. No tool description states whether repeated calls with same input are idempotent (safe to retry) or create duplicates. Violates pattern:idempotent-operation.
MISSING PERMISSION / SCOPE DOCUMENTATION: No tool declares what permissions it requires (e.g., 'write:database', 'read:filesystem'). No audit trail or caller identity context visible. Violates pattern:scope-declaration and pattern:audit-trail.
Add error recovery guidance to tool descriptions. Example for create_index: 'Returns error if index already exists. Call index_exists first, or drop_index before retrying. Common failure: out of memory on large tables, reduce m or ef_construction parameters.'
Document pagination and result limits in tool descriptions. Example for execute: 'Returns up to 1000 rows per call. For larger result sets, use LIMIT/OFFSET in your SQL query or paginate programmatically.'
Add idempotency declarations. Example for insert_document: 'IDEMPOTENT: If called twice with same filepath, the second call updates the existing document rather than creating a duplicate.'
Add permission/scope declarations. Example: train tool requires 'write:database' + 'read:database' scope. Document in tool description so agent configurations can enforce least privilege.
Add input validation & actionable errors. Return structured errors like {code: INVALID_ALGORITHM, message: 'Algorithm must be one of: random_forest, linear_regression, gradient_boosting', valid_options: [...]}. Never return raw stack traces.