Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
The server provides 9 tools with basic schemas and descriptions. All tools are explicitly registered with @mcp.tool() decorator in main.py and have input schemas visible. However, critical gaps exist: (1) tool naming is inconsistent and unclear, 'delete_expense_by_id_catogery' contains a typo ('catogery' vs 'category') and is functionally redundant with 'delete_expense_by_id'; (2) descriptions are minimal (10-20 chars) and lack context about when to use each tool or what data they return; (3) output schemas are not documented, callers cannot predict response structure; (4) error handling returns status/message dicts inconsistently; (5) no parameter descriptions for optional fields; (6) no guidance on destructive operations (delete_all_expenses has no confirmation or dry-run); (7) destructive tools lack permission checks or audit trails. The server lands in the 'Fair' range (C grade) due to moderate implementation but significant pattern gaps.
Tool naming collision and typo: 'delete_expense_by_id_catogery' is functionally identical to 'delete_expense_by_id' but includes a spelling error ('catogery' instead of 'category'). This violates the principle that each tool should do exactly one thing and that similar tools must have unambiguous names. LLMs will struggle to choose between these two delete operations.
Missing output schema documentation. Tools return dicts like {'status': 'success', 'id': ..., 'message': ...} or [{'id': ..., 'date': ..., 'amount': ...}] but the tool definitions do not document these schemas. LLMs cannot predict return types or plan downstream operations. All tools lack documented return schemas.
CRITICAL: Remove 'delete_expense_by_id_catogery' entirely. It is functionally identical to 'delete_expense_by_id' but with a typo in the parameter name. Keep 'delete_expense_by_id' as the canonical delete-by-ID tool. This resolves the tool naming collision.
Document output schemas for all tools. Example for get_all_expenses: 'Returns a list of expense objects, each with fields: id (integer), date (string, YYYY-MM-DD), amount (number), category (string), subcategory (string, may be empty), note (string, may be empty). Returns empty list if no expenses found.' For add_expense: 'Returns {"status": "success", "id": <new_expense_id>, "message": "..."} on success, or {"status": "error", "message": "..."} on failure.'
Expand tool descriptions to 50 - 200 characters. Include WHAT the tool does, WHEN to use it vs alternatives, and WHAT FIELDS it returns. Example: 'Retrieve all expenses from the database, sorted by date (earliest first). Returns a list of expense objects with id, date, amount, category, subcategory, and note fields. Use this to see the full expense history; use list_expenses_by_date to filter by date range, or summarize to aggregate by category.'
Add descriptions to optional parameters. Example for add_expense's 'subcategory': 'Optional subcategory for the expense (e.g., "Groceries" under "Food"). If not provided, defaults to empty string.' For 'note': 'Optional text description or memo (e.g., "lunch with client"). Defaults to empty string.'
Add input validation to date parameters. Validate that date strings match YYYY-MM-DD format. Return clear error: 'Invalid date: got "2025-13-45", must be YYYY-MM-DD format (e.g., 2025-01-15).' This lets LLMs self-correct.
Score history
Overall score trend
↓ 1 points across a rubric change (v1 → v2)
48/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
48
2026-07-28+
v2
2026-03-09
F
49
-
v1
58/100
Summarize expenses by category within a date range.
Minimal tool descriptions (10 - 30 characters). Example: 'Retrieve all expenses from the database.' (37 chars) lacks context about when to call it vs. list_expenses_by_date, what structure is returned, or whether it paginates. Descriptions should include WHAT, WHEN, WHAT DOES IT RETURN.
No parameter descriptions. Optional fields like 'subcategory' and 'note' in add_expense lack descriptions. LLMs cannot infer whether a param is required, what format it expects, or what it controls without explicit documentation.
Destructive operations lack confirmation or dry-run. delete_all_expenses (which deletes every record) has no confirmation step, no dry-run mode, no permission checks.
No error guidance or recovery hints. Errors return dicts like {'status': 'error', 'message': '...'} but do not guide the LLM on what to do next (retry? ask user? call a different tool?).
No input validation or constraint documentation. Example: add_expense accepts 'date' as a string in YYYY-MM-DD format, but the tool does not validate this format or describe the expected range. LLMs can pass invalid dates like '2025-13-45'.
No pagination support. Tools like get_all_expenses and list_expenses_by_date return all records without limit or pagination. Without pagination, large result sets blow the context window.
No audit trail or permission checks. Destructive tools (delete_*) do not log who called them, when, or with which parameters. No permission gates exist, any agent can delete all expenses.
Implement confirmation for delete_all_expenses. Add a new tool 'confirm_delete_all_expenses(confirmation_token)' that requires a unique token returned by delete_all_expenses in dry-run mode, or add a 'dry_run' boolean parameter that returns the count of expenses to be deleted without actually deleting them. This prevents catastrophic accidents.
Add permission checks. Before executing any delete operation, verify the calling agent/user has authority (e.g., check an allow-list or permission scope). Return 'Permission denied: you do not have authority to delete expenses.' This enables least-privilege agent configurations.
Add pagination to get_all_expenses and list_expenses_by_date. Add optional parameters 'limit' (default 20, max 100) and 'offset' (default 0). Return a response with 'expenses' (list), 'total' (count of all matching records), and 'offset' and 'limit' for the next call. This prevents context window bloat.
Add error guidance and recovery hints. Instead of 'Database error: ...' return actionable messages like: 'Database is in read-only mode. Check file permissions on /path/to/expenses.db. If permissions are correct, try again.' For not-found errors: 'No expense found with ID 123. Use get_all_expenses or list_expenses_by_date to find valid IDs.'
Add an audit log. Log all tool calls to stderr or a structured log file with timestamp, tool name, parameters (redacted if sensitive), and outcome. This enables compliance and incident response.
Document constraints for amount parameter. Example: 'Amount of the expense (positive number, e.g., 25.50). Limit to 6 digits before the decimal point and 2 after.' This prevents LLMs from passing absurd values like 1e10.