Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
ExpenseTracker has reasonable naming and complete schemas for all 12 tools, but descriptions are sparse and lack critical guidance for LLM selection. Parameter descriptions exist but are minimal (10-30 chars). Error handling returns status codes but lacks recovery guidance. No tool annotations (destructiveHint, idempotentHint, readOnlyHint) despite having a mix of READ_ONLY, REVERSIBLE, WRITE, and DESTRUCTIVE operations. The server lacks pagination for list results despite storing expenses in a growing database. Output schemas are implicit rather than documented.
Tools (12)
add_expensewritesource verified70/100
Add an expense to the tracker
budget_alertread onlysource verified68/100
Check if spending exceeds budget limit
daily_averageread onlysource verified67/100
Calculate the average daily spending within a date range
Missing tool annotations (destructiveHint, idempotentHint, readOnlyHint) despite risk field metadata in spec. Tools like delete_expense (DESTRUCTIVE) and update_expense (REVERSIBLE) should be explicitly annotated for LLM decision-making.
Tool descriptions are extremely terse (13-45 characters). E.g., 'Add an expense to the tracker' lacks context on date format, category validation, or when to use this vs bulk import.
Expand every tool description to 150 - 250 characters, answering WHAT (state modification), WHEN (vs similar tools), and WHAT IF FAILS (recovery guidance). E.g., 'Add an expense to the tracker' → 'Record a new expense with date (YYYY-MM-DD), amount, and category. Returns expense_id. Use summarize() to review totals by category. If date is invalid, returns error with valid format.'
Add tool annotations for every tool: readOnlyHint=true for list/summarize/report tools; destructiveHint=true for delete_expense; idempotentHint=true for read/list tools (safe to retry). Align with @mcp.tool() decorator or schema metadata.
Implement pagination for list_expenses, summarize, and top_spending_categories: add limit (1 - 100, default 20) and offset (default 0) parameters. Return {results: [...], total_count, has_more} so LLMs know when to fetch next page.
Fix monthly_report month-end logic: use calendar.monthrange(year, month) to compute actual last day of month instead of hardcoding 31. Document that function returns summary by category for the given month.
Refactor error responses: return {error_type: 'invalid_input'|'not_found'|'conflict'|'server_error', message: '...', recovery: '...'}. E.g., 'error_type: invalid_input, message: Invalid date format, recovery: Use YYYY-MM-DD (e.g., 2024-01-15)'.
Add dry-run parameter to delete_expense: dry_run=true returns what would be deleted without actually deleting. E.g., 'Found 5 expenses totaling $234. Call again with dry_run=false to confirm deletion.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
First recorded score · v2 rubric
64/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-23
C
64
<=2025-11-25
v2
Generate a monthly expense report
read_expenseread onlysource verified70/100
Read an expense by ID
summarizeread onlysource verified67/100
Summarize expenses by category within a date range
Get the top spending categories within a date range
update_expensereversiblesource verified70/100
Update an expense
list_expenses, summarize, and top_spending_categories have no pagination or result limits documented. As the expense database grows, returning all matching rows will blow context windows. Should accept limit/offset params and return total_count or next_cursor.
Error responses return generic status/error fields without recovery guidance. E.g., 'Invalid expense' tells LLM nothing about whether to retry, ask user, or skip. Should include classification (retryable/user-fixable/fatal) and actionable next step.
No dry-run or confirmation pattern for irreversible operations (delete_expense, export_csv). Agents should be able to preview consequences before committing destructive actions.
Parameter descriptions lack format constraints. E.g., 'Date in YYYY-MM-DD format' is mentioned in schema but 'start_date' and 'end_date' descriptions are vague. Should state: 'ISO 8601 date (YYYY-MM-DD, required)', 'Category must be one of: [list] or custom string', 'Limit must be 1-100 (default 3)'.
monthly_report has broken logic: constructs month-end dates with hardcoded '31' (e.g., '2024-02-31'), which will fail for February. Should use calendar module to compute actual month-end or document as accepting year/month only and computing correct range internally.
export_csv does not handle empty data set gracefully. Calls list_expenses which may return error or empty list; accessing data[0].keys() will crash if list_expenses returns empty. Should validate data before writing CSV and return actionable error.
Output schemas are implicit/undocumented. Tool responses return dicts like {status, id}, {status, error}, or raw result rows, but no formal schema is declared in tool metadata. LLMs cannot predict response structure and must infer from examples.
delete_expense provides no confirmation mechanism. Agents could accidentally delete all expenses in a loop. Should require explicit confirmation or return a preview of what will be deleted.
delete_expense
Document output schemas in tool metadata. Specify that add_expense returns {status: 'ok'|'error', id?: integer, error?: string}. List tools return [{id, date, amount, category, subcategory, note}, ...]. This enables LLM planning.
Implement rate limiting and timeouts for database operations. Add explicit timeout (e.g., 5 seconds) for list_expenses on large date ranges. Return 'Query timeout. Try narrowing date range or filtering by category.' if exceeded.
Add parameter constraints in schema: amount {minimum: 0.01, maximum: 999999.99}, limit {minimum: 1, maximum: 100, default: 20}, month {minimum: 1, maximum: 12}. These are machine-parseable and prevent LLM hallucination.