A college expense tracker and personal finance management system with social debt tracking, fitness logging, and spending analytics
Finance-Tracker has 10 tools with partial schema definitions and inconsistent documentation. Most tools have descriptions (7/10), but parameter documentation is sparse and inconsistent. No input schemas are formally visible in the source code, only inline type hints in Python function signatures. Output schemas are not documented. Error handling is minimal and does not guide recovery. The server relies on string parsing and database RPC calls without validation guidance. Tool naming is generally clear (verb_noun pattern), but composition suffers from overly smart tools that combine multiple concerns (e.g., log_personal_expense includes health clarification logic internally). Security: no evidence of secret injection, permission gates, or audit trails. Worst performers: log_protein_intake (instructs AI to estimate values without validation), check_social_finances (vague query_type parameter), log_debt (conflates user perspective with explicit borrower/lender).
Registers a new friend for tracking shared expenses/debts.
Analyzes personal spending by Category and Health.
Checks gym attendance and total protein for the last X days.
The Master Tool for social finances. Can answer history OR balance questions.
Teaches the system if a specific food item is healthy (True) or unhealthy (False). Use this when the user answers your clarification question.
Logs that one person owes another money. If user owes someone: borrower='Me', lender='Name'. If someone owes user: borrower='Name', lender='Me'
No formal JSON Schema registration visible. All 10 tools use Python type hints inline but do not show explicit MCP schema registration. This violates the 'schemas must be machine-readable' principle and makes it impossible for MCP clients to validate inputs before sending requests.
log_protein_intake instructs AI to estimate nutritional values without validation, confidence bounds, or server-side verification. This is a correctness anti-pattern, the tool trusts LLM reasoning for factual data (protein content in foods) and provides no way to validate or correct estimates. Users will receive hallucinated nutrition data.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Logs a personal expense. IMPORTANT: If you don't know if the item is healthy, pass is_healthy=None. The system will check the database and tell you if it needs clarification.
Logs food and its protein content. CRITICAL INSTRUCTION FOR AI: If the user says 'I ate 2 eggs' or 'I had a chicken breast', they will NOT provide the protein amount. YOU (the AI) must ESTIMATE the protein grams based on standard nutritional data and pass it to 'protein_g'.
Logs that you went to the gym.
Records a payment (payback) and automatically settles the oldest debts.
check_social_finances accepts free-form query_type (string) instead of enum. Should declare query_type as one of ['BALANCE', 'HISTORY'] to prevent invalid queries. Current implementation will fail silently or return confusing results if LLM passes unexpected values like 'SUMMARY' or 'BALANCE_HISTORY'.
log_debt uses human-friendly identifiers ('Me', 'Name') in borrower/lender params and requires agent to reason about perspective (user owes vs owes to user). This forces the LLM to infer semantics. Better pattern: accept user_id or canonical 'current_user' and separate tools (log_debt_owed, log_debt_owed_to_me) or a clearer enum direction parameter (e.g., direction='owed_by_me'|'owed_to_me').
Output schemas are not documented for any tool. Clients and LLMs cannot infer the structure of responses. check_social_finances and analyze_spending return unstructured multi-line strings that require parsing; check_fitness_stats does the same. Responses should be structured JSON with documented field types.
Error handling returns bare exception strings (e.g., 'Database Error: {str(e)}') without guidance on recovery, retryability, or next steps. LLM cannot determine if error is transient, user-fixable, or fatal. Missing error classifications and recovery instructions.
No input validation visible. Tools accept string inputs (names, item descriptions, query types) without sanitization, format validation, or length limits. No protection against SQL injection via RPC parameters, command injection, or path traversal if filesystem operations are later added.
log_personal_expense includes internal clarification logic (NEEDS_CLARIFICATION status) that returns instructions to the AI. This couples the tool interface to LLM behavior and makes the tool harder to understand statically. Better: return a structured response (e.g., {status: 'clarification_needed', prompt: 'Is X healthy?'}) and let the client/LLM layer decide how to respond.
check_social_finances and check_fitness_stats return potentially large result sets without pagination, limits, or total counts. No documentation on maximum results. If a user has 500 transactions, all will be returned as multi-line formatted strings, exhausting context window.
Numeric parameters (amount, protein_g, days) lack min/max constraints in descriptions. protein_g should validate [1, 500], days should validate [1, 365], amount should validate > 0. Without constraints, LLMs may pass absurd values (negative amounts, 10,000g protein).
No idempotency guarantees documented. If log_personal_expense or log_debt is called twice with identical arguments, will it create duplicates or be deduplicated? No guidance for agents on safe retry behavior.
No evidence of permission gates or audit trails. Tools modify sensitive financial data (expenses, debts, payments) with no authorization checks, no logging of who called what, and no compliance traceability.