A dual-mode MCP server providing expense tracking and arithmetic operations
Mixed quality across two separate MCP registrations. Expense tools have basic descriptions and parameter schemas, but output schemas are undocumented and error handling is generic. Math tools have minimal descriptions (docstrings only, under 20 chars effectively) and lack parameter context. Overall, the server follows basic MCP registration but falls short of production standards: no tool annotations, no pagination support for list operations, no structured error recovery guidance, and thin documentation.
Return a + b.
Add a new expense entry to the database.
Provide expense categories resource
Return a / b.
List expense entries within an inclusive date range.
Return a % b.
Return a * b.
Return a ** b.
Math tool descriptions are trivial (< 20 chars effectively). Docstrings 'Return a + b', 'Return a - b', etc. lack context for when an LLM should invoke them. No indication of use cases, error cases (e.g., division by zero), or relationship to downstream tools.
No output schemas documented for any tool. list_expenses and summarize return lists, but no documented structure for items, fields, or pagination info. add_expense returns a success dict with 'status', 'id', 'message', but this is inferred, not formally declared. LLMs cannot plan downstream calls without knowing return structure.
list_expenses and summarize lack pagination support. No offset/limit parameters, no total count, no next_cursor. If expense DB grows to thousands of records, returning all matches will blow context window and degrade LLM reasoning.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Return a - b.
Summarize expenses by category within an inclusive date range.
Error handling is generic and non-actionable. Both add_expense and list_expenses catch exceptions and return a flat dict with 'status' and 'message'. No guidance for retry, no error classification (retryable vs fatal), no recovery hints. E.g., 'Database error: ...' tells the LLM nothing about what to do next.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). add_expense modifies state (WRITE risk marked), but the tool definition lacks a destructive hint. Math tools are pure functions but lack readOnly annotations. Agents cannot infer retry safety or side effects.
'categories' tool is named as a noun, not a verb. Should be 'get_categories' or 'list_categories'. Resource-backed tools should still follow verb_noun convention in the tool registry. Current name is ambiguous (could be data, not an action).
Math tool parameters lack descriptions. 'a' and 'b' are bare names without context: are they floats, ints, or numeric strings? Code shows _as_number() coercion, but LLM sees only type hints 'a: float, b: float' with no description of what happens to strings or what errors are thrown.
No mention of constraint violations or invalid input handling in descriptions. E.g., div(a, 0) will raise ZeroDivisionError, but the tool doesn't mention this or guide the LLM on recovery. pow with negative exponents also risk overflow or precision loss.