Multi-server MCP setup with math and expense tracking tools integrated into a LangGraph-based chatbot with RAG capabilities
20 tools across two domains (expense tracking, math). All tools have descriptions and input schemas with typed parameters. However, descriptions are minimal (10-50 chars, well below the 194-char baseline), lack context on WHEN to use tools, and omit error guidance. No output schemas documented. Parameter descriptions are present but terse. No tool annotations (readOnlyHint/destructiveHint). Error handling is absent, no recovery guidance for invalid inputs. Composition is reasonable (single responsibility per tool), but descriptions do not explain dependencies or prerequisites. Average per-tool score: 52.
Add a new expense.
Show spending by category.
Delete expense by index.
Compute derivative. Example: x**3 + 2*x
Evaluate a mathematical expression. Example: 2 + 3*4
Expand an algebraic expression. Example: (x+2)*(x+3)
Factor an algebraic expression. Example: x**2 - 9
Descriptions are critically short (10-50 chars vs 194-char baseline). Most lack context on WHEN to use the tool, prerequisites, or what it returns. E.g., 'Calculate total expenses' does not explain whether it sums all expenses, filters by date, or returns a breakdown.
No output schemas documented. LLMs cannot infer what fields to expect from responses. E.g., does list_expenses return [{amount, category, description, id, date}] or just [amount, category]? This forces agents to guess and risks parsing failures.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | 2026-07-28+ | v2 |
Calculate the greatest common divisor of two integers.
Compute indefinite integral. Example: x**2
Calculate the least common multiple of two integers.
Return all expenses.
Calculate determinant.
Multiply two matrices.
Calculate the mean of a list of numbers.
Calculate the median of a list of numbers.
Return prime factorization.
Simplify an algebraic expression. Example: x**2 + 2*x + x**2
Solve equations. Example: equation = "x**2 - 5*x + 6" variable = "x"
Calculate the standard deviation of a list of numbers.
Calculate total expenses.
No error handling or recovery guidance. If delete_expense receives an invalid index, what happens? If evaluate receives malformed math syntax, does it return a parse error or crash? No tool documents error cases or tells the LLM how to recover.
Destructive tool (delete_expense) lacks confirmation or dry-run support. No tool annotation (destructiveHint) to signal to the LLM that this operation is irreversible. Agents may delete expenses without user confirmation.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). LLMs cannot distinguish safe read-only tools from destructive ones, risking unsafe agent behavior. E.g., list_expenses should be marked readOnly; delete_expense should be marked destructive.