This server has 11 tools with basic definitions but significant quality gaps. All tools ARE explicitly registered with @mcp.tool() decorators and have descriptions present, which is positive. However, descriptions are uniformly terse (6-50 chars), well below the 10-1024 baseline recommended in the rubric. Parameter descriptions exist but are minimal, e.g., 'ID of the gathering' provides no guidance on format, validation, or when to use alternatives. Input schemas are visible (JSON Schema with types), but lack validation rules, constraints, or enum definitions. Error handling is absent, the run_command wrapper returns generic JSON success/error blocks with no recovery guidance. No output schemas are documented. The architecture wraps a subprocess CLI tool rather than implementing business logic directly, which reduces visibility into actual error conditions and validation. Overall, this falls into the C-range (fair/significant gaps) due to missing parameter constraints, no output documentation, generic error handling, and descriptions that don't meet LLM-optimization standards.
Descriptions uniformly terse (6-50 chars, well below 10-1024 baseline). Example: 'List all gatherings' (19 chars) provides no guidance on pagination, result limits, or when to call it instead of show_gathering. LLMs cannot distinguish similar discovery tools.
Parameter descriptions lack validation rules and constraints. Example: gathering_id accepts 'string' with description 'ID of the gathering (format: yyyy-mm-dd-type)' but no enum, pattern, or length limits are enforced. LLMs cannot validate against the stated format. member_name and amount lack guidance on valid ranges (e.g., can amount be negative? must member_name be alphanumeric?).
Expand all tool descriptions to 50 - 200 characters, following LLM-optimization guidelines. Example: 'Create a new gathering to track expenses. Specify a unique ID (format: YYYY-MM-DD-<type>) and initial member count. Returns gathering details.' This tells the LLM WHAT, WHEN, and WHAT TO EXPECT.
Add parameter constraints: gathering_id should declare regex pattern (YYYY-MM-DD-.*), enum values, or length limits. member_name and amount should specify valid ranges (e.g., amount > 0, or amount can be negative for refunds). Document these as part of the parameter description AND JSON Schema pattern/minimum/maximum fields.
Document output schemas for each tool. Example: list_gatherings() should return {type: 'object', properties: {gatherings: {type: 'array', items: {type: 'object', properties: {gathering_id: {type: 'string'}, members: {type: 'integer'}, closed: {type: 'boolean'}}}, total: {type: 'integer'}}}. Agents need to know the exact field names and types.
Implement structured error responses with classification and recovery hints. Example: 'Gathering not found. available_gatherings: ["2024-01-15-lunch", "2024-01-20-dinner"]. Call list_gatherings() or check the gathering ID.' For validation errors: 'Invalid amount: -10. Amounts must be positive (record_payment supports negative values for refunds, use that instead).'
Add dry-run/confirmation support for destructive tools (delete_gathering, remove_member). Implement a pattern where the agent must call a separate confirm_<operation> tool or pass a --dry-run flag that returns what would happen without executing. Example: delete_gathering with force=true should return a confirmation request or detailed warning.
No output schemas documented. Tools return generic JSON objects via subprocess with keys like 'success', 'error', 'output' but no typed response structure is defined. LLMs cannot plan downstream calls (e.g., does list_gatherings return an array with 'gathering_id' or a dict with 'gatherings' key?). Output fields are not defined.
Error handling is absent. All tools wrap subprocess.run() and return generic JSON errors. No error classification (retryable vs fatal vs user-fixable), no recovery guidance, no actionable messages. Example: invalid gathering_id returns raw subprocess stderr without suggesting how to recover (e.g., 'call list_gatherings first').
Destructive operations (delete_gathering, remove_member) lack confirmation or dry-run support. An agent can delete a gathering with zero safeguards. No pattern:confirmation-request implemented.
Parameter 'force' in delete_gathering defaults to false, which is safe, but the tool definition lacks any mitigation pattern (dry-run, confirmation step) to prevent accidental deletion of active gatherings. Tool descriptions do not warn agents of the consequence.
No pagination or result limits. list_gatherings() returns all gatherings in one call. If a user has 1000+ gatherings, the response will be massive, exhausting context and increasing hallucination risk. Rubric baseline: tools returning lists should support limit/offset and return total count.
Subprocess architecture obscures actual validation and error messages. The run_command() wrapper calls 'python gatherings.py --json' and parses output, but we cannot see the actual CLI error messages or validation logic in the provided source (gatherings.py is truncated). This means we must assume error handling exists but cannot verify it.
Implement pagination for list_gatherings(). Add optional limit (default 20, max 100) and offset (default 0) parameters. Return {gatherings: [...], total: <count>, has_more: <boolean>, next_offset: <int>}. This prevents context explosion and follows the paginated-result pattern.
Validate all inputs in the MCP layer (not just subprocess). Check gathering_id format against the stated regex before calling the CLI. Check numeric ranges (e.g., amount > 0 unless refund). Return immediate, actionable errors rather than subprocess failures.
Add tool annotations: mark delete_gathering and remove_member with destructiveHint=true; mark show_gathering, list_gatherings, calculate_reimbursements with readOnlyHint=true; mark operations that produce the same output on retry with idempotentHint=true (create_gathering is NOT idempotent if it errors on duplicate ID). This helps agents reason about side effects.
Provide field documentation for commonly chained operations. If agents call create_gathering and then immediately add_expense, ensure create_gathering returns the gathering_id in the same format that add_expense expects. Document this dependency in both tool descriptions.