An MCP host that connects to multiple MCP servers (cookie jar, pantry, legacy), manages tool namespacing, implements host-side approval gates for destructive operations, and runs an agent loop with Claude API integration
Two tools with clear, action-oriented names and detailed descriptions. Both have complete input schemas with typed parameters and enums. cookie_jar uses verb_noun naming (good), descriptions are 150-200 chars (within baseline), and parameters are well-constrained. smash_jar includes a deliberate confirmation pattern (SMASH enum) to prevent accidental destruction. However, output schemas are not documented, LLMs cannot infer what fields are returned. Error handling guidance is absent; no recovery hints for failure cases. Tool composition is sound (separate concerns), but missing pagination/result limits documentation and response field details.
Look inside the cookie jar, put cookies in, eat cookies, or refill it. This is THE cookie jar for this app and its contents are stored in a real database, so the count is correct across machines and survives restarts. Use this for any question about how many cookies there are.
PERMANENTLY destroy the cookie jar: sets the cookie count to zero AND erases the jar's entire history of past events. This cannot be undone and there is no backup. Use only when the user explicitly asks to destroy, smash, wipe, reset from scratch, or erase the jar and its history.
Output schemas not documented. LLMs cannot infer what fields are returned by cookie_jar (count? history? timestamp?) or smash_jar (success? confirmation?). This forces agents to guess and wastes tokens on trial-and-error.
No error handling guidance. If cookie_jar fails (database down, invalid count), the response should guide recovery: 'Database unavailable, retry in 30s' or 'Count must be 1-500, got X'. Currently agents receive no actionable next steps.
smash_jar description lacks explicit state-change warning. While 'PERMANENTLY destroy' is clear, the description should state 'This operation cannot be undone and has no backup' more prominently to prevent accidental invocation.
cookie_jar 'count' parameter lacks context on what it represents. Description says 'How many cookies to add or eat' but does not clarify: is this absolute or relative? Does 'eat 5' mean remove 5 or set to 5? Ambiguity invites misuse.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 79 | 2026-07-28+ | v2 |
No pagination or result limits documented. If cookie_jar returns a history array, how many items? Is there a limit? Unbounded responses risk context window exhaustion.