Model Context Protocol (MCP) server for vocabulary management. Exposes vocabulary operations as MCP tools accessible to Claude and other AI clients. Supports OAuth 2.0 authentication with configurable authorization endpoints. Includes Heisig-specific hanzi study tools.
Server has 4 tools with clear naming and reasonable descriptions. Tool names follow verb_noun pattern (bulk_add_vocabulary, add_vocabulary, delete_session, add_hanzi). Descriptions are moderately detailed (120-280 chars), providing context on when to use each tool. However, critical gaps exist: (1) output schemas are entirely undocumented, no tool specifies what fields are returned or their types; (2) error handling is absent, no guidance on what happens when operations fail, how to recover, or how to classify errors; (3) parameter validation constraints are minimal, numeric ranges and enum values are largely missing; (4) destructive operations (delete_session) lack confirmation/dry-run patterns. Tool definitions are clearly visible and properly registered via @mcp.tool decorators with explicit parameter type hints, so no inference penalty applies.
Add or enrich Heisig hanzi study cards (max 50). Each card needs the hanzi, a single English keyword (meaning only), pinyin with tone mark, and the tone number 1-5. Re-calling on a character that already exists enriches it in place and preserves its review schedule. Pass session_name to group NEW cards; existing cards keep their session.
Add a single vocabulary word to the personal study app. Use this when the user wants to save one word with its definition. Pass session_name to assign it to a named study session. The definition field is for meaning/usage only — do NOT put pinyin in it.
Add multiple vocabulary words at once to the personal study app (max 50). Use this when the user has asked to save several words from a conversation, or when you've explained multiple words and want to offer to save them all. Pass session_name to group all words under a named study session (e.g. 'Japanese N5 Verbs'). The definition field is for meaning/usage only — do NOT put pinyin in it.
Delete an entire vocabulary session and all its words. Use this when the user wants to remove all words from a named session and clear the session itself. The session is identified by its exact name.
Output schemas completely undocumented. No tool declares what fields are returned, their types, or data structure. All four tools return `str`, but the actual response format is opaque to LLMs. This forces agents to reason about return types without guidance, increasing hallucination risk.
No error handling guidance. Tools return strings but provide no documentation on error cases, failure modes, or recovery paths. E.g., delete_session silently fails if session not found, or throws an exception, LLM cannot distinguish and retry intelligently.
Destructive operation lacks confirmation pattern. delete_session irreversibly removes a session and all words, but has no dry-run, confirmation, or undo mechanism. Agents making mistakes will cause data loss.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 57 | - | v1 |
Parameter validation constraints missing. 'max 50' is stated in description text only, not as a schema maxItems constraint for bulk operations. Numeric tone parameter (1-5) has minimum/maximum in schema but other tools lack equivalent constraints.
Enum constraints absent. Language code in VocabWord is a free-form string with no enum list. This invites hallucinated language codes ('Japanes', 'ja-JP', 'japanese') when constrained values would be self-documenting.
Pinyin handling documented but not enforced in schema. Tool descriptions explicitly warn 'do NOT put pinyin in definition' (repeated 3 times across descriptions), but no validation or regex constraint prevents it. Agents may still be confused.
Parameter relationships undocumented. In bulk_add_vocabulary, both individual word items and the top-level session_name have session_name fields with overlapping semantics. Description says 'top-level overrides per-word', but this dependency is not formalized in either parameter description.
Missing pagination and result limits. No tools document page/limit parameters or total counts. If bulk_add_vocabulary returns results for 50 words, the response structure is unknown, does it return success/failure per item, or a single aggregate response?