The server defines 3 tools with reasonable structure, but falls short of production quality. All tools have descriptions and input schemas, but several critical gaps emerge: (1) Parameter descriptions vary in completeness, some are thorough (e.g., 'sortBy' in list-words explains each enum value and when to use it), while others are minimal (e.g., 'query' in search-words has no guidance on valid formats). (2) Error handling is present in code but not documented in tool descriptions, LLMs cannot read try-catch blocks, so recovery guidance is missing. (3) Output schemas are NOT documented in tool definitions; callers must reverse-engineer the response shape. (4) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present, forcing LLMs to guess whether tools are safe to retry or have side effects. (5) Tool composition is reasonable (three distinct operations: add, list, search), but add-words-batch combines creation + audio + Anki integration, a single tool doing multiple things. Naming follows verb_noun convention adequately (add-, list-, search-), but descriptions could be more LLM-optimized for selection clarity.
Add words to vocabulary list and create Anki cards (supports both single and batch operations). For single word, pass an array with one item.
List words from Anki with various sorting options: Sorting Methods: - due: Sort by due date (less familiar words first) Best for: Regular review sessions, focusing on words that need immediate attention Use when: You want to practice words you're struggling with - recent: Sort by review frequency (most recently reviewed first) Best for: Reviewing your recent learning progress Use when: You want to see what you've been studying lately - difficulty: Sort by accuracy rate (most challenging words first) Best for: Targeted practice on problematic words Use when: You want to focus on words with high error rates - added: Sort by creation date (newest first) Best for: Reviewing recently added vocabulary Use when: You want to practice new words you've just learned - lapses: Sort by number of lapses (most forgotten first) Best for: Identifying consistently troublesome words Use when: You want to focus on words you frequently forget Note: Results are limited to prevent overwhelming output. Use the limit parameter to adjust.
Search words and definitions in Anki deck
Output schemas not documented. Tools return structured data (e.g., add-words-batch returns {content: [{type: 'text', text: ...}]}, list-words returns card objects with interval/reps/lapses fields) but LLMs cannot see this structure in tool definitions. Callers must reverse-engineer by inspecting code or trial-and-error.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are absent. LLMs cannot determine which tools are safe to retry (list-words is idempotent; add-words-batch is destructive and should not be auto-retried without confirmation). This invites retry loops that duplicate Anki cards.
Error handling is implemented in code (try-catch blocks log and return error messages) but error recovery guidance is NOT in tool descriptions. LLMs cannot read exceptions, they need descriptions like 'If Anki is unreachable, ensure AnkiConnect is running on http://localhost:8765' in the tool definition itself.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
search-words description is at the minimum threshold (20 chars) and lacks critical context. Does query do substring matching or exact word match? Case-sensitive? Is it searching word text or definition text or both? LLMs cannot infer this and may misuse the tool.
add-words-batch combines vocabulary curation (user provides words), audio synthesis (AudioService), and Anki card creation (AnkiService) in one tool. This violates single-responsibility principle. Consider splitting into separate tools or clearly documenting which steps can fail independently.
Parameter 'tags' description says 'alphanumeric, hyphens, and underscores only' but no regex pattern or length limit is provided. Code does not validate tag format, LLMs may pass invalid tags, which Anki rejects silently.
No pagination guidance. list-words accepts limit (1-100, default 20) but does not return a cursor or total_count. If Anki deck has >100 cards, agents cannot iterate through all results. Code sorts and slices locally, which may fail with very large decks.