MCP server for ChaosLimba — read-only access to Romanian learning platform data
ChaosLimba MCP server demonstrates solid tool definition practices with consistent naming conventions, descriptive docstrings for most tools, and well-structured JSON schemas using Zod. However, there are significant gaps in output schema documentation, some parameter descriptions lack actionable detail, and error handling guidance is minimal. The server registers 20 tools via the mcp-sdk with good initial coverage, but lacks documentation of expected return structures that would help LLMs chain calls effectively. Tool naming is consistently verb-prefixed (get_, add_, cl_) which aids discoverability. Most tools have 10-200 character descriptions, aligning with baselines. Per-tool analysis reveals strong schema definition in parameters but weak/missing output documentation.
Inserts a new content item into the content_items table. Use during dev sessions to seed reading passages, audio content, etc.
Inserts a new reading comprehension question into the reading_questions table. Provide a passage, question, answer options, and the index of the correct answer.
Cross-references grammar_feature_map against content_items.language_features to show which grammar features have content coverage and which are gaps. The core instructional design audit tool.
Returns a summary of fossilization interventions — which patterns are being escalated and whether they are resolving. Shows max tier reached, total interventions, and resolution counts.
Returns content items, optionally filtered by difficulty level, topic, or type. When no difficulty filter is set, results are stratified across difficulty levels for even coverage.
Returns aggregated error patterns from error_logs across all users (anonymized). Useful for understanding where learners actually struggle.
Missing output schema documentation for all tools. Tool implementations return JSON via generic TextContent blocks without documenting field names, types, or structure. LLMs cannot infer what fields are returned or chain tools together without this information.
No error handling guidance provided. Tools return generic database result JSON without error messages that guide LLMs on recovery (e.g., 'User not found. Try search_users() with a partial name'). A failed database query returns minimal context.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Returns aggregated feature exposure data — how many times each grammar feature has been seen by learners, with correctness rates. Anonymized. Useful for finding undertaught or poorly-performing features.
Returns aggregated stats on AI-generated content — what types are being generated, for which error targets, listening rates, and estimated TTS costs. Anonymized. Useful for auditing the AI tutor's output quality and cost.
Returns all entries from grammar_feature_map, optionally filtered by CEFR level. Shows feature keys, names, categories, descriptions, prerequisites, and sort order.
Returns AI-generated learning narrative summaries — periodic reflections on learner progress with stats. Anonymized. Useful for auditing narrative quality and checking if the reflection system captures meaningful patterns.
Returns mystery vocabulary items — words learners encountered and flagged for exploration. Shows word, context, definition, examples, grammar info, and whether the learner has explored it. Useful for understanding organic vocabulary discovery.
For a given grammar feature, returns its full recursive prerequisite chain so you can audit whether the sequencing is pedagogically sound. Caps depth at 10.
Returns proficiency score history over time — overall, listening, reading, speaking, writing scores by period. Anonymized across all users. Useful for tracking whether content improvements translate to learner gains.
Returns reading comprehension questions, optionally filtered by CEFR level. Shows passage text, question, answer options, and correct index. Useful for auditing question quality and level coverage.
Returns a summary of all tables and their columns in the ChaosLimba database. Use this to orient yourself when first connecting.
Returns aggregated session data — session counts by type, average duration, and content engagement. All data is anonymized (no user IDs returned). Useful for understanding how learners actually use the app.
Returns stress minimal pairs — words where stress placement changes meaning (e.g., CÁsă vs caSĂ). Core pronunciation training data.
Returns AI tutor conversation starter questions, optionally filtered by CEFR level or category. Useful for auditing prompt quality and topic coverage.
Returns TTS (text-to-speech) usage stats — characters consumed per day. Useful for monitoring costs and usage trends.
Returns tutor opening messages keyed by self-assessment level. Shows how the AI tutor greets learners at different proficiency levels.
cl_add_content and cl_add_reading_question lack error handling examples. Insertion failures (duplicate keys, validation errors) should return actionable messages, not raw SQL errors. The sourceAttribution.originalUrl and other optional fields should document validation constraints.
Some parameter descriptions are vague. E.g., 'filter by CEFR level' (cl_get_grammar_map), does this filter the returned records or apply a pedagogical constraint? 'Sorted by X' is unclear if the consumer controls sort order. Several tools describe what data exists but not when an LLM should call them vs. similar tools.
cl_get_content uses stratified sampling when no difficulty filter is set. This algorithmic detail is buried in the description and not explained in parameter guidance. An LLM may not understand why passing no filter yields unexpected diversity, or that limit is divided by CEFR_LEVELS.length.
No mention of pagination for large result sets. Tools accept a 'limit' parameter but do not document offset/cursor patterns or total counts. cl_get_content could return 50 items easily; no guidance on fetching the next page.
Tool response payloads include raw database column names and nested JSONB fields without flattening. E.g., 'language_features' is returned as opaque JSON. An LLM cannot reason about the internal structure without documentation of expected keys (grammar[], vocabulary.keywords[], structures[]).
Database connection string (CHAOSLIMBA_DATABASE_URL) is required but no guidance on failure modes or retry behavior. If the database is unreachable, users see 'Error: CHAOSLIMBA_DATABASE_URL environment variable is required' on server startup, not a diagnostic about database connectivity.