An MCP server that uses a Domain-Specific Language (DSL) to manage course content, including chapters, quizzes, exams, and resources. Parses course definitions using ANTLR-generated lexer and parser, builds an AST, generates Python code with embedded course content, and exposes MCP tools for course and student management.
This server exhibits critical definition quality gaps across the board. While 13 tools are defined with clear action verbs and basic input schemas, nearly all lack substantive descriptions, parameter-level documentation, and output schema definition. Tool descriptions are present but universally brief (20-40 chars), violating the 10-1024 character guideline and failing to explain WHEN to use each tool or WHAT it returns. Input schemas show type and description fields but parameter descriptions are generic one-liners (e.g., 'The student's unique identifier' for student_id repeated across 5 tools). No tool documents its output structure, forcing LLMs to guess what fields are returned. Error handling is completely absent, no recovery guidance, no categorization, no actionable error messages. The 'generated_code.py' backend is not visible, so actual tool implementations cannot be verified. Tools like submit_quiz_answer and complete_quiz_session modify state but descriptions do not warn of side effects. Composition shows reasonable verb_noun naming (get_, list_, create_, mark_, start_) but lacks the depth production tools require.
Complete a quiz session, calculate the final score, and determine if the student passed
Create a new student record with initial course enrollment
Retrieve detailed information about a specific chapter, including title, description, content items, and objectives
Retrieve course metadata including name, author, description, level, and tags
Retrieve exam details including title, instructions, questions, passing score, and time limit
Retrieve a quiz including its title, instructions, questions, settings, and metadata
Retrieve a student's progress through the course, including completed chapters and quiz scores
Output schemas completely undocumented. No tool declares what fields it returns. LLMs cannot plan downstream calls or extract results reliably. For example, get_student_progress returns unknown structure, is it {progress_percent, completed_chapters, quiz_scores}? This forces trial-and-error.
Tool descriptions are critically short (20 - 40 characters) and lack actionable context. Example: 'Retrieve a student's progress through the course, including completed chapters and quiz scores' (94 chars) is acceptable, but many like 'Start a new quiz session for a student, initializing the quiz attempt with tracking' (81 chars) omit WHEN to use the tool vs alternatives or what happens on error.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 34 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
List all chapters in the course with their titles and numbers
List all exams in the course with their names and titles
List all quizzes in the course with their titles and numbers
Mark a chapter as completed for a student and record the completion date
Start a new quiz session for a student, initializing the quiz attempt with tracking
Submit a student's answer to a quiz question and record whether it is correct
Parameter descriptions are generic and non-actionable. Every 'student_id' parameter says 'The student's unique identifier' without explaining format (is it an email, username, or opaque ID?) or whether lookup is supported. Parameter descriptions average ~20 chars; baseline is 72 chars. Example: submit_quiz_answer accepts 'student_answer' described only as 'The student's answer to the question', does it accept free text, numeric, multiple-choice letter, or JSON?
No error handling or recovery guidance. Tools that modify state (start_quiz_session, submit_quiz_answer, complete_quiz_session, mark_chapter_complete, create_student) do not declare what errors can occur or how to recover. For example, what happens if create_student is called with a duplicate student_id? Should the agent retry, call a different tool, or ask the user?
Destructive operations lack warnings or confirmation steps. submit_quiz_answer and complete_quiz_session modify course data but descriptions do not warn of side effects or offer a dry-run. If an agent submits wrong answers or completes a quiz prematurely, there is no undo path.
List tools (list_chapters, list_quizzes, list_exams) do not declare pagination, limits, or result count. If a course has 1000 chapters, does list_chapters return all 1000 or cap at 50? Are offset/limit parameters supported? Baseline pattern requires pagination for list operations.
Tool implementations in generated_code.py are not visible in source; tool definitions are inferred from tool registration metadata only. Actual behavior, error handling, and output structure cannot be verified.