Educational MCP server repository demonstrating MCP concepts including tools, prompts, and resources using FastMCP framework
This server exhibits significant definition quality gaps across multiple dimensions. Of 15 tools, 12 have basic descriptions (10-50 chars) that lack context for LLM selection and do not explain when to use each tool versus similar alternatives. Tool naming includes anti-patterns: 'generate_and_create_exercises' violates the 'one responsibility' rule (split into generate + create); 'run_in_terminal_cmd' is redundant. Input schemas are present for most tools but parameter descriptions are minimal or absent in several cases. Output schemas are not documented, LLMs cannot reason about what fields to expect from responses. Error handling is absent: no recovery guidance, no classification of retryable vs fatal errors. The server mixes multiple concerns across tool silos (exercises in tools_server.py, research in AIResearchHub/server.py, study progress in resources_server.py, terminal execution in server_part_2.py), creating composition hazards: there is no clear data flow between tools, no guarantees that tool A's output contains the IDs tool B needs. Critical security issue: 'run_in_terminal_cmd' accepts arbitrary shell commands with no sanitization, validation, rate limiting, or permission gating, LLMs can be tricked into running destructive commands. The server uses deprecated sampling pattern (tools_server.py calls ctx.session.create_message with SamplingMessage), this should be replaced with direct LLM provider integration.
Add a paper to a research entry
Add a repository to a research entry
Use this tool to create new Python exercises from a dictionary containing generated exercises.
Generate exercises using sampling and create them automatically.
Generate GitHub search commands for finding code implementations
Get information about a specific speaker's sessions at the AI Engineer Conference
Get study progress for a user.
Tool naming violates single-responsibility principle: 'generate_and_create_exercises' combines two separate actions (generate + create) into one tool name, making it ambiguous and forcing the LLM to reason about multiple concerns.
Critical security vulnerability: 'run_in_terminal_cmd' accepts arbitrary shell commands with zero input validation, sanitization, rate limiting, or permission gating. LLMs can be prompted to run destructive commands (rm -rf /, password exfiltration, etc.). Tool requires destructive annotation, dry-run mode, and explicit user confirmation.
Duplicate tools cause LLM selection ambiguity: 'get_study_progress' (#9) and 'get_users_progress' (#11) appear to do the same thing. No clear reason for both to exist; composition hazard if they return different schemas.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 35 | - | v1 |
Get the study progress for a user.
List all created exercises.
List all available exercises for a specific level.
Start researching a topic and get research ID
Run a command in the terminal.
Start a study buddy session for the user using the current exercises_db and progress.
Track the study progress of a user.
Update research status and add notes
No output schemas documented. LLMs cannot reason about response structure or plan follow-up tool calls. Responses are inferred from code but not declared in tool registration. Example: 'research_topic' returns research_id, total_research_topics, but LLMs do not know this without executing.
Parameter descriptions are missing or too brief (under 50 chars). Examples: 'list_exercises' has no input params but no explanation of output structure; 'get_speaker_session' speaker_name param has no description; 'update_research_status' status param has no enum constraint or valid value documentation (valid values 'pending, active, complete' are only in code).
Tool descriptions are too brief (under 50 chars avg) and lack discovery context. Examples: 'List all created exercises' (29 chars) does not explain discovery role or when to call; 'Get study progress for a user' (29 chars) does not differentiate from get_users_progress. Missing: WHEN to use, WHY to prefer one tool over another.
Deprecated sampling pattern used: tools_server.py calls ctx.session.create_message(SamplingMessage) to invoke the LLM. This pattern was deprecated in 2025-03-26 and removed from the current spec (2026-07-28). Migration: replace with direct integration to LLM provider API (OpenAI, Anthropic, etc.) to generate exercises without server-side sampling.
No error handling or recovery guidance. Tools return bare success/failure responses with no guidance on next steps. Example: 'research_topic' returns {success: True, ...}; if it fails, no hint on why (e.g., 'Research limit exceeded, try archiving old entries') or what the LLM should do next.
Composition hazards across multiple tool silos (tools_server.py, AIResearchHub/server.py, resources_server.py, server_part_2.py) with no guaranteed data flow. Example: exercises created by 'generate_and_create_exercises' are stored in exercises_db, but 'track_study_progress' references 'completed_exercises' (a list of strings) with no schema guarantee that these names match tool output. No tool chains documented.
Enum constraints missing for status/level parameters. 'update_research_status' accepts status as free-form string with valid values only in code ('pending, active, complete'). 'list_exercises_for_level' and 'track_study_progress' accept level as string with no enum (valid values 'beginner', 'intermediate', 'advanced' are not declared). LLMs will hallucinate invalid values.