CLI, REST API, and MCP server for the Deep Learning DIY course (dataflowr.github.io)
This server has 20 well-named, READ-ONLY tools with consistent verb-noun naming (list_, get_, search_, check_, suggest_, sync_). All tools have descriptions (100+ chars average) and JSON Schema with typed parameters. However, there are significant gaps in output schema documentation, parameter descriptions lack depth and constraints, and error handling guidance is absent. The tools follow a coherent curriculum domain (modules, sessions, notebooks, quizzes, transcripts) but responses are not constrained, list_modules could return all 25+ modules without pagination, and descriptions assume internal API knowledge rather than user intent. Tools like 'get_slide_content' and 'get_quiz_content' lack return type documentation. Parameter descriptions are present but generic (e.g., 'Module ID (e.g. '12', '2a')' repeats across 11 tools without explaining format constraints or valid ranges). No tool declares error recovery guidance. Overall, definitions are structured and readable but lack the precision required for confident agentic reasoning at scale.
Validate a student's quiz answer.
Get full course structure as context.
Get full details for a homework.
Get full details for a specific module. Includes description, notebooks with GitHub and Colab links, tags, and prerequisites. Use this when a student asks about a specific topic or module. Module IDs: '12' (Attention/Transformers), '2a' (PyTorch tensors), '18b' (diffusion), etc.
Fetch actual notebook cells from GitHub.
Fetch only exercise prompts + skeleton code from a notebook.
Get the GitHub and Colab URL for a specific notebook. Use this when a student wants to open or run a specific exercise. kind: 'intro' | 'practical' | 'solution' | 'bonus' | 'homework' (default: practical)
Output schemas not documented. Tools like get_slide_content, get_quiz_content, get_page_content do not document what fields/structure the response contains. LLMs cannot plan downstream processing or understand what data is available.
Parameter descriptions lack format constraints and valid ranges. 'module_id' appears in 11 tools with the same generic description '(e.g. '12', '2a', '18b')' but does not clarify: Are alphanumeric IDs like '12' and '2a' both valid? Is '2a' format specific? No regex, length limits, or enum provided. Repeated across many tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 28 | - | v1 |
Fetch module source markdown from dataflowr/website repo.
Get prerequisite modules for a given module.
Fetch quiz questions from dataflowr/quiz repo.
Get full details for a session, including all its modules, notebooks, and key takeaways. Use this when a student wants to understand what a session covers.
Fetch lecture slides from dataflowr/slides repo.
Fetch full content of a transcript concept note.
List all homeworks in the course.
List all modules in the dataflowr Deep Learning DIY course. Can be filtered by session number, tag, or GPU requirement. Returns module IDs, titles, descriptions, and tags.
List all course sessions with their titles and module IDs. Sessions group modules into ~2-3 hour teaching blocks.
Search modules by keyword across titles, descriptions, and tags. Use this to find relevant modules for a student's question or topic. E.g. query='attention' finds the Transformer module; 'generative' finds GANs, autoencoders, flows, diffusion.
Fuzzy search 318 concept notes from lecture transcripts.
What to study after a given module.
Compare catalog against website + slides repos.
No error recovery guidance. Tools provide no indication of what the LLM should do on failure. E.g., if get_module fails because module_id doesn't exist, there is no error message suggesting 'call search_modules or list_modules first to find valid IDs.' No recovery patterns.
Pagination not addressed for list tools. list_modules, list_sessions, list_homeworks, search_transcripts could return many results (25+ modules, 318 transcript notes) without limit/offset parameters or result-count constraints. LLMs may receive context-window-blowing responses.
Minimal descriptions for simple tools. list_homeworks ('List all homeworks in the course.'), list_sessions ('List all course sessions with their titles and module IDs.'), sync_catalog ('Compare catalog against website + slides repos.') are under 60 chars. Too terse, LLMs cannot determine why to choose this tool vs alternatives.
No idempotency or state mutation declarations. All tools are READ-ONLY (noted in the feature spec), but this is not reflected in tool descriptions or parameter annotations. LLMs cannot infer whether a tool is safe to retry without explicit annotation.
Inconsistent parameter naming for module ID. Used as 'module_id' in most tools, but no enum or pattern constraint. Similar tools use it without clarification on accepted format. 'kind' parameter in notebook tools uses free-form string instead of enum constrained to ['intro', 'practical', 'solution', 'bonus', 'homework'].
check_quiz_answer returns a boolean or validation result, but the actual return structure is not documented. Will the LLM receive {correct: true, feedback: '...'} or just a plain true/false? Unknown output schema.