Educational AI assistant backend with course materials search, learning plan management, evidence tracking, and external tool integration via MCP
CoursePilot 2.0 has 8 tools with mixed quality. Naming follows verb_noun convention well (search_materials, list_materials, get_plan, etc.). Descriptions are present for all tools but vary in completeness, some are context-rich (plan_update, emit_evidence) while others are minimal (list_materials, artifact_read). Input schemas are visible and mostly well-structured with type definitions, but many parameter descriptions are sparse or missing. Critical gap: no output schemas documented for any tool. Error handling is not evident in the visible code. The Chinese-language descriptions are detailed for domain-specific tools but do not substitute for comprehensive parameter documentation in English.
读取本会话最近的跨轮产物(如练习题目与私有答案要点)。
列出当前课程概念目录里的概念及其 id。归因证据前必须先调用它——concept_id 只能取自这里,目录外的概念一律用 topic_hint。
把一次可判定的作答结果写成证据事件。掌握度数值由服务端的确定性算法更新,不要在参数里给分数或掌握程度。
读取当前课程学习档案:最近的证据事件、各概念掌握度与弱项,用于回答学习进度类问题。
读取当前课程的学习计划(由服务端持久化,可能不存在)。
列出当前课程资料库中的教材文件名。
重写学习计划里今天及以后的待办条目。必须先 get_plan 拿到 expected_version;一次给出这段周期的完整条目(长期计划就一次给完),不要分多次追加。只有用户在对话里明确要求排计划或调整计划时才可调用。用户点名要清空某几天(例如出差、休息),那几天就一条学习任务都不许留,内容匀到别的日期上;匀不完就压缩每天的量,不要留一条在原地。
No output schemas documented for any tool. LLMs cannot plan downstream calls or validate responses without knowing what fields to expect.
list_materials and artifact_read have minimal descriptions (under 50 chars). Insufficient context for tool selection. list_materials says 'list files' but doesn't explain format, size limits, or when to call it vs search_materials.
get_plan and get_archive have sparse descriptions without explaining return structure, use-cases, or prerequisites. 'Read archive' and 'Read plan' lack actionable context.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 41 | 2024-11-05+ | v1 |
在当前课程的教材资料库中检索相关内容。系统已用用户原话检索过一次;当已有证据不足、用户追问细节、或需要换关键词(含中英互译)时再调用。
Parameter descriptions are often missing or trivial. 'keyword' in concept_search has no description of format or filtering behavior. 'items' in plan_update is documented but nested object properties lack inline guidance.
No error handling guidance visible. No documentation of retryable vs fatal errors, missing resource responses, or validation failures. LLMs cannot recover from failures without explicit recovery paths.
plan_update requires 'expected_version' for optimistic concurrency, but no documentation of what happens on version mismatch or how to recover. No conflict resolution guidance.
concept_search has an optional 'keyword' parameter but no documentation of what fields are returned or whether results are sorted by relevance.
emit_evidence requires 'concept_id' from concept_search or uses 'topic_hint' as fallback, but the fallback behavior and human-review SLA are not documented. LLMs may not understand when to use which.