LLM-powered personalized learning assistant for students. Provides test result analysis, weak area identification, study schedule generation, question answering, practice question generation, and content search.
LearnMate exhibits significant definition quality issues across multiple dimensions. While tool names generally follow verb_noun conventions (upload_test_result, generate_schedule, ask_question), descriptions are often generic and lack depth needed for LLM tool selection. Parameter schemas are present but incomplete, many parameters lack detailed descriptions or constraints. Most critically, output schemas are not documented, forcing LLMs to infer response structure. Error handling is minimal, with generic HTTPExceptions providing no recovery guidance. The server demonstrates foundational tool structure but falls well short of production-grade quality expected for agent integration.
Ask a question with vector database context retrieval. Returns answer with source citations and confidence score.
Ask an educational question and receive an explanation tailored to difficulty level and learning style.
Check a student's answer against the correct answer and provide detailed feedback.
Generate practice questions on a given topic at specified difficulty level.
Generate a personalized revision study schedule based on weak areas, available study time, and exam date.
Search educational content by query, topic, and difficulty level.
No documented output schemas for any tool. LLMs cannot determine what fields to expect from responses, forcing inference and downstream errors.
Generic, non-actionable error handling. HTTPException(status_code=500, detail=f'Failed to process: {str(e)}') provides no recovery guidance or error classification. LLMs receive stack traces instead of 'Try uploading a different PDF format' or 'User not found. Call search_users() first.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 30 | - | v1 |
Upload a PDF file, process it to extract chapters and pages, and load documents into vector database for semantic search.
Upload and analyze a PDF test result. Extracts test metrics, identifies weak areas, and provides recommendations.
Parameter descriptions are minimal or missing. 'learning_style' parameter in ask_question lacks enum constraints and usage context. 'difficulty_level' in multiple tools is not validated against acceptable values. No description of valid enums in ask() for 'subject' parameter.
Tool 'ask' has ambiguous naming. It competes with 'ask_question' for essentially the same purpose (answering questions). LLMs will struggle to distinguish when to use which tool. Violates single-responsibility principle.
Search tools (search_content) return structure is undocumented. Without knowing what fields are returned (title, url, snippet, score?), agents cannot extract or chain relevant data.
Parameter 'file' in upload_test_result and upload_pdf lacks validation guidance. File size limits, accepted MIME types, and maximum page counts are not documented.
generate_schedule expects 'weak_areas' as array of objects with 'confidence_score', but upload_test_result returns 'weak_areas' as array of strings (from the code, weak_areas list comprehension produces dicts). Parameter dependency not documented.
No pagination support on search_content. Free-form search queries could return hundreds of results, overwhelming context window. No documented limit or cursor mechanism.