An AI-powered tutoring platform with multi-agent classroom mode, project-based learning, skill assessment, and document parsing capabilities. Provides APIs for chat, grading, content generation, and social features.
This Next.js HTTP server exposes 13 tools via API routes, but exhibits significant definition quality gaps across the board. Tool naming is inconsistent (mix of verb_noun, noun_verb, and pure nouns). Most tool descriptions lack clarity on WHEN to use them vs. similar tools, and many are under the 50-char minimum for LLM discoverability. Input schemas are present but sparse, many parameters lack descriptions, and enums are missing where they should exist (e.g., 'plan' accepts 'free|basic|pro|admin' but no enum constraint). Output schemas are entirely absent from the source; the code generates responses but does not document what fields to expect. Error handling is minimal, most routes return generic HTTP status codes without actionable recovery guidance. Two WRITE tools (webhook handlers) lack permission gates and audit logging. The overall effect is that an LLM would struggle to reason about tool selection and would need to guess at parameter semantics.
Processes Afdian payment webhook notifications to upgrade user subscription plans based on plan ID mapping.
OAuth callback endpoint that exchanges authorization codes for sessions and handles password recovery flows.
Main chat endpoint supporting tutor-guided tutoring with Socratic method, persona-based learning, story context, cross-tutor memory, and custom prompts. Includes rate limiting, plan-based access control, and daily message limits.
Multi-agent conversation endpoint that orchestrates Teacher + Assistant + AI Students in a classroom setting. Returns SSE stream with multiple agent responses per turn.
Auto-detects system proxy and compares responses from multiple tutors on a single question using the Socratic method.
Generates self-contained interactive HTML5 simulations/experiments for hands-on learning. Uses the Code model for best code generation quality.
Output schemas completely absent. Code generates responses (agent messages, slides, HTML simulations, quiz grades) but LLMs have no formal specification of response structure, field names, or types. This forces LLMs to parse unstructured text or guess at field semantics.
No error handling guidance. Routes return HTTP status codes (404, 500) but do not tell LLMs what to do next, retry, ask user, call a different tool, or give up. Recovery paths are completely absent.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Generates a set of educational slides about a skill topic. Returns structured slide data with title, content, speaker notes, and layout type.
AI grades short answer responses using the Generation model. Evaluates student answers against expected answers with leniency for understanding.
Processes LemonSqueezy subscription webhook events to manage user plan upgrades/downgrades and track payments.
Extracts text from PDF/TXT uploads and generates course outlines with suggested quizzes and knowledge points.
Project-Based Learning chat API providing AI-guided project collaboration. Uses Teaching model for project mentor role to guide students through project milestones.
Generates social content including group chat between tutors and a student diary entry reflecting on the learning session.
Converts text to speech using OpenAI TTS (if configured) or falls back to Web Speech API. Auto-detects language and selects appropriate voice.
Three WRITE tools (afdian_webhook, lemonsqueezy_webhook, auth_callback) lack permission gates and audit logging. No checks that the calling entity has authority to upgrade subscriptions or modify auth state. No traces of who triggered the action or what changed.
Webhook endpoints (afdian_webhook, lemonsqueezy_webhook, auth_callback) should not be exposed as agent-callable tools. These are for receiving external events (POST from payment gateway, OAuth provider), not for agents to invoke. Exposing them as tools creates security and semantic confusion. They should be internal background services.
Many parameter descriptions are missing or vague. E.g., 'history' in classroom_chat is documented only as 'Array of previous classroom messages' with no detail on required fields (role, name, emoji, content). 'customPrompt' in chat lacks max length constraint. 'milestones' in pbl_chat array structure is undocumented.
Generic tool naming reduces discoverability. 'chat' is ambiguous when five distinct chat-like tools exist (chat, classroom_chat, pbl_chat, compare_tutors). Tool names do not disambiguate intent or context, forcing LLMs to read descriptions and guess.
Enum constraints missing for bounded values. 'plan' accepts 'free|basic|pro|admin' but no enum declared in schema, LLMs may hallucinate invalid plan names. Same for 'model', 'role', 'type' in webhook endpoints.
Numeric parameters lack range constraints. 'slideCount' (default 8) has no min/max documented, can an agent request 1 or 1000 slides? 'currentMilestone' in pbl_chat is a number with no bounds checking.