Agentic AI Backend with MCP Protocol for Smart App. Features RAG (Retrieval-Augmented Generation) from rag_data.txt, MCP-style tool calling (search, calculate, get_time), Ollama LLM integration, and FastAPI REST endpoints.
This MCP server has 4 tools with basic implementations, but falls well short of production quality across naming, descriptions, schema precision, and error handling. Tool names are generic (search_emsi, get_current_time, calculate, get_grade_requirements). Descriptions exist but are minimal (10-80 chars, baseline 194 chars). Input schemas are extremely sparse, the source shows only parameter names and types in informal notation (e.g., 'query: string'), NOT proper JSON Schema with constraints, enums, or validation rules. Critically, there is NO documented output schema for ANY tool, which violates the baseline requirement that 100% of A+ tools have documented return types. Error handling is entirely absent, no actionable recovery guidance, no error classification, no input validation beyond a trivial 'allowed' character set. The tools lack composition discipline: search_emsi, calculate, and get_grade_requirements are domain-specific and useful, but get_current_time is a generic utility that should not coexist with domain tools. No pagination, no structured response format guidance, no idempotency declarations, no field naming consistency with parameter names. Overall, this reads like a prototype, not a production-ready MCP implementation.
Perform a mathematical calculation
Get the current date and time
Get requirements for a specific grade level
Search EMSI knowledge base for information
NO documented output schemas for ANY tool. Code returns strings (search_emsi), timestamp (get_current_time), result string (calculate), and dict (get_grade_requirements), but NONE are formally declared in tool definitions. Baseline: '100% of A+ tools have documented return types.' This is a HARD MISS.
Input schemas are informal and incomplete. Tools expose bare parameter names + types (e.g. 'query: string') but lack JSON Schema structure with constraints, minLength, maxLength, pattern, enum, or format.
Descriptions are too short. Baseline: 194 chars average (p10=34, p90=392). No guidance on WHEN to use each tool or prerequisites. Per pattern:tool-description, 'Write descriptions as if prompt-engineering. State WHAT the tool does, WHEN to use it, and any prerequisites.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 38 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 31 | - | v1 |
Parameter 'level' in get_grade_requirements enumerates valid values IN the description ('beginner/intermediate/advanced/expert/master/legend') instead of using a JSON Schema enum constraint. Free-form strings invite hallucinated values.' This forces LLMs to parse text instead of selecting from a formal constraint.
Error handling is absent or non-actionable. calculate returns bare strings 'Invalid expression' and 'Calculation error' (13-18 chars) with no recovery guidance. Per pattern:recovery-guide: 'Error responses must tell the LLM what to do next: "User not found. Try search_users() with a partial name." A raw error code or stack trace gives the agent nothing to act on.' No error classification (retryable vs fatal), no input validation details, no suggestion of alternatives.
Tool composition is weak. get_current_time is a generic utility that pollutes a domain-specific tool collection. Per pattern:tool: 'Each tool should do exactly one thing.' This tool does nothing domain-specific and adds noise to an EMSI-focused agent. Tools do not reference each other or chain naturally, no output field alignment per pattern:tool-chain.
No idempotency or safety declarations. search_emsi and get_grade_requirements are read-only (marked in Risk field), but this is NOT documented in tool descriptions or input schemas. Per pattern:command-tool: 'If the tool modifies state (creates, updates, deletes, sends), the description must say so. Agents need to know which calls are safe to retry and which have irreversible consequences.' Neither the safe tools nor hypothetical unsafe tools have this guidance.
Parameter naming lacks type suffixes and consistency. 'query' (search_emsi), 'expression' (calculate), 'level' (get_grade_requirements) are bare nouns without clarification of format/type. A bare "user" parameter leaves LLMs guessing.' This server's params are bare and rely on description text to explain semantics.