MCP server exposing the Education Agent Skills Library as callable tools
The server defines 4 tools with reasonable naming (verb-prefixed: list_, get_, find_, suggest_) and adequate descriptions. All tools have input schemas with type and description fields. However, output schemas are not formally documented in the definition layer, responses are returned as unstructured text wrapped in { content: [{ type: 'text', text: ... }] }. Parameter descriptions are present but generic in some cases (e.g., 'domain' in list_skills lacks format/constraint detail). Error handling is basic: missing skills return error strings, but no recovery guidance is offered. The server lacks output schema documentation, pagination guidance, and detailed parameter constraints that would help LLMs compose calls reliably.
Search skills by tag, domain, evidence strength, or free text across skill names and descriptions.
Get full metadata for a specific skill including evidence sources, input/output schemas, and chaining information.
List all available education skills grouped by domain. Returns skill ID, name, evidence strength, tags, and estimated teacher time.
Describe what you're trying to do in plain English and get 3-5 relevant skill recommendations. The entry point for users who don't know what skills exist.
Output schemas not formally documented. All tools return unstructured text responses wrapped in { content: [{ type: 'text'; text: string }] }. LLMs cannot plan downstream calls or parse structured fields (e.g., skill_id, domain) reliably. Responses need JSON-formatted output schemas that document expected fields, types, and relationships.
Parameter constraints underspecified. 'domain' and 'evidence_strength' parameters lack enum declarations or valid-value documentation. The 'query' parameter in find_skills offers no guidance on search syntax (exact vs. substring, case sensitivity). LLMs will hallucinate invalid values.
Error handling provides no recovery guidance. When a skill is not found (get_skill_details) or no matches returned (find_skills, suggest_skills), responses are bare error strings. No suggestion of alternatives (e.g., 'Did you mean: skill-id-1, skill-id-2?'). LLM has no path forward.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | <=2025-11-25 | v2 |
No pagination or result-limiting mechanism. list_skills and find_skills return all matching skills without limit or cursor. If 500+ skills exist, responses bloat context window and degrade LLM reasoning. Rubric baseline: cap at 20-50 and offer pagination.
Tool descriptions are adequate but generic. 'List all available education skills grouped by domain' does not explain WHEN to call this vs. find_skills or suggest_skills. Missing guidance on tool selection and prerequisites. Rubric baseline: descriptions should be 50-200 chars and answer WHAT, WHEN, and WHAT IT RETURNS.