Educational repository containing 12 different AI agent implementations, including simple chatbots, tool-calling agents, RAG agents, web-browsing agents, code-generation agents, memory agents, and an MCP server implementation. Not a single MCP server but a collection of agent examples.
This educational repository demonstrates basic tool patterns but falls well short of production-grade quality. While tool names generally follow verb_noun convention and schemas are present with JSON Schema format, there are critical gaps: (1) Parameter descriptions are minimal or missing entirely across all tools. (2) Output schemas are not documented, tools return strings or dicts with no declared structure for LLM consumption. (3) Error handling is absent, no guidance on recovery or categorization. (4) Multiple tools with identical names (get_weather appears in tools.py and mcp_server.py; calculator/calculate naming inconsistency) indicate fragmentation. (5) Tools return mock/fake data with hardcoded responses, limiting real-world applicability. (6) No error recovery patterns, no input validation guidance, no security declarations. The repository is instructional (showing how to build agents) rather than production-ready (defining professional-grade tools).
Evaluate a math expression safely.
Safely evaluate a math expression, e.g. '2 + 3 * (4 - 1)'.
Extract absolute links from a page.
Extract plain text from a page, limited for model context.
Fetch page title + text content from URL.
Get the current local date and time.
Get current local server time.
Get mock weather for a city.
Parameter descriptions are missing or trivial across all tools. Examples: 'city' has only 'City name' (9 chars), 'query' has only 'Search query' (12 chars), 'expression' has only 'Math expression' (15 chars). LLMs cannot infer context from names alone, these descriptions violate pattern:tool-description.
Output schemas are not documented. Tools return unstructured responses (strings like 'calculator error: {exc}', dicts with arbitrary fields) with no declared structure for LLM consumption. Agents cannot plan downstream calls or extract fields reliably. Violates pattern:response-shaper.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Get mock weather data for a city.
Search Google and return top organic results parsed from HTML.
Return 3 mock web search results for a query.
Duplicate tool definitions with inconsistent names across modules: 'get_weather' appears in both 02-tool-calling-agent/tools.py and 07-mcp-agent/mcp_server.py with slightly different schemas (one has description 'Get mock weather data for a city', other 'Get mock weather for a city'). 'calculator' vs 'calculate' naming in different modules. This fragmentation violates pattern:tool, single canonical definitions required.
Error handling is absent across all tools. When calculator encounters an invalid expression, it returns 'Calculator error: {exc}', no guidance on recovery, no categorization (retryable vs fatal), no suggestion for next steps. Violates pattern:recovery-guide and pattern:error-classification.
No input validation constraints documented. Parameters like 'city' and 'query' are free-form strings with no enum, format, or length constraints. LLMs hallucinate invalid values with no guidance. Violates pattern:constrained-input.
Mock/hardcoded responses limit real-world applicability. get_weather returns deterministic fake data (samples = [('sunny', 25), ('cloudy', 19), ...]). search_web returns template results with no real search. These are pedagogical, not production tools. Violates idempotence and real-world reliability expectations.
No tool composition support. Tools like search_google, fetch_page, extract_links, extract_text appear in separate agents but are not chained together (no 'search, fetch top result, extract links' workflow). Pattern:tool-chain violated.
Tool descriptions are below rubric baseline (194 chars avg). Examples: 'Get mock weather data for a city.' (31 chars), 'Return 3 mock web search results for a query.' (45 chars), 'Get the current local date and time.' (35 chars). These provide minimal context for LLM selection. Descriptions should explain WHAT, WHEN, and any prerequisites.