A multi-module educational framework demonstrating Agent, LLM, LangChain, LangGraph, RAG, MCP server implementations, and OpenAI Agents SDK with practical examples
This is a tutorial/educational repository, not a production MCP server. Tool definitions are scattered across multiple files with severe quality issues: 5 duplicate tool names (get_weather appears 5 times with different schemas), missing or trivial descriptions, incomplete parameter documentation, and no structured output schemas. The repository demonstrates LangChain/LangGraph patterns but does not constitute a cohesive, deployable MCP server. Most tools lack parameter type specifications in visible code, and error handling is absent. Per the rubric baseline, median community servers score 45-55; this one falls well below due to scattered definitions, duplication, and lack of production-grade structure.
查询座位余量
创建新订单,自动重试保障成功率
执行退款逻辑
模拟获得用户名字
模拟获取天气
模拟获得天气信息
查询指定城市的实时天气。如果是此时此刻的天气请求,调用此工具。
模拟获得天气信息
Critical: Five instances of 'get_weather' tool with different input schemas (location vs city parameter names). This creates ambiguity for LLM tool selection and violates the single-responsibility principle. Duplicate tool names force LLMs to guess which variant to call.
High: execute_refund and check_seat have empty input schemas ({}) with minimal descriptions (under 20 chars when accounting for substance). No parameter documentation exists to guide LLM calls.
High: Output schemas are not documented for any tool. Callers do not know what structure to expect. LLMs need to know what fields to expect so they can plan downstream tool calls and extract the right data.'
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 37 | - | v1 |
查询指定城市的实时天气。如果是此时此刻的天气请求,调用此工具。
查询《战争与和平》小说中的内容,包括人物、情节、历史事件等
模拟发送邮件
High: Parameter descriptions are inconsistent or missing. 'location' vs 'city' naming for get_weather variants makes tool chaining impossible. Tools like create_order accept 'query' but do not specify format, length limits, or valid values.
Medium: No error handling guidance. Tools return strings like 'ERROR: 服务暂时不可用' but do not categorize errors as retryable, user-fixable, or fatal. LLMs cannot determine if they should retry, ask the user, or give up.
Medium: No pagination or result limits documented. search_war_and_peace and other retrieval tools do not specify max results or how to handle large datasets.
Low: Descriptions in Chinese. While culturally appropriate, descriptions should be in English for international LLM compatibility. Most production LLMs default to English system prompts, risking misinterpretation.