A practice repository containing multiple MCP server implementations: a research paper server using arXiv integration, a leave management server, and multi-server chatbot clients using Anthropic Claude.
This is a learning/practice repository with significant quality gaps. Of 5 tools, all have basic descriptions (20-100 chars), but critical issues emerge: (1) schemas are inferred from fragments rather than explicitly visible in tool registration code; (2) descriptions lack context on WHEN to use tools and prerequisites; (3) no parameter descriptions beyond name; (4) no error handling guidance; (5) no output schema documentation; (6) missing tool composition planning (e.g., search_papers stores data but extract_info retrieves it, relationship unexplained). The codebase shows client examples (mcp_chatbot_for_multi_server.py) but the actual server implementations (chatbot_example.py, mcp_server_by_youtube/main.py) are not fully included, only fragments are visible. This means tool definitions cannot be fully audited.
Apply leave for specific dates (e.g., ["2025-04-17", "2025-05-01"])
Search for information about a specific paper across all topic directories.
Check how many leave days are left for the employee
Get leave history for the employee
Search for papers on arXiv based on a topic and store their information.
Tool implementations not fully visible in source. Schema definitions inferred rather than explicitly shown. Cannot verify input/output contract.
No parameter descriptions provided. Every parameter has only a name and type, LLMs cannot infer valid formats, ranges, or semantics. E.g., is employee_id a UUID, email, or numeric ID? What format is paper_id (arXiv:2401.12345 or just 2401.12345)?
No output schemas documented. Agents cannot know what fields to expect or how to chain tools. E.g., does search_papers return paper_id? Does get_leave_history return a list or single object?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 33 | - | v1 |
No error handling or recovery guidance. What happens if apply_leave fails due to insufficient balance or conflicting dates? How should the agent respond?
Tool relationships unexplained. search_papers 'stores' data; extract_info 'searches across directories'. Are these the same data? Must search_papers run before extract_info? Dependency undocumented.
No constraints on numeric parameters. max_results defaults to 5, is there a maximum? Can an LLM request 10000 results?
apply_leave signals WRITE risk but lacks confirmation/dry-run pattern. Destructive operations (deleting leave, modifying historical records) should support confirmation.
Vague description in extract_info: 'Search for information... across all topic directories' contradicts the verb 'extract'. Does this search or extract? When should users prefer this over search_papers?