Collection of MCP servers demonstrating tool creation across multiple days: docs search (Day04, Day08, Day09), filesystem operations (Day08), web search (Day08, Day09), and database query tools (Day10)
This is a collection of 22 tools across a multi-day tutorial project ('10-days-of-agents'). While many tools have descriptions and basic schemas, the quality is inconsistent. Critical issues: (1) Tool name duplication and lack of uniqueness, 'calculator' appears twice (Day01, inferred day elsewhere), 'search_local_docs' appears 4 times with nearly identical signatures, 'file_write_safe' appears 3 times, 'web_search' appears 2 times. This violates the single-responsibility and naming-clarity patterns; LLMs cannot distinguish between functionally identical tools. (2) Descriptions are inconsistent in length and depth, some are detailed (e.g., calculator at ~80 chars, local_search at ~100 chars), others are minimal (e.g., 'echo' is vague, query is underspecified). The rubric baseline is 194 chars average; most tools here fall short. (3) Schemas are present but sparse, many tools expose only 1-2 parameters with minimal constraints. No enums, no min/max bounds on numeric params, no validation rules. (4) Output schemas are documented in descriptions (as return types) but not formally in the tool registration. (5) Error handling is informal, tools return {'ok': True/False, 'error': str} but provide no recovery guidance or actionable error messages. (6) Security: file_write_safe and run_shell_safe do implement sandboxing (commendable), but run_shell_safe is overly permissive (echo, ls only, barely useful) and lacks timeout enforcement visibility. (7) Composition issues: the 22 tools span 10 separate 'day' directories with duplicate functionality (search_local_docs x4, file_write_safe x3), suggesting this is not a cohesive server but a collection of tutorial examples. This is educational material, not production-grade tooling.
Evaluate safe math expressions with support for percentages. Returns {"ok": True, "result": float} on success or {"ok": False, "error": str} on failure
Safely evaluate arithmetic expressions (supports %, ( ), + - * /)
Get the schema (column names and types) for a specific table
Load files from data/, chunk, embed, and persist to Chroma. Returns simple stats about indexing: files_indexed, chunks_added, files list, index_dir, collection name
Retrieve top-k chunks relevant to the question with citations. Returns results with source, page, snippet, and score
Echo a message, with optional dry_run mode
Tool name duplication: 'calculator' appears 2x, 'search_local_docs' appears 4x, 'file_write_safe' appears 3x, 'web_search' appears 2x. LLMs cannot distinguish identical tool names across the server; they will select ambiguously or fail.
Descriptions lack depth and consistency. Rubric baseline is 194 chars average. Tools like 'echo' (40 chars), 'query' (55 chars), and many 'search_local_docs' variants (50 - 55 chars) fall significantly short. Descriptions do not explain WHEN to use the tool, what prerequisites exist, or what distinguishes it from similar tools.
Input schemas lack constraints. No enums declared for known-value params (e.g., 'mode' in file operations could be read/write/append; should be enum). No min/max bounds on numeric params (top_k, timeout_s). No format/pattern constraints on string params. Rubric requires explicit constraints to prevent LLM hallucination.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 49 | - | v1 |
Safely read UTF-8 text from a file under sandbox (out/ directory). Returns file content
Safely write UTF-8 text to a file under 'root' (allowlist directory). Returns the RELATIVE path that was written
Safely write UTF-8 text to a file under sandbox (out/ directory). Returns path that was written relative to out/
Write UTF-8 text under Day05/out only; size-limited and sandboxed. Returns {"ok": True, "bytes": <int>} or {"ok": False, "error": <str>}
Create a GitHub issue. Returns issue number + url. Respects dry_run and retries on transient errors
List all table names in the DuckDB database
Search ./data/*.md for query terms and return filenames + snippets. Returns {"ok": True, "results": [{"file": "...", "snippet": "..."}]} or {"ok": False, "error": "..."}
Execute a safe read-only SQL query against DuckDB. Only read-only queries are allowed
Run whitelisted shell commands: echo, ls. Reject anything else. Returns command exit code and output (trimmed to 2000 chars)
Get a complete summary of all tables and columns in the database
Search ./data for a keyword and return top snippets
Search ./data for a keyword and return top snippets. Returns {"hits": [{"path": "...", "snippet": "..."}]}
Keyword search in Day05/data; returns title, snippet, path, score
Search ./data for a keyword and return top snippets. Returns {"hits": [{"path": "...", "snippet": "..."}]}
Search the web using Tavily and return top results (title + URL)
Search the web and return top results. Uses stub database for demo (mcp topic returns 3 hardcoded results)
Output schemas not formally documented in tool registration. Tools return structured objects (e.g., {'ok': bool, 'result': float, 'error': str}), but these schemas are described only in text within descriptions, not as formal output schema declarations. LLMs cannot parse these descriptions reliably.
Error handling lacks recovery guidance. Tools return {'ok': False, 'error': str} but errors are generic (e.g., 'Invalid input length') without actionable next steps. Rubric requires: suggest alternative tools, clarify constraints, or indicate retry eligibility.
'echo' tool is vaguely named and described. Description 'Echo a message, with optional dry_run mode' does not explain why an LLM would call this. Is it for logging? Testing? Confirmation? The name 'echo' is generic and does not follow verb_noun convention (e.g., 'log_message', 'test_output').
'run_shell_safe' is overly permissive for two commands (echo, ls) but the description does not explain why these are safe or what the LLM should use them for. Expected use cases are unclear.
'query' tool lacks input validation and error guidance. Description says 'Only read-only queries are allowed' but does not specify how the server enforces this, what error is returned if a mutation is detected, or how the LLM should respond.
Parameter naming inconsistency. 'github.create_issue' uses 'owner' and 'repo' (clear), but many search tools use 'query' vs 'searchString' vs 'search_query' (3 different names for the same concept). Rubric requires consistent naming to avoid LLM confusion.