Learning MCP by building a simple server with multiple MCP implementations (Air Fryer, Calculator, Gmail, Google Drive) and a Flask chat backend
Mixed quality across 18 tools. Naming is generally action-verb-based and clear (cook, add, sum, multiply, divide, list_messages, search_messages, read_message, create_draft, etc.), following the verb_noun pattern well. Descriptions are present for all tools but vary in quality, most are 20-60 characters, which is on the lower end of the ideal 50-200 char range. All visible tools have input schemas with proper types and parameter descriptions. However, critical gaps exist: (1) no output schemas are documented, LLMs cannot plan downstream calls without knowing what fields to expect; (2) no error handling guidance in any tool, no recovery messages or classification; (3) no pagination support on list/search tools despite potentially large result sets (list_messages, search_messages, list_files, search_files, list_folders, recent_files); (4) no tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear semantics (create_draft is WRITE, others are READ_ONLY); (5) create_draft tool lacks confirmation/dry-run despite being a state-mutating operation. The calculator tools (add, sum, sum_many, multiply, divide, explain_calculation) are simple and well-formed but extremely generic, no domain-specific context. Gmail and Google Drive tools are more substantial but lack pagination and clear error scenarios.
Add two numbers together.
Cook food in the air fryer for a specified time.
Create an email draft (does not send).
Divide one number by another.
Get help explaining a mathematical calculation step by step.
Get detailed information about a specific file.
Get the count of unread messages in the inbox.
No output schemas documented. LLMs cannot determine what fields to expect from tool responses, blocking downstream composition and forcing discovery calls.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Tools are categorized by risk metadata, but this is not reflected in tool definitions, preventing agents from making safe retry decisions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 51 | - | v1 |
List files in Google Drive.
List folders in Google Drive.
List messages in Gmail inbox.
Multiply two numbers together.
Read the content of a file from Google Drive.
Read the full content of a specific message.
Get recently modified files.
Search for files in Google Drive.
Search for messages in Gmail.
Add two numbers together.
Add multiple numbers together.
List and search tools lack pagination. list_messages, search_messages, list_files, search_files, list_folders, recent_files have no offset/limit or cursor support, risking context window exhaustion if many results exist.
No error handling guidance. Tools return error messages (e.g., cook validates time_seconds > 0) but lack actionable recovery guidance. LLMs cannot determine whether to retry, ask the user, or abandon the goal.
create_draft is a state-mutating tool with no confirmation or dry-run. Agents can create unintended email drafts without warning. No recovery mechanism offered.
Descriptions are too brief (mostly 20-60 chars, below ideal 50-200 range). Many lack context for when/why to call the tool. E.g., 'Read the full content of a specific message' lacks context on when to use read_message vs list_messages.
No chaining metadata in responses. Tools like search_files or search_messages do not document what IDs/references they return, preventing LLMs from composing follow-up calls (e.g., search → read).