Multi-tool MCP server suite with Amazon scraper, Gmail integration, Neo4j graph access, weather, math, and Twitter analysis tools
This server has significant structural and definition quality issues. Out of 17 tools, nearly all lack proper parameter type definitions visible in the schema, descriptions are minimal or missing structured context, and output schemas are entirely undocumented. Tools are spread across multiple independent modules without coordination. The rubric baseline for A-grade tools requires ALL params to have types AND descriptions, documented return schemas, and error handling guidance, this server has almost none of that. The tool definitions appear to rely on LangChain StructuredTool introspection (via tool_def_maker.py), but the actual schema properties are not visible in the provided code. Several tools have only trivial one-line descriptions (e.g., 'Find the current date.' for get_current_date). Parameter descriptions are present for some tools (e.g., wikiSearch, scrape_product) but missing entirely for others (get_date_time has no parameters anyway, but get_current_date is equally sparse). No output schemas are documented anywhere. Error handling is absent, no recovery guidance, no retryable vs fatal classification. This lands in the F - D range; the presence of descriptions for a few tools and basic naming (mostly verb-noun) saves it from a complete F.
Add two numbers
Find the current date.
Displays the current date and time in this specific format: Day(day suffix) Month Year, Hour:Minute AM/PM
Get nodes for a label with optional exact-match properties. If match is empty or null, returns all nodes with that label.
Fetch current weather + short forecast for a city. Returns a compact dict.
Returns whether OAuth is configured and token is present/valid.
Search and list messages. Returns lightweight cards: id, threadId, headers, snippet.
Output schemas completely undocumented. No tool documents what it returns, field types, or structure. LLMs cannot plan downstream tool calls or extract required data.
Multiple tools with trivial one-line descriptions under 20 characters (get_date_time, get_current_date, gmail_auth_status, list_labels, list_relationship_types). These fail the rubric minimum, descriptions must explain WHAT, WHEN to use, and any prerequisites. LLMs cannot infer tool purpose from a bare label.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Read a single message by id. Returns headers + best-effort text body.
Returns count of unread messages in INBOX.
Initialize and verify the Neo4j connection. Use this once to check that the MCP server can reach the database.
List all node labels in the database.
List all relationship types in the database.
Multiply two numbers
Run any Cypher query. Returns a list of record dicts.
Scrape product information from Amazon product URL including name, price, rating, reviews, availability, and description
Search for products on Amazon and return a formatted list of results with name, price, rating, image URL, and product URL
Searches Wikipedia for the given query and returns a summary of the most relevant article.
Tool 'mutiple' has a typo in the name (should be 'multiply'). Tool names must be unambiguous and correctly spelled, LLMs treat names as ground truth for function lookup.
Parameter descriptions mostly absent or minimal. Many tools (e.g., get_date_time, get_current_date, gmail_auth_status, init_neo4j, list_labels) have no parameters described, or params are bare names without context. The rubric baseline: 100% of A+ tools have parameter descriptions.
No error handling, recovery guidance, or error classification visible. If a Gmail auth fails, scraping fails, or Neo4j is unreachable, tools return nothing actionable. Agents cannot self-correct or understand next steps.
No tool composition guidance. For example, gmail_list and gmail_read are separate tools but no documentation explains the expected workflow: call gmail_list to get message IDs, then gmail_read with a message_id. The gmail_list description does not mention the field name that gmail_read expects.
Gmail tools may expose auth tokens or session metadata in responses. No evidence of stripping sensitive fields before returning to the agent. If a token appears in tool output, it enters the LLM context and risks being logged or echoed.
No pagination or result-limiting strategy documented. search_products accepts max_results but no guidance on default, minimum, or maximum values. gmail_list has a max_results with default=10 but no documentation of the result cap or next_cursor/pagination field. Agents can request huge datasets that blow context windows.
Destructive tool (run_cypher, with 'write' risk classification) has no confirmation step or dry-run mode. An agent can DELETE or UPDATE the entire Neo4j database without a safety gate. Agents are error-prone and this tool should require explicit confirmation or dry-run validation.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in the schema. While run_cypher is marked 'write' in the risk field, there is no formal destructiveHint annotation. Idempotency of tools like 'add' (math) vs 'search_products' (API call with side effects) is not declared.