Multi-tool MCP server suite providing math solving, weather, translation, web search, email, Spotify recommendations, GitHub integration, and research capabilities via FastMCP
This server exhibits significant quality gaps across naming, descriptions, schema completeness, and error handling. While some tools have basic structure, most lack proper LLM-optimized descriptions, comprehensive input schemas, and actionable error guidance. The server mixes read-only tools (arxiv search, math solving) with high-risk write operations (send_email, GitHub issue creation, code execution) without clear permission gates or warnings. Tool names are inconsistent in style, some use action verbs (solve_math, send_email) while others use nouns (github_tool, code_executor). Input schemas exist but are sparse in constraint documentation. Output schemas are largely undocumented, forcing LLMs to guess the structure of returned data. No security hardening visible for credential injection or sensitive parameter validation.
Search Internet Archive for digitized books, historical texts, web archives, audio/video recordings, and other non-academic content. NOTE: For academic research papers (AI, ML, physics, CS, math), prefer arxiv_research_search instead — Archive.org does not reliably index them.
Search arXiv for academic papers (ML, AI, physics, math, CS, etc.) and return an LLM-summarized research overview. Use this tool for any query about AI / ML papers, Physics, mathematics, computer science research, Named papers or authors in academic contexts.
Sandbox: run Python code safely, capture stdout/stderr/errors
AI-style structured code generation with explanations
Integrated: fetch code from GitHub → execute → analyse mistakes/improvements
github_tool combines 8 distinct actions (search_repos, get_readme, read_file, list_issues, etc.) into a single tool. LLMs cannot reason effectively about such multi-responsibility tools; they require separate, focused tools for each action.
code_executor tool accepts arbitrary Python code with minimal validation and no timeout enforcement visible in description. This is a severe security risk, agents could execute malicious code, access file systems, or exhaust resources. No permission gates or sandbox verification in the schema.
send_email and read_emails expose Gmail integration without explicit mention of credential handling or permission scope. No documentation of what scopes (read:email, write:email, send:email) are required or how credentials are injected. Potential for token leakage in logs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | - | v1 |
Interact with GitHub repositories. Actions: search_repos — Search GitHub for repositories matching `query`; get_readme — Fetch the README of `owner/repo`; read_file — Read a specific file at `path` in `owner/repo` (on `branch`); list_issues — List open issues in `owner/repo`; get_issue — Get a single issue by `issue_number` in `owner/repo`; create_issue — Create a new issue in `owner/repo` with `issue_title` and `issue_body`; list_files — List files/folders at `path` in `owner/repo` (on `branch`); repo_info — Get metadata about `owner/repo` (stars, language, description, etc.)
Reads emails from Gmail based on a search query.
Sends an email using the user's Gmail account.
Comprehensive math solver. Give it any math problem in natural language or expression form. Handles: derivatives, integrals, equation solving, simplification, evaluation.
Get an archived Wayback Machine snapshot of a URL, with an LLM summary.
No output schemas documented for any tool. LLMs cannot predict the structure of returned data (e.g., what fields does arxiv_research_search return? Is it a list, an object with a results array?). This forces LLMs to make assumptions and fail on unexpected structure.
code_executor and code_writer descriptions are vague and underdescriptive. 'Sandbox: run Python code safely' (40 chars) does not explain when to use it, what safety guarantees exist, or what happens on error.
github_tool parameter 'action' is a free-form string with no enum constraint listed in the schema. The description mentions 8 valid values (search_repos, get_readme, read_file, etc.), but LLMs will hallucinate other values like 'fork_repo' or 'delete_repo' which are not supported.
arxiv_research_search sort_by parameter is a free-form string. The description says valid values are 'relevance', 'lastUpdatedDate', 'submittedDate', but these should be declared as an enum in the schema to prevent hallucinations.
archive_research_search mediatype parameter is under-constrained. Description mentions 7 valid values, but no enum in schema. LLMs will invent values like 'document', 'manuscript', or 'paper' which are not supported.
No error classification or recovery guidance visible in any tool description. If arxiv_research_search times out, returns no results, or hits a rate limit, the description does not tell the LLM whether to retry, ask the user, or skip the tool.
read_emails accepts a 'query' parameter (free-form string) with no format or syntax documentation. Does it support Gmail search syntax (is:unread from:alice@example.com)? Boolean operators? Date ranges? Without this, LLMs will pass malformed queries and get failures.
code_executor timeout parameter description says 'Execution timeout in seconds' but does not specify minimum, maximum, or default values. LLMs may pass 0 (instant timeout), negative numbers, or 999999 (hang indefinitely).