Multi-purpose MCP server suite providing tools for arXiv paper search, AWS cost/log/storage analysis, clinical trials, ChEMBL chemistry, PubMed research, code execution, image generation, and RAG capabilities
Scoring was not performed
Tool naming anti-pattern: 'search_clinical_trials_and_save_studies_to_csv' combines two distinct responsibilities (search + file I/O). Violates single-responsibility principle. Should split into 'search_clinical_trials' and 'save_results_to_csv'.
Typo in tool name: 'compount_activity' should be 'compound_activity'. This orthographic error forces LLMs to spell-correct or risk tool selection failure.
Parameter descriptions missing or trivial for 60+ parameters across the server. Examples: 'days' in get_daily_cost lacks bounds (should be 1-365 or similar); 'region' lacks validation hint; 'filename' in save tools lacks path safety guidance. LLMs cannot infer constraints from names alone.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 26 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Output schemas not documented. No tool describes its return type structure. Examples: 'search_papers' returns a list but no schema for paper object fields is visible; 'get_logs' returns logs but structure undefined. LLMs must guess at response structure, causing downstream tool chaining failures.
Error handling absent or minimal. No tool provides recovery guidance. Example: 'download_paper' and 'read_paper' return {'error': 'Paper not found'} with no suggestion to try 'search_papers' first. LLMs have no way to self-correct or plan recovery.
Code execution tools (repl_coder, repl_drawer) lack confirmation/dry-run capability. These tools can execute arbitrary Python and have IRREVERSIBLE side effects (e.g., network calls, data corruption). No safeguards, rate limiting, or execution timeout visible in source.
File I/O tools (save_to_csv, load_csv_data) lack path validation. No visible guards against path traversal attacks. 'filename' parameter is passed directly to filesystem operations. LLM could be tricked into writing outside intended directory.
Pagination and result limits missing/unclear. 'search_papers' has max_results=10 default but no upper bound documented. 'pubmed_search' defaults to 10 but no guidance on token cost of large result sets. Response size could exhaust context windows.
Parameter type ambiguity. 'region' appears in 10+ tools with description 'AWS region' but no enum constraint or validation hint. LLMs can pass invalid values like 'US' or 'west' instead of 'us-west-2'. Should use enum constraint.
Missing field chaining references. Example: 'search_papers' returns paper_id but no subsequent tool accepts paper_id directly, 'download_paper' expects 'paper_id' (consistent), but output lacks 'resource_uri' consistently. Broken chains force discovery detours.