A comprehensive MCP application platform providing RAG capabilities, AWS integrations, arXiv search, and various data retrieval tools through multiple MCP servers
This MCP server has significant quality gaps across naming, descriptions, schema completeness, and error handling. While tool names generally follow verb conventions (search_papers, list_papers, read_paper), several critical issues emerge: (1) Parameter descriptions are inconsistent, some tools like 'unified_email_search' have detailed param descriptions, while 'aws_cli' exposes dangerous parameters without proper validation guidance. (2) Output schemas are entirely undocumented, no return type specifications visible for any of the 11 tools, violating the baseline that 100% of A+ tools document return types. (3) Tool composition violates single-responsibility: 'outlook_database_query' is a raw SQL executor with no safeguards, and 'aws_cli' bundles all AWS operations into one tool instead of domain-specific wrappers. (4) The 'aws_cli' tool is a critical security and usability anti-pattern, it accepts arbitrary AWS service + operation + parameters, inviting SQL injection-like prompt injection attacks and requiring users to know AWS API surface. (5) No evidence of error classification, recovery guidance, or idempotency documentation. (6) Natural-language identifiers are absent, email tools require 'account' as a string but don't hint at fuzzy matching or validation. Baseline for average params per tool is 4; most tools here meet that, but parameter quality (types + descriptions) is sparse. Descriptions for tools are present but brief (40-80 chars), below the 194-char production average. Overall, the server reads as a rapid integration of multiple existing APIs without proper LLM-facing abstraction.
Execute AWS service operations using boto3 with comprehensive error handling and validation.
Download a paper from arXiv.
Get email volume statistics grouped by time period.
Get statistics about email folders.
Get a list of the latest papers in a specific category.
Get a comprehensive overview of the mailbox.
Execute a custom read-only SQL query against the Outlook database.
aws_cli tool violates single-responsibility and security best practices. Accepts arbitrary 'service_name', 'operation_name', and 'parameters' as free-form inputs. This invites prompt injection attacks and prevents meaningful LLM reasoning about AWS operations. No input validation, no permission gating, no operation-specific schemas.
outlook_database_query exposes raw SQL interface. Despite claiming 'SELECT only', there is no visible input validation, no SQL parsing to enforce read-only semantics, and no error handling for malformed queries. This is a major injection vector.
No output schemas documented for any of the 11 tools. The rubric baseline states 100% of A+ tools document return types. Without documented schemas, LLMs cannot plan downstream tool calls or extract specific fields, forcing them to guess or parse unstructured responses.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Read the content of an arXiv paper.
Search for papers on arXiv with advanced filtering.
Get comprehensive statistics about email senders including per-folder breakdowns.
Unified search tool for finding emails with advanced filtering.
Parameter descriptions are inconsistent and often lack actionable format/constraint guidance. E.g., 'date_filter' accepts strings like 'today', 'this week', 'last 30 days', '2025-06-01..2025-06-30' but the description does not enumerate valid formats or regex patterns. LLMs will guess.
No error classification or recovery guidance visible. Tools do not document what happens on failure (retryable? user-fixable? fatal?). No examples of actionable error messages like 'User not found. Try search_users() with a partial name.'
Tool composition violates single-responsibility principle. 'aws_cli' bundles all AWS services and operations. Should split into domain-specific tools: 's3_list_buckets', 's3_upload_file', 'ec2_list_instances', etc. Monolithic tools prevent focused LLM reasoning and increase hallucination risk.
No evidence of idempotency semantics. Tools like 'download_paper' and 'read_paper' could have side effects (caching, rate limits). No documentation on whether repeated calls with same input produce the same result or side effects. Agents will retry on ambiguous failures, non-idempotent tools risk duplicate operations.
Email tools accept 'account' as a bare email string with no validation guidance. Description does not indicate whether fuzzy matching is supported or if the exact email must be in the system. Users operating via chat say 'my Gmail account' or 'the shared mailbox', tool does not abstract this lookup.
No pagination documentation for list tools. 'unified_email_search' has 'limit' and 'offset' but no mention of total count, next_cursor, or has_more indicator. 'list_papers' has 'max_results' but no guidance on pagination. Without explicit pagination, large result sets blow context windows.