Multi-agent MCP server ecosystem with specialized agents for financial analysis, web search, database queries, travel planning, and data ETL operations
This multi-server suite contains 17 tools with highly inconsistent quality. ETL tools (1-8) have explicit schemas and basic descriptions but lack production-grade error handling and LLM-optimization. Agent proxy tools (9-17) have minimal input schemas, trivial descriptions, and hide complexity behind CrewAI agents rather than exposing composable operations. No tool demonstrates proper error recovery guidance, permission gating, or audit trail. Most descriptions are under 50 characters and lack context for LLM tool selection. Output schemas are undocumented for most tools.
Search the web and scrape relevant content using Brave Search MCP and a CrewAI-powered agent.
Ensures each column in data has the correct type according to the provided type mapping.
Analyze context7 to retrieve any information about any documentation using CrewAI-powered agent.
Checks for anomalies in the data based on provided rules (e.g., min/max for columns).
Proxy a user question into your Docker-based MCP server via CrewAI.
Enforces constraints such as not-null and unique on columns.
Analyze github repositories data using CrewAI-powered agent.
Nine 'analyst' tools (multi_analyst, brave_web_search, context7_analyst, docker_mcp_tool, github_analyst, selenium_scraper_tool, supabase_analyst, yfinance_analyst) accept only a generic 'question' string parameter with minimal schema detail. They hide all internal complexity (agent routing, tool composition, API calls) behind natural language. This violates the Single Responsibility Pattern, agents cannot decompose multi-step workflows or handle partial failures. No way to request specific output formats, pagination, or fallback behavior.
Descriptions for 'analyst' tools are trivial (20 chars or less: 'Analyze X to retrieve any information', 'Search the web and scrape'). These provide no context for WHEN to use the tool, HOW it differs from similar tools, WHAT prerequisites exist, or WHAT the output structure looks like. LLMs cannot reliably select between brave_web_search and context7_analyst based on these descriptions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 7 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Handles missing values according to provided strategy.
Handle financial and DB questions using unified tool access.
Reads a CSV file from the given path and returns its contents as a list of dicts.
Removes duplicate rows in the data.
Use Selenium MCP to scrape structured data from websites based on navigation instructions.
Applies string standardization rules per column (e.g., lowercase, strip).
Analyze supabase tables and answer questions about out data using CrewAI-powered agent.
Applies transformations to the data (e.g., rename columns, add new columns).
Orchestrates travel planning using multiple specialized agents for flights, accommodations, and local experiences.
Analyze yfinance library and answer questions about out data using CrewAI-powered agent.
No tool documents its return type or output schema. ETL tools do not declare what fields users should expect in the response dict (e.g., does read_csv_file return {'data': [...], 'columns': [...], 'errors': []}?). Agent proxy tools return unstructured CrewAI Task results that LLMs must parse as free-text. This makes downstream tool chaining impossible and forces LLMs to guess at field names.
ETL tools (check_data_types, detect_and_report_anomalies, etc.) document that 'Agent must ask user for parameters if not provided', but this is LLM-hostile. A tool requiring complex structured params (type_mapping dict, anomaly_rules dict) should fail fast with a clear error message if params are missing, not rely on the LLM to infer the structure and ask the user. No recovery guidance is provided, errors like 'Column cannot be converted to int' do not suggest corrective actions.
transform_data tool uses df.eval() on user-supplied expressions without sanitization. The docstring contains 'WARNING: eval can be unsafe if the string comes from user input!' but provides no mitigation. This enables code injection attacks. An agent tricked via prompt injection could execute arbitrary Python code on the server.
No tool validates inputs before processing. ETL tools accept 'data' (List[dict]) but do not validate that rows are dicts, that column names match type_mapping, or that data is not empty. Invalid input causes cryptic pandas errors that LLMs cannot recover from. No min/max bounds on array sizes (could accept millions of rows and exhaust memory).
Agent proxy tools (multi_analyst, brave_web_search, etc.) accept user_id or no user context at all. No permission gating is visible, agents have full access to supabase_analyst, github_analyst, and docker_mcp_tool regardless of user identity. No audit trail (who called what, when, with what result). Destructive tools like docker_mcp_tool (WRITE risk) have no confirmation step or dry-run mode.
travel_planner tool accepts structured input_data with nested properties (departure, destination, start_date, end_date, num_travelers, attractions, accommodation_type) but provides no enum constraints or format validation. An LLM could pass invalid dates, negative traveler counts, or unrecognized accommodation types. No error guidance if the travel planning fails (e.g., 'No flights found').
ETL tool names start with action verbs, which is good (read, check, detect, remove, handle, standardize, enforce, transform), but agent proxy tools use vague names like 'multi_analyst', 'context7_analyst', 'docker_mcp_tool'. These names do not convey what action happens, is it fetching data, creating something, or analyzing? LLMs struggle to differentiate tools when names are ambiguous.