Collection of data analysis agent implementations integrating with Keboola MCP Server using CrewAI, LangChain, and DSPy frameworks
This server defines 6 tools with moderate naming but significant gaps in schema documentation, parameter validation, and structured output. All tools accept a single 'query' parameter as a string rather than typed, structured schemas. While descriptions exist (ranging 72-182 chars), they lack specificity about input formats, error handling, and output structure. The tools are implemented as LangChain BaseTool subclasses but the MCP registration layer is not visible in the provided source, suggesting tools are inferred rather than explicitly registered with MCP schemas. Error handling is basic (try/catch with string returns) and lacks recovery guidance. No tool annotations, no pagination, no documented output schemas.
Perform basic data analysis on datasets. Input should be JSON with 'data' (list of dicts) and 'operation' (describe, correlate, summarize).
Create data visualizations. Input should be JSON with 'data' (list of dicts), 'chart_type' (bar, line, scatter, pie), and 'title'.
Fetch financial data for stocks, ETFs, or other securities. Input should be a ticker symbol (e.g., 'AAPL', 'GOOGL').
Scrape data from web pages. Input should be a URL to scrape basic text content from.
Search the web using DuckDuckGo for information
Search Wikipedia for information
All 6 tools accept a single 'query' parameter typed as string with no structured schema. No parameter type definitions, ranges, enums, or validation rules visible. Input schema is effectively absent for all tools.
No documented output schemas for any tool. Return values are JSON strings (via json.dumps()) or plain error strings, but LLMs are given no specification of expected fields, types, or structure. Agents cannot plan downstream calls.
data_visualization and data_analysis require complex JSON payloads as the query string (e.g., '{"data": [...], "chart_type": "bar"}'), but no schema validation, format documentation, or enum constraints are provided. LLMs must construct JSON manually, prone to malformed input.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 29 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 21 | - | v1 |
Error handling returns plain strings ('Error: No data provided') with no recovery guidance or classification. LLMs cannot determine if an error is retryable, user-fixable, or fatal. No guidance on next steps.
Tool registration not visible in provided source code. Tools are implemented as LangChain BaseTool subclasses, but no MCP server registration (via tool definitions with input/output schemas) is shown. Tool definitions are inferred from class attributes.
data_visualization returns filenames but no structured result. Agents cannot reliably parse the response or extract metadata. Response should include file path, data point count, chart metadata as structured JSON.
web_search and wikipedia_search descriptions are vague (35-40 chars). Do not explain what structure is returned, how many results, or pagination limits. LLMs cannot predict output shape.
financial_data documentation does not mention potential API failures (yfinance timeouts, ticker not found, API rate limits). No guidance on retry logic or fallback behavior.
web_scraping accepts any URL with no validation or safety checks. No documentation of URL scheme requirements, timeout limits, or handling of 404/403 errors. Agents could pass malformed URLs or trigger downstream failures.