Multi-tool MCP server integrating Korean weather data, Tavily search, AWS operations, stock trading information, code interpretation, document retrieval (RAG), and file system access
This server exhibits significant definition quality gaps across multiple dimensions. While 9 tools are present, most lack proper schemas, descriptions are often vague or generic, and parameter documentation is sparse. The server mixes high-risk operations (AWS, code execution, drawing) with READ_ONLY tools without consistent error handling patterns. Tool naming is sometimes unclear (e.g., 'repl_coder' vs 'repl_drawer', why not 'execute_python_code' and 'execute_python_visualization'?). No tool annotations present for destructive/irreversible operations despite AWS and code execution risks. Output schemas are not documented anywhere in the visible code.
Draw a graph of the given trend.
Execute Python code to perform calculations or data processing. You MUST provide the 'code' parameter.
Execute a Python script to draw a graph. You MUST provide the 'code' parameter.
Query the keyword using RAG based on the knowledge base.
Returns the last ~period days price trend of the given company name as a JSON string.
Performs a web search using Tavily's AI search engine and generates a direct answer to the query, along with supporting search results.
No input schemas visible for repl_coder, repl_drawer, retrieve, draw_stock_trend, and retrieve_stock_trend. Code execution tools (repl_*) have only a 'code' parameter defined, but no output schema, error recovery guidance, or constraints on execution time/memory.
High-risk operations (code execution, AWS) lack safety annotations. No toolAnnotations present in feature flags. Tools like repl_coder, repl_drawer, and use_aws should include destructiveHint:true and idempotentHint:false to signal to agents that these operations are irreversible and non-idempotent.
use_aws tool has no input validation or parameter constraints visible. The 'parameters' field is a free-form object with no schema validation. The description says 'comprehensive error handling' but no error recovery guidance is documented. No enum constraints for service_name or operation_name, LLMs will hallucinate invalid AWS service names.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Searches recent news articles using Tavily's specialized news search functionality.
Performs a comprehensive web search using Tavily's AI-powered search engine. Excels at extracting and summarizing relevant content from web pages, making it ideal for research, fact-finding, and gathering detailed information.
Execute AWS service operations using boto3 with comprehensive error handling and validation.
Code execution tools (repl_coder, repl_drawer) lack execution safety constraints. No visible timeout, memory limit, or list of allowed/forbidden imports. Description says 'Execute Python code' but does not warn about side effects, file I/O risks, or network access. Agents could accidentally run destructive code (os.system, file deletion, etc.).
retrieve and retrieve_stock_trend tools have vague descriptions ('Query the keyword using RAG...', 'Returns the last ~period days price trend...'). No guidance on what 'RAG' means, what knowledge base is queried, or what format the stock trend JSON contains. LLMs cannot infer downstream field names for tool chaining.
No documented output schemas for any tool. Without knowing the return type and field structure, agents cannot reliably chain tools or extract the right data. E.g., does retrieve return {results: [{title, url, snippet}], total_count}? Does retrieve_stock_trend return {dates: [], prices: []}?
Tavily search tools have good descriptions and parameter schemas (best in the set), but still lack documented output structure. Are results paginated? What fields does each result object contain? Does the LLM need to call again for more results?
draw_stock_trend and repl_drawer have overlapping responsibilities with retrieve_stock_trend/repl_coder. Tool naming does not distinguish between 'execute arbitrary code' (repl_coder) and 'draw a chart' (repl_drawer), both use the same mechanism. Consider renaming to execute_python_code and execute_python_visualization, or merging into execute_python with a type parameter.
No error handling or recovery guidance visible in descriptions. What happens when repl_coder hits a Python SyntaxError? When use_aws receives invalid credentials? When retrieve's knowledge base is empty? Descriptions do not mention retryable vs fatal errors.
Stock trend tools use hardcoded Korean defaults ('네이버', period=30). No enum for company_name, which means LLMs will hallucinate invalid company names. The 'period' parameter has no min/max constraints, an agent could pass period=99999.