A LangChain playground using TypeScript with agent-based infrastructure investigation capabilities including AWS ECS/RDS, New Relic, Sentry, and code research domains
This server has critical gaps in tool definition quality. Of 11 tools examined, none have visible input schemas in the provided source code. Descriptions are present but often vague and lack actionable guidance for LLM selection. Many tool names violate single-responsibility principles (e.g., 'investigate_and_analyze_*' combines multiple concerns). Tools 8-11 are proxied MCP adapters with inferred definitions, capping their individual scores at 50. No parameter descriptions, no output schemas documented, and error handling guidance is absent. This is a category D/F server.
Combined tool that fetches logs from New Relic and analyzes them in one step. Keeps raw log data internal to reduce token usage when passing context to other domain agents.
Use LLM to generate NRQL queries for New Relic logs
Use LLM to generate query for fetching trace logs from New Relic
Execute API calls without LLM inference to fetch New Relic investigation context
Combined data gathering AND analysis in one call for AWS ECS tasks. Keeps raw AWS API data internal and returns only the analysis summary, reducing token usage when supervisor passes context to other agents.
Combined data gathering AND analysis in one call for AWS RDS instances. Keeps raw AWS API data internal and returns only the analysis summary, reducing token usage when supervisor passes context to other agents.
No input schemas visible for any tool. Source code shows tool definitions in TypeScript files but no JSON Schema parameter definitions are provided or referenced. This violates the fundamental requirement that input parameters be formally specified with types and constraints.
Multiple tool names violate single-responsibility principle by combining investigation AND analysis (e.g., 'investigate_and_analyze_ecs_tasks', 'investigate_and_analyze_rds_instances', 'fetch_and_analyze_logs'). Per Agentic Tool Patterns, tool names should reflect one action. These should be split into separate discovery and analysis tools.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 27 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Combined data gathering AND analysis in one call for Sentry issues. Keeps raw Sentry API data internal and returns only the analysis summary, reducing token usage when supervisor passes context to other agents.
Web search tool via Brave Search MCP adapter
Tool to resolve library identifiers via Context7 MCP adapter
Read-only Kubernetes tool to describe resources via kubectl through MCP adapter
Read-only Kubernetes tool to get resources via kubectl through MCP adapter
Descriptions are present but lack LLM-actionable guidance. Descriptions like 'Combined data gathering AND analysis in one call for AWS ECS tasks' (59 chars) do not explain WHEN to use this tool vs an alternative, what parameters it expects, or what the response structure contains. Missing: parameter descriptions, output schema documentation, and selection rationale.
No output schemas documented. LLMs cannot plan downstream tool calls without knowing response structure. Tools like 'get_investigation_context' and 'generate_log_nrql_query' have no visible response documentation.
Proxied MCP adapter tools (tools 8-11) have minimal descriptions and inferred definitions. Tool names like 'mcp__brave-search__brave_web_search' and 'mcp__kubernetes-readonly__kubectl_get' provide no context about parameters, expected behavior, or integration patterns. Descriptions are <30 chars for most adapters.
No error handling guidance visible. Tools like 'generate_log_nrql_query' and 'generate_trace_logs_query' do not document failure modes, recovery strategies, or what error messages the LLM should expect. Per pattern:recovery-guide, errors must tell LLMs what to do next.
Tool names lack verb clarity. 'investigate_and_analyze_*' does not match standard verb_noun convention (get_, create_, update_, delete_, search_, list_). This makes it harder for LLMs to infer intent from the name alone.