Stateless MCP server with live node sync, bulletproof validation, and embedded LLM for n8n workflow automation. Supports both stdio (Claude Desktop) and HTTP modes.
Mixed quality across 12 tools. Strong points: most tools have descriptions and input schemas with some parameter constraints (enums, defaults, ranges). Critical gaps: descriptions lack actionable 'when to use' guidance; many parameters missing type information or descriptions; no visible output schemas documented; error handling does not guide LLM recovery; several tools combine multiple concerns (e.g., execute_agent_pipeline chains 4 operations); tool names do not consistently follow verb_noun convention (e.g., 'query_graph', 'nano_llm_query' are vague vs. 'search_graph', 'search_nodes').
Run full agentic GraphRAG pipeline: (1) discover workflow patterns from goal, (2) query knowledge graph for node relationships, (3) generate n8n workflow from patterns + graph insights, (4) validate generated workflow. Returns complete workflow ready to deploy. IMPORTANT: This is the primary way to use nano agents - provide a natural language goal and get back a complete, validated n8n workflow.
Query the n8n knowledge graph for node relationships. Useful for understanding which nodes work well together. Returns relevant nodes, edges showing relationships, and a summary of the subgraph. Run this to explore node combinations for your workflow.
Discover workflow patterns matching the goal. Useful for understanding common workflow structures. Returns pattern matches with confidence scores and suggested nodes. Run this FIRST to understand what patterns exist for your goal.
Generate n8n workflow JSON from discovered patterns. Useful for creating workflows based on specific patterns. Returns complete workflow with nodes, connections, and settings. Run after pattern_discovery to generate a workflow from a matched pattern.
Get status of nano agent orchestrator and individual agents. Returns initialization status, current tasks, and configuration. Useful for monitoring and debugging agent execution.
Tool names lack clear verb_noun convention. 'query_graph', 'nano_llm_query', 'execute_agent_pipeline', 'execute_pattern_discovery', 'execute_graphrag_query', 'execute_workflow_generation' are vague about what they do. LLMs cannot infer intent from names alone. Should be: 'search_graph', 'search_nodes_by_intent', 'generate_workflow', 'discover_patterns', 'search_graph_relationships', 'generate_workflow_from_pattern'.
Output schemas NOT documented in tool descriptions. The source shows input schemas with Zod validation but no visible return type documentation. LLMs cannot plan downstream calls without knowing what fields are returned. Examples: nano_llm_query returns quality scores and node tiers (mentioned in description) but structure not formally specified; execute_agent_pipeline claims to return 'complete workflow ready to deploy' but no schema shown.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-04-07 | F | 26 | - | v1 |
n8n_list_mcp_workflows - List workflows created/managed by MCP Use cases: - List workflows created by Claude via MCP - Filter out manually created workflows - Track MCP workflow adoption - Provide workflow suggestions to agents NEW in n8n 1.113.3: MCP workflow metadata support
n8n_monitor_running_executions - Monitor currently running executions Use cases: - Monitor active workflow executions - Prevent duplicate runs - Display real-time execution status - Detect stuck executions NEW in n8n 1.113.3: Enhanced execution filtering
🔄 Retry a failed or stopped n8n workflow execution (NEW in v3.0.0) Use this when: - An execution failed due to temporary issues (credentials, network, etc.) - You've fixed the underlying problem and want to retry - You want to re-run a workflow with the same input data The tool provides intelligent retry suggestions based on the execution status and error details.
⭐ GET NODE PERFORMANCE TIERS: View empirically-derived node valuations from the learning system. Returns performance tiers for n8n nodes based on: - Success rates from real usage - Temporal difference (TD) learning credit assignment - Quality scores from past queries - Execution efficiency Tiers: gold (>= 0.80), silver (0.60-0.79), bronze (0.40-0.59), standard (< 0.40) Use to understand which nodes perform best for your use cases and to build workflows with proven, reliable nodes.
📊 GET SYSTEM OBSERVABILITY: View metrics, traces, and statistics from the nano LLM system. Returns: - Prometheus-compatible metrics (counters, gauges, histograms) - Distributed traces (Jaeger/Zipkin format) - System statistics (traces collected, nodes valued, strategies evaluated) Use for monitoring system health, debugging issues, and understanding system behavior.
🤖 INTELLIGENT NODE SEARCH with intent routing, quality assurance, and learning. This is the primary interface to the dual-nano LLM system which: 1. Understands your query intent (direct lookup, semantic search, workflow pattern, property search, integration, recommendation) 2. Routes to optimal search strategy 3. Validates and assesses result quality (5 dimensions: quantity, relevance, coverage, diversity, metadata) 4. Automatically refines queries up to 3 times if needed 5. Learns from results via reinforcement learning with TD(λ) credit assignment 6. Returns results with quality scores and node performance tiers (gold/silver/bronze/standard) Use this when you need intelligent node recommendations with guaranteed quality. The system learns and improves over time.
Graph-based retrieval from the locally cached n8n knowledge graph. Returns 3–5 relevant nodes, edges, and a concise subgraph summary (<1K tokens). Always uses the local cache; no live n8n calls.
execute_agent_pipeline combines 4 distinct operations (pattern discovery, graph query, generation, validation) into one tool. Violates single-responsibility principle. Agents cannot compose steps selectively or recover from mid-chain failures. Should split into: discover_patterns, search_graph, generate_workflow, validate_workflow.
Error handling guidance missing. Tool descriptions state what they do but not what errors can occur or what LLM should do next. Examples: n8n_retry_execution can fail if execution cannot be retried (source shows 'canRetry' check) but description does not mention this; nano_llm_query can return low-quality results but no guidance on recovery (retry, refine query, etc.). Pattern: recovery-guide
Parameter descriptions inconsistent or minimal. Examples: n8n_monitor_running_executions 'workflowId' param lacks guidance on whether this is the n8n workflow ID or an MCP-managed ID; query_graph 'top_k' documented as 'Max number of nodes to return' but no guidance on performance trade-offs or memory impact; nano_llm_query 'userExpertise' has enum values but no explanation of how each level affects results (beginner→ simpler explanations?, expert→ more advanced patterns?). Rubric: 'A parameter named 'type' is ambiguous, the description disambiguates it'
No pagination documented for list-returning tools. n8n_list_mcp_workflows accepts 'limit' but no mention of 'offset', 'cursor', 'page', or total count. Rubric: 'Tools returning lists should accept page/offset and limit parameters and return a total count or next_cursor.' nano_llm_node_values has 'limit' (default 50) but no pagination control. Result sets may exceed context window.
Tool descriptions contain vague language and examples that may be treated as literals. Examples: 'send slack message when airtable record updated', 'generate report from data' are use-case examples but LLMs often treat example values as blueprints. Rubric: 'Do not put example values in the description. LLMs tend to reuse example values literally'.
Missing security annotations. No visible mention of permission gates, scope declarations, or rate limits in tool definitions. nano_llm_observability exposes 'traces' and 'metrics' which could leak internal system information without access control. Pattern: scope-declaration, permission-gate
Inconsistent parameter naming conventions. Examples: 'top_k' vs 'topK' (execute_graphrag_query uses 'topK', query_graph uses 'top_k'); 'includeStats' vs 'includeHistory' (both boolean flags but different semantics not explained); 'limit' used across multiple tools but some have min/max constraints (1 - 100) while others do not (nano_llm_node_values: default 50, no stated upper bound). Rubric: 'When multiple tools operate on the same resource, their names must make the distinction obvious.'