A FastAPI-based MCP server that provides RAG (Retrieval-Augmented Generation) capabilities for n8n workflows and documentation using ChromaDB vector search, with support for multiple LLM providers (OpenAI, Anthropic, Google).
This MCP server exposes 8 tools for n8n RAG functionality. While all tools have descriptions and most have basic input schemas, the definitions exhibit significant gaps in parameter documentation, output schema specification, and error handling guidance. Only 3 of 8 tools have adequate parameter descriptions; 5 tools lack proper output schema documentation. Parameter names are action-verb-based (good), but many parameters lack type clarity and validation constraints. Error handling is minimal, tools return generic error objects without recovery guidance. The server uses FastAPI with HTTP transport, which is acceptable, but tool definitions are scattered across the codebase without a unified, MCP-compliant registration pattern.
Build a flexible prompt for n8n workflow generation, supporting both structured (goal, triggers, integrations, steps, constraints) and plain text inputs with RAG context injection.
Generate query expansion terms for semantic search using heuristic term matching (webhooks, Airtable, Slack, Gmail, Cron, etc.).
Extract JSON from LLM text response, handling markdown code fences and malformed JSON.
Perform keyword search on ChromaDB collection using source metadata and text substring matching.
Pack retrieved workflow and documentation chunks into a context string with token budget awareness, interleaving workflow and doc chunks at a 2:1 ratio.
Refine/rewrite a prompt using the LLM-based rewrite_prompt function.
Output schemas are not documented for any tool. Callers cannot determine the structure, type, or fields of responses. This forces LLMs to infer structure from trial-and-error or internal code inspection.
Parameter names and types are inconsistent or problematic. 'collection' param in keyword_search accepts an internal ChromaDB object (not a natural identifier). 'prompt_input' in build_prompt_flexible is an ambiguous union type (string|object) without clear guidance on when to use each. 's' in tcount is a single-letter parameter name, violating readability. 'q' across multiple tools is acceptable but inconsistent with common 'query' naming.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 22 | - | v1 |
Retrieve relevant n8n workflow and documentation chunks using hybrid vector and keyword search, with query expansion and context deduplication.
Count tokens in a string using tiktoken cl100k_base encoding.
Many parameter descriptions lack constraints (min/max for integers, allowed values for strings, format specifications). E.g., 'budget_tokens' in pack_context has no min/max; 'n_results' in keyword_search has no limits; 'prompt' in refine_prompt has no length constraints. LLMs will pass arbitrary values that may break the tool.
Tool descriptions are vague or technical. 'pack_context' doesn't explain when to call it. 'keyword_search' is overly technical and doesn't differentiate from 'retrieve'. 'tcount' name is an abbreviation; LLMs won't understand 'tcount' vs 'count_tokens'. 'expand_queries' lacks clarity on OUTPUT format.
Error handling is minimal or absent. Tools return generic error objects ('error': str(e)) without recovery guidance. E.g., refine_prompt catches all exceptions and returns {'error': '...'}, but doesn't tell the LLM whether to retry, call a different tool, or ask the user.
Parameter descriptions are minimal or missing context. Most describe WHAT the parameter is, not WHY the LLM might want to call the tool. E.g., 'retrieve' says 'The query/search term to retrieve context for' but doesn't explain what 'context' means or when this is better than other tools.
No pagination, result limits, or truncation guidance. 'retrieve' doesn't document how many results are returned. 'keyword_search' doesn't enforce a max result count. 'pack_context' doesn't explain what happens if context exceeds the token budget.