A boilerplate for building multi-agent AI clusters using LangGraph with FastAPI, featuring supervisor architecture, MCP tool integration, conversation management, and crew-based agent coordination
Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This is a FastAPI-based boilerplate server that exposes 26 tools wrapping multi-agent orchestration, conversation management, crew/agent lifecycle, and cloud storage operations. Tool definitions are visible in the route files, but quality is uneven. Strengths: tool names follow verb-noun convention (get_, create_, update_, delete_, list_), schemas declare types and formats for most parameters, descriptions are present for all tools. Weaknesses: descriptions are often generic and lack context about when to use each tool; parameter descriptions are sparse or missing constraints; no documented output schemas; error handling guidance is absent; risk annotations are declared in metadata but not formalized in MCP schema; many parameters lack the LLM-optimized constraints (enums, ranges, patterns) needed for reliable agent invocation. The boilerplate structure suggests incomplete MCP integration, tools are HTTP endpoints wrapped by FastAPI, not yet native MCP tools with full schema documentation.
Generic descriptions lack LLM optimization. Most tool descriptions (e.g., 'Create a new crew', 'Get a list of all agents') are under 50 characters and do not explain WHEN to use the tool vs. similar alternatives, WHAT side effects occur, or WHY the agent should select this tool. This forces LLMs to guess intent.
Missing parameter constraints. Pagination parameters (skip, limit) appear in multiple tools but lack documented ranges or limits. No enum constraints on `role` parameter in add_message despite defining valid values (USER, ASSISTANT, AGENT). No regex patterns or length restrictions documented for string parameters like crew name, agent description, or file keys. LLMs cannot infer these constraints from parameter names alone.
Recommendations
Expand all tool descriptions to 50-200 characters. For each tool, explicitly state: (1) What it does. (2) When to use it (e.g., 'Use create_crew before creating agents; agents belong to crews'). (3) What data is returned (e.g., 'Returns conversation object with id, created_at, and user_id fields for use in chat calls'). (4) Any prerequisites or side effects. Example: 'Create a new conversation between a user and a crew. Returns the conversation ID for use in chat and get_messages calls. Each conversation tracks message history independently.'
Add enum constraints to parameters with known-valid values. For add_message, change role description to 'Role of the message sender. Must be one of: USER (user input), ASSISTANT (AI assistant response), AGENT (agent-generated action). Agents must set role=AGENT for tracking.'
Document all pagination parameters with explicit bounds and return-value guidance. Example for get_conversations: 'skip: Number of records to skip (0 - 10000). limit: Max records to return (1 - 100, default 20). Returns: {conversations: [...], total: <int>, has_more: <bool>}' Include a note: 'Large result sets slow LLM reasoning; use limit ≤ 50.'
Add documented output schemas to every tool. For create_conversation, add: 'Returns: {id: uuid, user_id: string, crew_id: uuid, title: string, created_at: ISO8601, updated_at: ISO8601}'. Use ISO 8601 dates, not timestamps, to avoid LLM calculation errors.
Formalize optional parameter structures. For metadata fields, specify sub-schema: 'metadata: Optional object with custom key-value pairs (e.g., {source: "web", priority: "high"}). No nesting; values must be strings or numbers.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 2 points across a rubric change (v1 → v2)
54/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
54
<=2025-11-25
v2
2026-03-09
D
52
-
v1
delete_crewdestructivesource verified72/100
Delete a crew
delete_filedestructiveauthsource verified72/100
Delete a file from storage
download_fileread onlyauthsource verified70/100
Download a file from storage
get_agentread onlysource verified72/100
Get a specific agent by ID
get_agentsread onlysource verified72/100
Get a list of all agents, optionally filtered by crew
get_conversationread onlysource verified75/100
Get a specific conversation by ID
get_conversationsread onlysource verified73/100
Get a list of conversations with optional filters
get_crewread onlysource verified72/100
Get a specific crew by ID
get_crewsread onlysource verified72/100
Get a list of all crews
get_file_urlread onlyauthsource verified72/100
Get a presigned URL for a file
get_messagesread onlysource verified72/100
Get messages for a specific conversation
health_checkread onlysource verified75/100
Health check endpoint to verify the API is running
No documented output schemas. Tool definitions declare input schemas but do not specify what fields are returned. LLMs cannot plan multi-step sequences (e.g., create_conversation followed by chat) without knowing what data is available. Example: create_conversation should document that it returns an 'id' field for use in subsequent chat calls.
Missing error handling guidance. No tool description indicates what errors might occur, how to recover, or whether the operation is retryable. Destructive tools (delete_*) do not warn about irreversible side effects or suggest confirmation steps. This prevents agents from making safe recovery plans.
Risk/permission annotations present in metadata but not in schema. Tools declare Risk flags (READ_ONLY, WRITE, DESTRUCTIVE) in the data structure but do not expose them as MCP tool annotations (readOnlyHint, destructiveHint). This prevents MCP clients from enforcing permission-aware tool selection without parsing tool descriptions.
Parameter type ambiguity in optional fields. Parameters like 'metadata' (object type) lack specificity about expected sub-fields, structure, or validation. This forces LLMs to guess what can be stored and risks malformed requests.
No pagination guidance in list tools. get_conversations, get_crews, and get_agents accept skip/limit but do not document maximum result counts, whether results can be empty, or what metadata is returned (total count, next_cursor). Large result sets can blow context windows.
Incomplete tool composition. The 'chat' tool operates on a conversation_id and message string but does not document what a 'crew' or 'agent' is, why an agent must be assigned to a conversation, or what the chat output contains. This forces LLMs to infer the relationship between conversations, crews, agents, and messages.
chatcreate_conversationassign_tool_to_agent
Add error guidance to destructive operations. For delete_conversation: 'Deletes the conversation permanently and all associated messages. This cannot be undone. If you need to preserve history, use update_conversation with is_active=false to archive instead. Common errors: conversation_id not found (suggest using get_conversations to locate); permission denied (user lacks authority over this conversation).'
Expose Risk annotations in MCP tool schemas using readOnlyHint, destructiveHint, and idempotentHint properties. For delete_* tools, set destructiveHint=true. For get_* and list_* tools, set readOnlyHint=true. This enables MCP clients to enforce permission-aware tooling.
Document the crew/agent/conversation model. Add a root description or discovery tool (e.g., 'A crew is a collection of agents. Each agent is configured with a model and tools. Conversations bind a user to a crew and track message history. Before chatting, create a conversation, then use chat to send messages to the crew.')
Add pagination limits to list results. Cap default limit at 20 - 50 and max limit at 100. For large datasets, implement cursor-based pagination (return next_cursor string) rather than offset/limit to prevent O(N) scanning.
Document idempotency. For create_conversation, specify: 'Idempotent if called with the same conversation_id; otherwise creates a new record. If you receive a conflict error, call get_conversation to verify the record exists.'
Add timeout and retry guidance. For chat (which may call external LLMs): 'Timeout: 60 seconds. If timeout occurs, the crew may have partially processed the message. Retry is safe; the tool will return the cached response if the conversation was already updated.'
Provide natural-language identifier support. Update descriptions to indicate: 'crew_id can be a UUID or crew name (e.g., "support-team"). If multiple crews share a name, UUID is required for disambiguation. Use get_crews to discover crew names and IDs.'