A FastAPI-based chatbot server that integrates MCP (Model Context Protocol) with RAG (Retrieval-Augmented Generation), GitHub context, memory management, and Groq LLM for personalized assistant capabilities
Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
This MCP server has severe definition quality issues across nearly all dimensions. Tool descriptions are minimal (averaging 30-50 chars, well below the 194-char baseline). Most parameters lack descriptions entirely, the schema files show only type and default, with no explanation of what the parameter controls. Input/output schemas are largely undocumented, making it impossible for LLMs to understand what data structures to pass or expect. The tool set shows significant overlap and composition problems: multiple tools do similar things (get_context vs get_context_items, add_*_message vs add_context_item, get_relevant_context vs get_relevant_context_for_query). Error handling is absent, no recovery guidance, no categorization of retryable vs fatal errors, and no invalid-value feedback. No security annotations or permission scopes are declared. The codebase reveals this is a chatbot wrapper around personal assistant functionality, not a production-grade agent tool suite.
Severe parameter description gaps: 18+ tools have parameters with no description text, only type and default. Examples: 'github_data' ("Dictionary containing GitHub API data...") and 'rag_data', 'preference', 'pattern', 'item', 'updates' lack actionable guidance on structure, required fields, or constraints. LLMs cannot reason about which fields to populate.
Massive tool overlap and composition violations: At least 3 pairs of tools do the same thing, 'get_context_items' + 'get_relevant_context', 'add_context_item' + 'add_github_context'/'add_rag_context', 'get_important_memories' defined twice (tools 13 and 22). LLMs waste reasoning cycles choosing between near-duplicate tools. No clear boundaries between context, memory, and protocol layers.
Deduplicate overlapping tools: Consolidate 'get_context_items' and 'get_relevant_context' into a single tool with optional query parameter. Remove the duplicate 'get_important_memories' (tools 13 and 22). Clearly separate memory (long-term learnings about user) from context (current conversation state) with distinct naming and purpose.
Expand all descriptions to 150 - 250 characters and include WHEN/WHY: Replace 'Add GitHub data to the context' with 'Adds GitHub repository metadata and recent activity to enrich the conversation context. Call this when the user asks about projects or recent work. Returns: None, updates internal context state.' Include dependency hints.
Document all object parameter schemas with nested field definitions: For 'github_data', specify required fields (repositories: [{ name, description, language, readme }], recent_activity: [string]). For 'rag_data', list expected structure (found: bool, content: string, sources: [{ title, url }]). Use JSON Schema properties + additionalProperties: false to enforce structure.
Add formal constraints to numeric parameters: Set bounds in schema (limit: { type: integer, minimum: 1, maximum: 100 }, days_threshold: { type: integer, minimum: 1, maximum: 3650 }). Document in description: 'limit (required, 1 - 100; default 5): Maximum number of items to return. Exceeding 100 may slow queries.'
Document output schemas explicitly in tool descriptions and use structured returns: E.g., 'get_relevant_memories returns an array of MemoryItem objects with fields: {id: string (UUID), type: string, content: string, importance: float (0.0 - 1.0), created_at: ISO 8601 date}. Results are sorted by importance descending.' Add a schema in the tool definition for return type.
Generic and vague tool descriptions (avg 30-50 chars, vs 194-char baseline): 'Add GitHub data to the context' and 'Get memory items, optionally filtered by type' do not explain WHEN to use the tool, WHAT it returns, or HOW it differs from similar tools. Descriptions are too short to guide LLM selection (baseline p10=34, p90=392 chars; most here are under 60).
No output schema documentation: Tool descriptions do not state what fields or structure is returned. E.g., 'get_relevant_memories' returns what? A list of MemoryItem objects? Strings? JSON with nested structure? LLMs cannot plan downstream calls or extract needed data without knowing the output schema.
Zero error handling and recovery guidance: No tool description mentions what errors are possible, how to handle them, or what the LLM should do on failure. No distinction between retryable errors (timeout, rate limit) vs user-fixable errors (invalid parameter) vs fatal errors (permission denied). No error messages in code show actionable guidance (e.g., 'Invalid importance score: must be 0.0 - 1.0').
No security annotations or scope declarations: No tool declares what permissions it requires (read:memory, write:memory, delete:memory). No mention of audit logging or who can call destructive tools like 'forget_old_memories'. Agents have no guidance on least-privilege configuration.
Destructive tool 'forget_old_memories' lacks confirmation/dry-run pattern: A tool that deletes memory records should support a dry-run preview before permanent deletion, yet the schema has no 'dry_run' parameter. Agents can accidentally erase data.
No idempotency guarantees stated: Tools like 'add_important_fact', 'add_user_preference' do not state whether repeated calls with identical input produce duplicate records or update existing ones. Agents retrying on ambiguous failures risk creating duplicates.
Parameter type constraints absent: Integer parameters like 'limit' (tools 5, 13, 21) have no bounds (min/max). 'days_threshold' (tool 14) has no range stated in description. 'importance' (tool 8) states '0.0 - 1.0' in description but this should be a formal constraint, not prose. Unbounded params let LLMs pass invalid values (limit=999999, days_threshold=-100).
Object parameters with no internal schema: 'github_data', 'rag_data', 'preference', 'pattern', 'item', 'updates' are all type 'object' with vague descriptions like 'Dictionary containing...'. No nested field schema (required fields, field types, valid keys) is provided. LLMs cannot construct valid input.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in schema: The MCP spec supports tool annotations to mark tools as read-only, destructive, or idempotent. This server does not use them, forcing LLMs to infer safety from descriptions alone (which are too vague to rely on).
all
Add error handling and recovery guidance to every tool description: 'If memory_id is not found, returns error code 404 with message: "Memory item {id} not found. Available memories: [list top 3 by importance]." Call get_important_memories() to see stored memories.' Include categorization: retryable (timeout, rate_limit) vs user-fixable (invalid_id, constraint_violation) vs fatal (permission_denied).
Introduce tool annotations: Mark tools with readOnlyHint (get_* tools), destructiveHint (forget_old_memories), and idempotentHint (tools that safely repeat). Update tool definitions to include these annotations in the MCP schema.
Add idempotency notes to write tools: 'Calling add_important_fact twice with identical fact and importance will update the existing fact's last_accessed timestamp, not create a duplicate.' Provide deduplication field (e.g., fact_hash) if needed.
Add dry-run support to destructive tools: 'forget_old_memories' should accept optional dry_run: boolean parameter. When true, returns a preview: {deleted_count: int, sample_deleted: [{id, content, age_days}]} without performing deletion. Description: 'Set dry_run=true to preview deletions before committing.'
Split complex object updates into separate field-specific tools: Instead of update_context_item with open-ended 'updates' object, offer update_context_item_relevance_score(item_id, new_score) and update_context_item_metadata(item_id, metadata). This prevents invalid field names and makes tool purpose explicit.
Add security scope declarations: Each tool description should state required scopes, e.g., 'Requires scope: read:memory. If scope check fails, returns 403 Forbidden: User does not have permission to read memory items.' Provide a scope reference table in server documentation.
Add pagination support to list tools: Tools like 'get_context_items' and 'get_memory_items' should accept 'offset' and 'limit' parameters, and return 'total_count' and 'next_offset' in response. Describe pagination: 'Pagination: offset (default 0), limit (default 10, max 100). Response includes total_count and next_offset to fetch subsequent pages.'
Create a data types reference document: Define ContextItem, MemoryItem, GitHubData, RAGData as reusable schema objects. Reference them in tool descriptions: 'Returns: array of MemoryItem (see MemoryItem schema).' This reduces duplication and clarifies structure.
Add usage examples to tool descriptions: Not example values (LLMs reuse those), but example use cases. E.g., 'get_relevant_memories: Called when user mentions previous preferences. E.g., "Did I say I prefer weekly summaries?" triggers this tool to find user_preference memories with high relevance to 'summaries'.'
Implement validation error messages: When a tool receives invalid input (e.g., importance > 1.0), return structured error: '{"error": "constraint_violation", "field": "importance", "received": 1.5, "constraint": "0.0 - 1.0", "recovery": "Provide importance between 0.0 and 1.0."}' instead of a generic 400 error.