Transform any GitHub repository into searchable vector embeddings. MCP server with smart indexing, voyage-context-3 embeddings, and semantic search for Claude/Cursor IDEs.
EmbeDocs MCP provides 4 tools with good parameter schemas and descriptions that explicitly guide agent behavior. However, descriptions are verbose and prescriptive rather than concise and LLM-optimized. Tool names are clear and action-oriented (search, fetch, status). Parameter schemas are well-defined with types, descriptions, and constraints (min/max, defaults). Output schemas are not documented, the server does not specify what fields are returned from searches or status checks, forcing LLMs to infer structure. Error handling is implicit (tool descriptions warn about limitations) but lacks recovery guidance. The 'mongodb-fetch-full-context' tool is critical for the workflow but is not automatically invoked, the agent must learn to chain calls, which increases failure risk.
MANDATORY TOOL - ALWAYS use after search to get COMPLETE files! PURPOSE: Reconstruct COMPLETE file content from chunks. Solves the chunking limitation! ⚠️ CRITICAL INSTRUCTION: You MUST use this tool IMMEDIATELY after ANY search! ⚠️ NEVER skip this step - chunks are INCOMPLETE and MISLEADING without full context! WHEN YOU MUST USE THIS TOOL: • IMMEDIATELY after mongodb-search finds files • IMMEDIATELY after mongodb-mmr-search finds files • For EVERY code file (*.js, *.ts, *.py, *.java, *.md, etc.) • When chunks show "..." or appear truncated • BEFORE providing ANY code examples to user • BEFORE explaining how something works MANDATORY WORKFLOW: 1. Search returns filename like "auth.js" → You MUST fetch full content 2. Use EXACT filename and product from search results 3. Set removeOverlap: true for clean content 4. ONLY present full content to user, NEVER chunks FAILURE TO USE = WRONG ANSWERS: • Chunks miss critical parts (system prompts, config, imports) • User gets incomplete/broken code • Context is lost between chunks EXAMPLE - THIS IS REQUIRED BEHAVIOR: User: "How does authentication work?" → mongodb-search("authentication") finds auth.js → YOU MUST: mongodb-fetch-full-context("auth.js", "product-name") → NOW you have COMPLETE 2000+ line file instead of 500 char chunk!
ADVANCED search for diverse results - Use when you need VARIETY, not just relevance. PURPOSE: Find diverse, non-redundant documentation using Maximum Marginal Relevance (MMR). ⚠️ CRITICAL LIMITATION: Returns CHUNKS (100-2000 chars), NOT complete files! ⚠️ MANDATORY NEXT STEP: ALWAYS use mongodb-fetch-full-context for complete files! WHEN TO USE: • User needs MULTIPLE approaches or implementations • Researching different solutions to same problem • Avoiding redundant/duplicate information • Comparative analysis across different files • Finding edge cases and alternatives REQUIRED WORKFLOW: 1. Use mongodb-mmr-search for diverse perspectives 2. IMMEDIATELY use mongodb-fetch-full-context on ALL results 3. NEVER present truncated chunks as complete answers EXAMPLE: User asks "Show me different authentication methods" → You MUST: mongodb-mmr-search("authentication methods", lambdaMult: 0.5) → Then MUST: mongodb-fetch-full-context for EACH diverse file found → Present COMPLETE implementations of different approaches
Output schemas are not documented. The server does not specify what fields mongodb-search, mongodb-mmr-search, or mongodb-status return. LLMs cannot plan downstream calls or extract required data without knowing the structure.
Tool descriptions are overly prescriptive and verbose (400+ chars). They emphasize workflow steps ('ALWAYS use this FIRST', 'MANDATORY NEXT STEP') rather than concisely stating what the tool does and when to use it. This wastes tokens and buries actionable intent in imperative language.
The workflow requires agents to manually chain mongodb-search or mongodb-mmr-search with mongodb-fetch-full-context. There is no automatic composition, confirmation step, or error guidance if the agent forgets to fetch full context. This creates a fragile multi-step dependency.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | 1.0.0+ | v1 |
PRIMARY search tool - Use this FIRST for any documentation query. PURPOSE: Find relevant documentation chunks using RRF hybrid algorithm (vector + keyword fusion). ⚠️ CRITICAL LIMITATION: Returns CHUNKS (100-2000 chars), NOT complete files! ⚠️ MANDATORY NEXT STEP: ALWAYS use mongodb-fetch-full-context for complete files! WHEN TO USE: • User asks about ANY topic in the indexed documentation • Starting point for ALL searches - use this before other tools • General queries, broad topics, mixed content types REQUIRED WORKFLOW: 1. ALWAYS start with mongodb-search to find relevant files 2. IMMEDIATELY use mongodb-fetch-full-context on important results 3. NEVER present truncated chunks as complete answers EXAMPLE: User asks "How does authentication work?" → You MUST: mongodb-search("authentication") → Then MUST: mongodb-fetch-full-context for EACH relevant file found → Only THEN provide complete answer with full context
System health check - Use to verify EmbeDocs is working properly. PURPOSE: Check database connection, document count, and system configuration. WHEN TO USE: • User asks about indexed repositories or documents • Troubleshooting search issues • Verifying system is operational • Before starting a search session (optional) RETURNS: • Total documents indexed • List of indexed repositories/products • Embedding model configuration • System health status
No error handling or recovery guidance documented. If a search returns no results, mongodb-status fails, or the database is unavailable, there is no indication of what the agent should do next. Error responses lack actionable next steps.
The 'limit' parameter defaults to 5 for search tools, but no rationale is provided. Baseline data shows typical limit defaults are 10-20. For semantic search over documentation, 5 may be too restrictive and force follow-up calls.