Documentation Code Extractor with MCP integration for documentation crawling and code search
CodeDox provides 5 well-named tools with mostly complete schemas and reasonable descriptions. Tool naming follows verb_noun conventions (init_crawl, search_libraries, get_content, get_snippet, get_page_markdown). Schemas are present with type definitions and constraints (enums, min/max bounds). However, descriptions vary in depth and some critical parameter documentation is sparse. Error handling is not evident from the provided code. Output schemas are mentioned in descriptions but not formally documented. The server's composability is good, tools chain well (search_libraries → get_content → get_snippet/get_page_markdown). Pagination support is present in search and get_content tools. Overall, this is a solid B-grade implementation with room for improvement in error recovery guidance and output schema formalization.
Search code snippets in a library. Results include SOURCE URLs for full docs via get_page_markdown tool. Modes: 'code' (default, with markdown fallback) or 'enhanced' (always searches markdown). Library can be name or UUID. Snippets truncated to 500 tokens; use get_snippet tool for full content.
Get full documentation markdown by URL (from get_content SOURCE) or snippet_id. Use query to search within the page. Control size with max_tokens or paginate with chunk_index. Provide url OR snippet_id, not both.
Get a specific code snippet by ID. Control size with max_tokens (100-10000, default 2000). Use chunk_index for large snippets. Returns formatted code with SOURCE URL.
Initialize a new web crawl job for documentation
Search or list available libraries. Returns library_id (for use with get_content), name, description, and snippet_count. Supports pagination via page/limit params. Empty query lists all.
Error handling guidance absent from tool descriptions. No indication of what errors are retryable, user-fixable, or fatal. No recovery paths documented (e.g., 'If library not found, call search_libraries() first').
Output schemas not formally documented. Descriptions mention return fields ('library_id', 'name', 'description', 'snippet_count', 'SOURCE', 'error') but no structured output schema is provided in the tool definition. LLMs cannot reliably parse unstructured responses.
get_page_markdown has mutually exclusive parameters (url vs snippet_id) but the description does not explicitly state this constraint. LLMs may pass both, causing ambiguous behavior.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 56 | 2024-11-05+ | v1 |
init_crawl accepts 'metadata' as a free-form object. No schema for the metadata structure. Documentation lists examples ('repository', 'description') but does not define required/optional fields or types. This invites unpredictable LLM behavior.
get_content description mentions 'SOURCE URLs' and 'truncated to 500 tokens' but no formal output schema documents the field names, types, or structure. Downstream tools (get_snippet, get_page_markdown) depend on this output but the contract is informal.