Documentation site crawler and MCP server for AI tool integration. Crawls documentation sites, builds full-text search indexes, and exposes crawled content via MCP tools.
A solid Go/MCP implementation with 10 well-defined tools. All tools have clear descriptions (194-character average aligns with baseline), structured input schemas with proper types, and good composition. Tool annotations are correctly applied (readOnlyHint, destructiveHint, idempotentHint). However, output schemas are not documented in source code, and there is no evidence of error handling guidance or recovery patterns in the handler implementations shown. The descriptions are clear and contextual, avoiding example values. Naming follows verb_noun conventions consistently.
Cancel a running or pending crawl job by job ID. Has no effect on jobs already in a terminal state.
Start a background crawl for a configured site. Returns immediately with a job ID.
Returns server identity, configured sites, and recent crawl jobs in one call. Call this first to orient yourself: it consolidates what would otherwise require list_sites plus several get_job_status calls. The MCP tool list is already advertised by the protocol so it is not duplicated here.
Return the most recent crawl summary for a site (last_crawl_started_at/ended_at, total_pages, mode, age_seconds) plus output/state dir presence and any running job. Use this to decide whether to query the existing crawl or run crawl_site first.
Get the status of a crawl job
Fetch a URL live over the network and return its content as markdown. This is an on-demand fetch independent of any crawl: it does not read the stored crawl output, and it ignores the site's configured content_selector and scope.
Output schemas not documented in source code or tool definitions. While tool descriptions explain what each tool returns conceptually, LLMs need explicit return type schemas to plan downstream calls and extract fields reliably.
No evidence of error handling guidance in handler implementations. Error responses should categorize failures as retryable, user-fixable, or fatal, and guide the LLM on next steps (e.g., 'Job not found. Call list_sites() to see available sites'). Currently no recovery guides visible.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 66 | <=2025-11-25 | v2 |
List crawled pages for a site, paginated and sorted by URL. Returns metadata only (URL, title, depth, crawled_at, content_length). Pass any URL returned here to read_page to get its stored markdown; get_page re-fetches a URL live rather than returning the crawled copy.
List all configured sites available for crawling
Return a page's markdown from the stored crawl output, without any network access. This is the counterpart to get_page: read_page serves the crawled copy that already had the site's content_selector applied, while get_page re-fetches the URL live. Use list_pages to discover URLs, then read_page to read them. Large pages are truncated at max_bytes; follow next_offset to read the rest.
Full-text search across all crawled documentation, ranked by relevance (BM25 with stemming), with zero network access. Results carry the page URL, its section heading path, and a snippet with match terms marked [like this]; follow up with read_page for the full page. Supports FTS5 syntax: quoted phrases, OR, and trailing * for prefix matching. Searches the stored index only; run crawl_site first for uncrawled sites.
Tools accept numeric parameters (max_results, offset, max_bytes, limit) without documented bounds. get_page max_bytes accepts up to 1MB; list_pages max_results caps at 1000. These constraints should be explicit in parameter descriptions to prevent LLMs from passing invalid values.
get_page tool returns live-fetched content but read_page returns stored/cached content. The distinction is clear in descriptions, but there is no parameter or output field marking which version was returned, forcing LLMs to track this mentally across calls.