MCP server that fetches GitHub repositories and documentation sites, extracting them into agent-ready context blobs with token budgets and search capabilities
Six well-named, read-only tools with clear descriptions (avg 180 chars, within baseline 194). All tools have complete JSON Schema with typed parameters and descriptions. Tool names follow verb_noun pattern (pack_repo, search_context, list_files, get_file, pack_docs, search_docs). Schemas are properly structured with required fields and constraints (e.g., max_tokens, depth 0 - 3, max_pages 1 - 30). However, output schemas are not documented, LLMs cannot see what fields to expect from responses. Error handling is absent: no guidance on retryability, user-fixable vs fatal errors, or recovery paths. No tool annotations (readOnlyHint, idempotentHint). Parameter descriptions lack format/range details in some cases (e.g., 'ref' accepts branch/tag/SHA but doesn't specify format). No batch variants despite tools that agents may call in loops.
Return the full text of a single file in a repository, by path. Use after list_files or search_context to drill in.
List the text files ctx would include for a repository (after stripping binaries, lockfiles and build dirs), with their byte sizes. Cheap way for an agent to see the layout before packing.
Crawl a documentation site from a start URL and return it as one agent-ready context blob: each page extracted to clean Markdown, concatenated with its URL. Stays within the same site section; bounded by depth and page count. Use max_tokens to fit a budget.
Fetch a GitHub repository and return it as one agent-ready context blob: text files concatenated with clear "==== path ====" headers, binaries/lockfiles/build dirs stripped, with a token estimate. Use include/exclude globs to focus, and max_tokens to fit a budget.
Search a GitHub repository and return only the passages that match a query — each with its file path, line number and a relevance score. Far more token-efficient than pack_repo when you need a specific detail (e.g. "where is auth handled?").
Output schemas not documented. LLMs cannot see what fields pack_repo, search_context, list_files, get_file, pack_docs, and search_docs return. This forces LLMs to guess field names and types, risking failed downstream tool calls and wasted context.
No error handling guidance. Tools lack descriptions of failure modes, retryability, and recovery paths. E.g., what happens if a GitHub repo is private, a URL is unreachable, or max_tokens is exceeded? LLMs receive raw errors with no actionable next steps.
Parameter descriptions lack format/range details. 'ref' accepts branch/tag/commit SHA but doesn't specify format. 'include' and 'exclude' accept glob patterns but don't clarify ** vs * vs ? semantics. 'context_chars' range (80 - 4000) is in schema but not in description text.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 81 | <=2025-11-25 | v2 |
Crawl a documentation site and return only the passages that match a query — each with its page URL, line and a relevance score. Token-efficient way to answer a question from docs without packing the whole site.
No tool annotations. Tools lack readOnlyHint (all are read-only) and idempotentHint (all are idempotent). Annotations help LLMs reason about safety and retry logic.
No batch variants. Agents may call search_context or search_docs in loops over multiple queries. A batch_search_context(queries: string[]) would reduce token waste and latency.