Save up to 80% tokens when AI reads code — MCP server for token-efficient code navigation, AST-aware structural reading instead of dumping full files into context window
Token Pilot defines 25 code-reading tools with generally strong naming and descriptions. Tool names follow verb_noun convention (smart_read, read_symbol, find_usages, etc.) and clearly indicate their purpose. Descriptions are detailed and contextual, explaining when to use each tool and what it returns. However, schema quality is uneven: while parameters are named and described, the source code does not show explicit JSON Schema type definitions for inputs or documented output schemas. The rubric requires visible schema definitions; inferred schemas cannot be scored at full value. All tools are READ_ONLY with no destructive operations, which simplifies error handling but limits the server's utility for write operations. Tool composition is excellent, tools chain naturally (e.g., smart_read → read_diff → read_for_edit before write) and accept session_id for deduplication, showing thoughtful design. Descriptions average 150-200 characters, well within the 10 - 1024 target range. Parameter descriptions are present and explain purpose, though formal constraints (enums, patterns, min/max) are sparse in the visible code.
Show call tree for a function or method
Code quality audit — TODOs, deprecated patterns, structural issues
Explore directory structure and contents
Starting work on a directory — outline + imports + tests + git log in one call
Dead code detection — find unreferenced symbols across project
Find where a symbol is defined, imported, or used — semantic search (definitions + imports + usages)
Module architecture — dependencies, dependents, public API
Output schemas not documented. While parameter descriptions are present, the source code does not show explicit JSON Schema definitions for tool return types. LLMs cannot plan downstream tool calls without knowing what fields will be returned.
Generic tool name 'explore' lacks specificity. It does not clearly distinguish its purpose from 'explore_area'. LLM may conflate these tools when making selection decisions.
Input parameter schemas lack formal constraints (enums, min/max, patterns). For example, 'mode' parameters in find_usages and explore_area accept string values but do not enumerate valid options in the schema, forcing LLMs to infer from description alone.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Transitive dependency path(s) between two modules
List all symbols (classes, functions, methods) in a directory in one call
Get overview of a new codebase/unfamiliar project
Verify edits after editing — shows only changed hunks (REQUIRES smart_read BEFORE editing)
MANDATORY before ANY Edit/Write tool call on existing code file — returns exact old_string for Edit's old_string parameter
Read a specific line range from a file
Read markdown/yaml/json/csv section — loads one heading/key/row-range (not the whole file)
Read ONE function/class body (not the whole file) — loads only that symbol
Read MULTIPLE function/class bodies from same file in one batch call instead of N separate calls
Understand file dependencies — imports, importers, tests — ranked by relevance
Session analytics — token savings by tool, intent classification, decision traces
Query session token budget and spending
Capture session state as compact <200 token block — goal, decisions, confirmed facts, files, next step
Review git changes mapped to functions/classes (not raw git diff) — structured commit changes
Structured commit history (not raw git log) — organized with categories
Read a code file with AST-aware structural compression, returning structure instead of full content — 60-80% fewer tokens than full-file read
Batch read up to 20 files at once instead of making separate calls
Run tests and return structured pass/fail results (not raw test output)
Error handling guidance not visible. The source code shows tool definitions but no documented error recovery patterns. Tools should specify what errors are retryable, what indicates user misconfiguration, and what guidance to provide to the agent.
session_id parameter used inconsistently and without clear explanation. Many tools accept session_id for deduplication, but the purpose and behavior are not documented in parameter descriptions. Is it required or optional? What happens if omitted?