Static source inference · medium confidence · evidence: structured output
No deprecated protocol patterns detected
Summary
This server exposes 14 tools with mixed quality. Most have descriptions and basic parameters, but schema documentation is incomplete or missing for critical tools. Tool naming follows verb_noun patterns (preflight_list_bundles, preflight_create_bundle, etc.), which is good. However, several high-complexity tools like preflight_memory (12 parameters across 3 layers) and arxiv_search (15 parameters) lack detailed constraint documentation, forcing LLMs to infer valid combinations. Error handling and output schemas are not documented in the provided source. The memory tool has overlapping action enums that could cause confusion. Most tools score 50-70; the arxiv_search tool is better-specified than the core preflight tools.
Tools (14)
arxiv_searchread only50/100
Search arXiv papers by preset, query, or ID. Quick start: `{"preset": "ai_mainstream", "daysBack": 2, "brief": true}`. Get details: `{"idList": ["2301.07041"]}`. Tip: Use brief=true for listing (saves tokens), idList for full paper details.
preflight_checkread onlysource verified48/100
Run duplicates, deadcode, circular dependency, and complexity checks.
preflight_create_bundlewritesource verified53/100
Create a bundle from GitHub, local paths, PDFs, markdown, or web docs.
Input schemas completely missing for 6 tools (preflight_list_bundles, preflight_create_bundle, preflight_check, preflight_lsp, preflight_rag). These tools appear in toolRouter.ts with descriptions but no visible input parameter definitions.
preflight_memory tool combines 8 distinct actions (add, update, search, reflect, stats, list, delete, gc) with 3-layer semantics (episodic, semantic, procedural) into a single tool. This violates single-responsibility principle. LLMs must reason about valid action-layer combinations without explicit guidance. No error documentation for invalid combinations like add+procedural or reflect on episodic.
Output schemas not documented for any tool in provided source. arxiv_search description mentions 'paper count and titles' in brief mode and 'full paper details' in ID list mode, but response structure (fields, format, pagination, chaining IDs) not specified. Other tools give no hints about return values, field names, or downstream tool compatibility.
Recommendations
Add explicit input schemas (JSON Schema with type, enum, minLength, maxLength, pattern constraints) for all 14 tools in server registration code. At minimum: bundleId (string), query (string, 1-1000 chars), action (enum with exact valid values), limit (number, 1-100), format (enum).
Document output schemas for each tool. State what fields are returned, their types, and which IDs can chain to other tools. Example: 'preflight_search_and_read returns {results: [{bundleId, filePath, snippet, lineNumber}], total, nextCursor}.' This enables downstream tool composition.
Split preflight_memory into separate tools by action: memory_add (episodic/semantic/procedural), memory_search, memory_update, memory_delete, memory_stats. Each tool has one clear intent; LLM doesn't have to reason about action-layer validity.
Expand descriptions for low-clarity tools to 50 - 200 characters. For each, answer: What does it do? When should the agent call it (vs similar tools)? What are prerequisites? Examples: 'preflight_rag: Query semantic search index for project context (use after preflight_create_bundle indexes a bundle; returns ranked snippets from documentation and code). Combines multiple repos; use preflight_search_and_read for keyword search.' and 'preflight_memory: Add persistent knowledge to long-term context (episodic = conversation state snapshot for quick retrieval; semantic = experience/facts for similarity search; procedural = global rules). Use to bootstrap agent memory at session start or capture learnings between tasks.'
Score history
Overall score trend
↑ 5 points across a rubric change (v1 → v2)
46/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
46
2025-06-18+
v2
2026-03-09
F
41
1.25.1+
v1
preflight_lspread onlysource verified52/100
Find definition, references, symbols, and diagnostics.
preflight_memorywritesource verified70/100
3-layer long-term memory system for persistent user context (ChromaDB). Layers: episodic (context quick injection), semantic (knowledge base/experience), procedural (global rules/preferences). Actions: add, search, list, update, delete, reflect, stats, gc.
preflight_rag tool has extremely vague description: 'Index/query semantic search content.' No explanation of when to use it, what 'content' means, how it differs from preflight_search_and_read (full-text) vs preflight_memory (long-term context), or what parameters it accepts. LLM cannot determine if it reads or writes.
arxiv_search parameter 'idList' and filter combinations (preset, query, daysBack, fromDate, toDate) lack documented interaction rules. Can both preset and query be provided? What happens if idList is set alongside filters? Description says 'preset overrides query' but doesn't clarify behavior with date filters.
preflight_delete_bundle is marked as DESTRUCTIVE but description is minimal: 'Delete a bundle.' No confirmation step, no dry-run mode, no recovery guidance. Given irreversible nature and agents' tendency to retry, missing error handling and confirmation pattern is a risk.
preflight_check tool description says it 'requires: path' but no documentation of what path format, whether absolute or relative, whether it must exist locally or in a bundle. bundleId is the primary MCP abstraction; requiring 'path' is inconsistent. Unclear if this operates on indexed content or fresh local filesystem.
Error handling and recovery guidance absent from all tool descriptions. No hints like 'If bundle not found, call preflight_create_bundle' or 'Search may timeout for large bundles, narrow query scope.' LLMs cannot self-correct on failures.
Add parameter descriptions where missing. For arxiv_search 'preset', explain each enum: 'ai_mainstream = major LLM/AI papers from top venues (default); ai_full = comprehensive AI; llm = language models; ml = machine learning; cv = computer vision; multimodal = multi-modal models.' For preflight_memory 'layer', add: 'episodic = session-scoped context (facts about this conversation); semantic = long-lived knowledge base (project patterns, best practices); procedural = global agent rules (always X when Y).'
Add error guidance to all destructive tools. Example: 'preflight_delete_bundle: DESTRUCTIVE, cannot be undone. Before calling, confirm with user. If deletion succeeds, returns {bundleId, deletedAt}. If bundle not found, returns error: 'Bundle not found, call preflight_list_bundles to verify ID.' No recovery tool exists; advise user to re-create from source.'
Document preflight_check path parameter: 'Local filesystem path (absolute or relative to PREFLIGHT_STORAGE_DIR) containing source code. Runs static analysis: duplicate code detection, dead code identification, circular dependency detection, cyclomatic complexity. Returns JSON report or error if path invalid or inaccessible.'
Add idempotency guidance. Example: 'preflight_generate_card: Idempotent (safe to retry). If regenerate=false (default), returns cached card if exists; if regenerate=true, rebuilds from scratch. Each call returns same result for same bundleId+repoId.' This tells agents which tools can be safely retried.
Document parameter dependencies for arxiv_search: 'If preset is provided, query is ignored. If both query and preset are set, preset takes precedence. Filters (daysBack, fromDate, toDate) apply in all modes. idList mode (fetching specific papers by ID) bypasses search and filters; provide idList alone to fetch paper details.'
Add 'nextSteps' hints to TOOL_REGISTRY entries. Example for preflight_search_and_read: 'nextSteps: ["Use preflight_read_file to open full file at path returned", "Use preflight_lsp for definition/reference details", "Use preflight_generate_card to extract knowledge summary"].' These guide multi-step workflows.