MCP server for consulting large context window models to analyze extensive file collections via OpenRouter
The 'consultation' tool has comprehensive, well-structured parameter documentation with detailed descriptions, an explicit JSON Schema with all required properties defined, and a long, multi-section description that provides context on use cases, limitations, and recovery strategies. Naming is verb-based and clear. However, the tool lacks explicit error handling guidance, output schema documentation, and tool annotations (readOnlyHint, destructiveHint, idempotentHint). The description, while thorough, is 2100+ characters, well above the 1024-character guideline, which wastes tokens. The server defines only one tool, so composition concerns are minimal but unavoidable.
Analyze files with an LLM - provide absolute file paths, query, model, and mode. STATELESS: Each call must contain complete absolute paths. No context is remembered. TIPS: - Hard questions: Spawn 3 parallel calls with varied query formulations - Long instructions: Put them in a file, include in files list, keep query short - WARNING: a query that is BOTH long AND densely packed with special/math characters (< > | & =, parens, LaTeX) can make the call fail with a misleading "'model' is a required property" error (trailing fields dropped). Put bulk/symbolic detail in a file and keep query short and prose-only. Quick mnemonics: - gptt = openai/gpt-5.6-sol + think (latest GPT, deep reasoning) - gemt = google/gemini-3.1-pro-preview + think (Gemini 3.1 Pro, flagship reasoning) - grot = x-ai/grok-4.6 + think (Grok 4.6, deep reasoning [effort xhigh]; 500K context — for bigger bundles use x-ai/grok-4.20 [2M context] instead) - oput = anthropic/claude-opus-4.8 + think (Claude Opus, adaptive thinking) - opuf = anthropic/claude-opus-4.8 + fast (Claude Opus, no reasoning) - fabt = anthropic/claude-fable-5 + think (Claude Fable, deepest reasoning [effort xhigh]; premium, hard problems only) - fabm = anthropic/claude-fable-5 + mid (Claude Fable, high-effort reasoning; premium) - gemf = google/gemini-3-flash-preview + fast (Gemini 3 Flash, ultra fast) - ULTRA = call GEMT, GPTT, GROT, and OPUT IN PARALLEL (4 frontier models for maximum insight) - FUSE = openrouter/fusion (one call: a frontier panel deliberates, a judge synthesizes; mode sets web-research depth). 128K context cap — for hard questions, not giant bundles Performance Modes (use 'mode' parameter): - fast: No reasoning, fastest - mid: Moderate reasoning - think: Maximum reasoning for deepest analysis TIMEOUT TIP: If 'think' times out, retry with 'mid' (especially GPT-5.6). For FUSION this won't help (the cost is the panel of models, not reasoning depth) — instead retry with a single model (e.g. openai/gpt-5.6-sol, google/gemini-3.1-pro-preview) or split the question. On timeout consult7 returns the partial output with a [TRUNCATED] marker rather than discarding it. Files: Absolute paths, wildcards only in filenames (e.g., /path/*.py not /*/path/*.py) Ignores: __pycache__, .env, secrets.py, .DS_Store, .git, node_modules Limits: Dynamic per model - each model optimized for its full context capacity
Tool description exceeds 1024 characters (measured ~2100 chars), violating token efficiency guideline. Description includes extensive model mnemonics, timeout recovery tips, and file handling rules that could be moved to README or inline comments in client SDKs.
No output schema documented. The tool description mentions 'partial output with a [TRUNCATED] marker on timeout' but does not formally specify the response structure (e.g., is it a string, object with fields, array?). LLMs cannot plan downstream actions without knowing what fields to expect.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. The tool is read-only (no mutation), but this is implicit in the description rather than explicitly declared via schema annotations. LLMs benefit from explicit hints.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
No error handling guidance in tool definition. While the description mentions timeouts and truncation, there is no guidance on what error codes/messages the tool returns, how to classify them (retryable vs. fatal), or what the LLM should do in response.
The 'model' parameter accepts a free-form string rather than an enum. While MODEL_EXAMPLES is centralized in tool_definitions.py and descriptions list valid options, the schema does not constrain input to an enum, allowing LLMs to hallucinate invalid model names (e.g., 'gpt-999'). Enums are self-documenting and prevent invalid calls.
The 'files' parameter accepts wildcards (e.g., '/path/*.py') but does not validate that absolute paths are provided. Documentation states 'Absolute file paths or patterns' but the JSON Schema does not enforce this constraint. LLMs may pass relative paths or invalid patterns.