Unified MCP server routing Gemini, Codex, Claude, Grok, Ollama, and Antigravity CLI tools behind runtime availability checks. Provides code review, second opinions, analysis, and AI-to-AI collaboration via multiple LLM providers.
The server implements 20 tools with substantial variation in quality. Most tools have descriptions and parameter schemas, but naming conventions are inconsistent, schema completeness is mixed, and error handling guidance is minimal. The 'ask-*' tools are reasonably well-documented but lack output schema definitions, making it unclear what structured data agents receive. Multiple ping tools and duplicate tool names (ask-gemini, ask-codex, ask-grok, ask-ollama, ask-antigravity appear twice with slightly different descriptions) signal composition issues and naming ambiguity. Tool descriptions range from 80 - 300+ characters, generally falling within acceptable bounds, but parameter descriptions for optional fields are sometimes vague. No evidence of idempotency markers, confirmation patterns for destructive operations, or recovery guidance in error messages. The schema definitions use proper JSON Schema types and minLength/maxLength constraints, but lack explicit output schema documentation, forcing agents to infer response structure.
Consult Google's Antigravity CLI (agy) through Ask LLM's canonical executor. Requires a supported authenticated agy installation. Output is bounded to Pi's 50KB/2000-line limits.
Send a prompt to Google's Antigravity CLI (agy) for a subscription-backed second opinion, code review, or analysis. Requires agy >=0.43.0, installed and logged in once; unsupported versions fail before model invocation. Defaults to the gemini-3.1-pro model at high reasoning effort, falling back to gemini-3.5-flash on a rate limit. Returns human-readable text plus a structured response.
Send a prompt to Anthropic Claude Code CLI for an independent second opinion, code review, or architecture critique. Defaults to claude-3-7-sonnet with claude-3-5-haiku fallback. Claude runs in safe mode with only Read, Glob, and Grep tools, so it can inspect context but cannot edit files or execute commands. Supports native sessions via sessionId and returns a structured AskResponse.
Consult OpenAI Codex through Ask LLM's canonical executor. Read-only by default; use workspace-write only for an explicit write flow such as codex-image. Output is bounded to Pi's 50KB/2000-line limits.
Send a prompt to OpenAI Codex CLI (defaults to codex-4 with automatic fallback on quota errors). Use for code review, second opinions, analysis, and AI-to-AI collaboration. Returns both human-readable text and structured response (provider, model, sessionId, usage). Calls are ephemeral by default; pass sessionId: "" on the first call to persist its Codex thread_id, then pass the returned sessionId on follow-up calls to continue the conversation.
Duplicate tool names with inconsistent descriptions. Tools 1 and 16 are both 'ask-gemini', tools 4 and 15 are both 'ask-codex', tools 9 and 17 are both 'ask-grok', tools 11 and 18 are both 'ask-ollama', tools 13 and 19 are both 'ask-antigravity'. The primary versions have detailed descriptions (e.g., 'Send a prompt to Gemini CLI...'), while Pi canonical versions have truncated, generic descriptions ('Consult Gemini through Ask LLM's canonical executor...'). This forces agents to disambiguate at selection time and wastes reasoning cycles. Identical names violate the tool composition pattern.
Excessive ping tools. Five separate ping implementations (tools 3, 6, 8, 10, 12, 14) test connectivity to different providers. This is tool sprawl, a single 'health_check' or 'get_status' tool could return a map of provider statuses, reducing cognitive load and complexity. Multiple near-identical tools violate the composition pattern.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2025-06-18+ | v2 |
| 2026-03-09 | D | 58 | - | v1 |
Send a code edit request to OpenAI Codex CLI and get structured search/replace edit blocks back (via codex --output-schema). Codex reads the existing files (read-only) and proposes precise, applyable changes for Claude to apply. Use this when you want Codex to suggest specific code modifications to existing files rather than just analysis.
Use Cursor Agent as a model-neutral read-only harness. Provider (claude, codex, gemini, grok) and exact model ID are separate and must agree; Auto or noncanonical catalog IDs are refused. Prompts above 16KB are piped over stdin. Requires an authenticated Cursor CLI and may consume included usage or on-demand spend; no spend settings or fallback are changed.
Consult Gemini through Ask LLM's canonical executor (`gemini-3.1-pro-preview` → `gemini-3.8-flash` on quota), including validation, sessions, and structured response. Output is bounded to Pi's 50KB/2000-line limits.
Send a prompt to Gemini CLI for code review, second opinions, analysis, and AI-to-AI collaboration. Defaults to gemini-3.1-pro with automatic fallback to gemini-3.5-flash on quota errors. Returns both human-readable text and structured response (provider, model, sessionId, usage).
Send a code edit request to Gemini CLI and get structured OLD/NEW edit blocks back. Gemini analyzes the files and returns precise, applicable code changes. Use this when you want Gemini to suggest specific code modifications rather than just analysis.
Send a prompt to Grok via xAI API or CLI for second opinions, code review, and analysis. Requires XAI_API_KEY environment variable and incurs metered API charges. Returns structured response with provider, model, and usage information.
Consult Grok through Ask LLM's canonical xAI API executor. Requires XAI_API_KEY and may incur metered API charges; no billing changes or model fallback are performed. Output is bounded to Pi's 50KB/2000-line limits.
Consult the configured local Ollama model through Ask LLM's canonical executor. No external provider data transfer; output is bounded to Pi's 50KB/2000-line limits.
Send a prompt to a local Ollama LLM for second opinions, code review, and analysis. No external data transfer; fully local processing. Defaults to the configured Ollama model. Returns structured response with provider, model, and usage information.
Test connectivity with the Claude MCP server and check whether Claude Code CLI is installed
Test connectivity with the MCP server
Test connectivity with the Grok MCP server and check whether Grok CLI or xAI API is available
Test connectivity with the Gemini MCP server and check whether gemini CLI is installed
Test connectivity with the Ollama MCP server and check whether Ollama is running
Test connectivity with the Antigravity MCP server and check whether agy is installed
Missing output schema documentation. Tools declare input schemas clearly but nowhere in the provided code is the output schema documented. Agents cannot know what fields a successful response contains, e.g., does ask-gemini return {text, model, sessionId, usage} or a different structure? This forces agents to explore responses dynamically, risking failed downstream chains.
Weak ping tool descriptions. All ping variants have descriptions like 'Test connectivity with the Gemini MCP server and check whether gemini CLI is installed' (40 - 60 chars). These lack actionable detail: When should an agent call ping vs catching errors from ask-gemini? What does a successful ping look like vs a failure? Should agents retry, fall back to another provider, or report to the user? Descriptions do not guide agent decision-making.
Parameter descriptions lack constraint detail. Many optional parameters have minimal guidance. Example: ask-gemini's 'model' param says 'DO NOT set this parameter. The tool automatically uses gemini-3.1-pro and falls back to gemini-3.5-flash on quota errors. Only set this if the user explicitly requests a specific model.' This is useful, but ask-codex's 'reasoningEffort' enum declares ['low','medium','high','xhigh','max','ultra'] without explaining what 'max' or 'ultra' actually do or when to use them. Agents cannot reason about which level suits a given task.
No idempotency or error recovery guidance. None of the ask-* tools document whether they are idempotent. If ask-gemini is called twice with the same sessionId and prompt, does it create a duplicate message or resume the session? Tools that support sessionId should clarify: does resuming a session replay prior context automatically, or does the agent need to re-send prior messages? Without this, retry logic is ambiguous.
Pi canonical tools have truncated descriptions that lack selectivity guidance. E.g., ask-codex (Pi canonical) says 'Consult OpenAI Codex through Ask LLM's canonical executor. Read-only by default; use workspace-write only for an explicit write flow such as codex-image. Output is bounded to Pi's 50KB/2000-line limits.' This explains what it does but NOT when to choose it over other ask-* tools or why the agent should prefer this variant. Descriptions do not answer 'When should the LLM select THIS tool instead of THAT one?'
Missing error handling guidance. Descriptions mention fallbacks (e.g., 'falls back to gemini-3.5-flash on quota errors') but do not tell the agent what to do if the fallback also fails. Should the agent retry with a different provider? Report an error to the user? Fall back indefinitely? No tool description includes recovery guidance like 'If this returns a 429 Quota Exceeded error, try ask-gemini-edit with a shorter prompt or switch to ask-ollama for local processing.'
sandbox parameter in ask-codex only partially documented. The description says 'Codex sandbox mode for this call. Defaults to 'read-only'... Set 'workspace-write' ONLY as an explicit opt-out for flows that need Codex to write files itself, e.g. image generation.' But it does not explain: What is the security implication of workspace-write? Can it modify files outside the working directory? Is it safe to enable for untrusted prompts? Agents need to understand risk before opting in.