MCP server that exposes Google Gemini as tools for Claude Code
The server exposes 4 tools with reasonable descriptions and proper Zod schemas. All tools are named with action verbs (gemini_*), descriptions are present and contextual (avg ~100 chars), and input schemas use Zod with type information. However, there are critical gaps: no output schema documentation, no parameter constraints where enums could apply (e.g., model names), error responses return raw exception messages without recovery guidance, and no idempotency or retry semantics documented. The tools are simple wrappers around the Gemini API with minimal composition, each tool reads from the API and returns structured text output, but the response structure is undocumented for downstream chaining.
Send code or text to Gemini for detailed analysis. Useful for code review, security audit, architecture analysis, or second opinions on complex code.
Ask Google Gemini a question or give it a task. Good for second opinions, large context analysis, or leveraging Gemini's strengths.
Multi-turn conversation with Gemini. Send a conversation history for contextual responses.
List available Gemini models.
Model parameter accepts free-form strings instead of enum of known options (gemini-2.5-pro, gemini-2.5-flash, gemini-2.0-flash). This invites hallucinated model names and API failures.
No output schemas documented for any tool. Code shows responses are structured ({name, displayName, inputTokenLimit, outputTokenLimit} for gemini_models; plain text for others), but LLMs cannot infer this and cannot plan downstream tool calls or data extraction.
Error handling returns raw exception messages ('Error: <msg>') without recovery guidance. No classification of retryable vs. user-fixable vs. fatal errors. Example: API 429 (rate limit) is retried server-side with exponential backoff, but if max retries exhausted, the LLM receives 'Error: ...' with no indication it should retry or backoff.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Distinction between gemini_ask and gemini_chat is unclear. Both send prompts to Gemini; gemini_chat accepts message history as array of {role, text} objects, but the description does not explain when an LLM should choose one over the other. This violates the pattern that similar tools must disambiguate in names and descriptions.
gemini_models description is 27 chars, below the 34-char p10 baseline and well below 50-200 char ideal for LLM-friendly descriptions. It does not state WHEN an LLM should call it or what it enables.
No idempotency semantics documented. If an LLM retries gemini_analyze due to transient failure, will the tool produce identical output or have side effects? Agents need this clarity to manage retries safely.