A zero-dependency MCP server that wraps the Gemini CLI. Speaks JSON-RPC 2.0 over stdio.
This MCP server defines 2 tools with explicit JSON schemas and descriptions. However, several critical quality gaps prevent a higher score: (1) Tool names lack clear action verbs, 'gemini' and 'gemini-reply' do not follow verb_noun pattern and are ambiguous about what action occurs; (2) Tool descriptions are adequate but minimal (34-47 chars), below the 50-200 char baseline for LLM-optimized descriptions; (3) Parameter descriptions are present but sparse, 'The delegation prompt' and 'Session ID returned by a previous gemini call' lack actionable guidance on format, constraints, or error recovery; (4) No output schema is documented, responses return 'content' array and 'threadId' but the agent doesn't know expected structure, field types, or how to chain calls; (5) Error handling is minimal, errors are returned as text in 'isError: true' responses with no guidance on recovery or categorization; (6) Missing safety patterns, both tools accept 'sandbox' and 'cwd' parameters that control process execution, but no permission gates, audit trails, or rate limits are evident; (7) Input validation exists but error messages are terse and don't suggest alternatives or fixes. The code is structurally sound (proper JSON-RPC 2.0 handling, explicit enum constraints on sandbox), but falls short of production-grade documentation and safety patterns expected in Arcade's 54 patterns.
Start a new Gemini expert session
Continue an existing Gemini session
Tool names do not follow verb_noun convention and are ambiguous. 'gemini' and 'gemini-reply' do not clearly signal the action (create_gemini_session, continue_gemini_session would be clearer). LLMs cannot infer from the name alone whether this creates a session, invokes a model, or manages state.
Tool descriptions are below the 50-200 char optimized range. 'Start a new Gemini expert session' (34 chars) and 'Continue an existing Gemini session' (35 chars) lack context on WHEN to use each tool, what 'expert session' means, and what distinguishes this from a direct Gemini API call.
No documented output schema. The code returns {content: [{type: 'text', text: response}], threadId: threadId, isError?: true} but the agent doesn't know: (a) what fields are guaranteed in 'content'; (b) whether 'threadId' is always present; (c) what structure 'response' (the text content) has; (d) how to chain the threadId to gemini-reply. This forces agents to guess or fail on type mismatches.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | 2024-11-05+ | v1 |
Parameter descriptions lack actionable detail. 'developer-instructions' is described as 'Expert system instructions', what format? How long? Can it contain newlines? The description for 'cwd' says 'Current working directory' but doesn't specify: must it exist? Relative or absolute? What happens if it doesn't exist? LLMs cannot infer these constraints.
Error handling is minimal and non-prescriptive. Errors return {content: [{type: 'text', text: 'Error: ...'}], isError: true} with no guidance on whether to retry, whether the error is user-fixable, or what the agent should do next. Example: 'Parse error: No JSON response found' tells the agent nothing actionable. Should it retry? Ask the user? Switch models? This violates the recovery-guide pattern.
Security: No permission gates, audit trails, or rate limits on process execution. Both tools accept 'sandbox' and 'cwd' parameters that control spawned processes. There is no verification that the calling agent has authority to execute Gemini CLI with custom working directories or write access. No logging of who called what, with what params, and what happened. This violates pattern:permission-gate and pattern:audit-trail.
Input validation errors are terse and don't suggest fixes. Example: sendError(id, -32602, "Invalid params: 'prompt' is required") tells the LLM the param is missing but not how to provide it or what format is expected. Should be: "'prompt' is required and must be a non-empty string. Example: 'Analyze this Python code for performance issues.'"
No idempotency or confirmation pattern. Both tools execute processes immediately (spawn Gemini CLI). If an agent retries due to ambiguous failure, the session may be created twice or prompts may be re-executed. No dry-run mode, confirmation step, or idempotency token support. This violates pattern:idempotent-operation and pattern:confirmation-request.
Parameter dependency not documented. The 'model' parameter is only valid for 'gemini' tool, not 'gemini-reply'. The code validates this, but the parameter descriptions in tools/list response don't mention: 'model is only supported by the gemini tool.' Agents may try to pass 'model' to gemini-reply and fail.
No handling of Gemini CLI absence or installation guidance. The error 'Gemini CLI not found. Please install it with npm install -g @google/gemini-cli.' is clear, but this is a fatal dependency error that occurs at runtime. This should be checked at initialize time or logged as a startup warning so clients know this server is non-functional before making tool calls.