MCP server for Google Gemini AI - text generation, web search, YouTube analysis
This MCP server provides four tools for Gemini AI integration with good naming conventions (all verb-noun format: gemini_generate, gemini_messages, gemini_search, gemini_youtube). Tool descriptions are present and reasonably detailed (150-250 chars each), explaining WHAT each tool does and WHEN to use it. Input schemas are fully defined with JSON Schema types, constraints (min/max for numbers, minLength for strings), and parameter descriptions. However, output schemas are not documented, the server returns text responses but does NOT specify the structure of those responses, which violates pattern:tool documentation requirements. Error handling exists (handleGeminiError function) but lacks per-tool guidance or actionable recovery steps. Parameter naming is consistent (temperature, max_tokens, top_p, top_k) across all tools, enabling composition. No critical security issues detected, API key is server-side injected via environment variable, not exposed as a parameter.
Generate text using Google Gemini AI with a simple input prompt. This tool sends a prompt to Google Gemini and returns the generated text response. It is ideal for single-turn interactions, creative writing, code generation, analysis, and general AI assistance tasks. Args: - input (string, required): The prompt or question for Gemini - model (string, optional): Model to use (defaults to GEMINI_MODEL env or gemini-flash-latest) - temperature (number, optional): Randomness 0-2 (higher = more creative) - max_tokens (number, optional): Maximum output length - top_p (number, optional): Nucleus sampling threshold 0-1 - top_k (number, optional): Top-k sampling parameter Returns: Generated text response from Gemini. Examples: - "Explain quantum computing in simple terms" - "Write a Python function to sort a list" - "Summarize the key points of machine learning" Note: Each call may produce different results due to model randomness.
Generate text using Gemini with structured multi-turn conversation messages. This tool enables multi-turn conversations by accepting an array of messages with alternating user/model roles. Use this for contextual conversations where previous exchanges inform the response.
Search the web using Google Search with Gemini grounding. This tool searches the web using Google's search capabilities and grounds the response with real, up-to-date information. Responses are generated by Gemini and include citations to the sources used.
Output schemas are completely undocumented. All four tools return responses, but the structure is never formally specified. LLMs cannot plan downstream tool calls or extract fields without knowing what fields are present. This violates the critical pattern:tool-description requirement that every tool must document its return type.
gemini_search and gemini_youtube lack result limits and pagination guidance. The descriptions do not state whether results are capped (e.g., top 10 search results, first 5 min of video analysis). Without explicit limits, LLMs may expect unbounded results, wasting tokens and risking context overflow.
Error handling is generic (handleGeminiError function) and returns brief messages without recovery guidance. E.g., 'Error: Invalid API key' does not guide the LLM to check GOOGLE_API_KEY or retry. Per pattern:recovery-guide, errors must tell the agent what to do next.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 26 | - | v1 |
Analyze YouTube videos using Gemini's multimodal capabilities. This tool downloads and analyzes YouTube videos using Google Gemini's multimodal understanding. You can ask questions about the video content, get summaries, extract key information, and more. Note: Video downloads are temporary and automatically cleaned up.
gemini_youtube parameter 'query' is optional but no default behavior is documented. If omitted, does it summarize the video? Analyze all content? Return metadata only? LLMs need explicit guidance on what happens when optional parameters are not provided.
No documentation of the CHARACTER_LIMIT (50,000 chars) enforced in the code. If a tool's response is truncated due to this limit, the LLM has no way to know, and may assume the response is complete and act on incomplete information.
Model fallback behavior (ACTIVE_MODEL, MODEL_FALLBACK_USED) is silently applied without error reporting. If the configured model is invalid and the server silently falls back to gemini-flash-latest, the user/agent may not know and could expect different behavior/capabilities.