Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server has severe definition quality issues across nearly all dimensions. Tool naming lacks action verbs, descriptions are often generic or misleading, parameter schemas are incomplete, and error handling is absent. The codebase shows signs of rushed development (see fix_tools.cjs with profanity comments). Most tools are defined as simple objects without proper handler implementations visible, making it impossible to verify actual behavior. The 15 tools span wildly disparate domains (codebase analysis, business ROI, component generation, licensing) without clear integration or composition. Output schemas are undocumented. Security concerns are significant: license activation and trial signup are WRITE tools but lack permission gates or audit trails.
Tools (15)
mcp__gemini__activate_licensewrite50/100
Activate a license key to unlock premium features
mcp__gemini__analyze_codebaseread onlyauth43/100
Comprehensive codebase analysis with AI insights
mcp__gemini__chat_plusread onlyauth42/100
Advanced collaborative AI chat with automatic model switching and context optimization
mcp__gemini__check_upgraderead only50/100
Check upgrade options and pricing for your current tier
mcp__gemini__debug_analysisread onlyauth48/100
AI-powered debugging assistance
mcp__gemini__financial_impactread onlyauth35/100
ROI analysis and cost-benefit calculations for technical decisions with business impact quantification
CRITICAL: Most tool names do not start with action verbs or are noun phrases (team_orchestrator, thinkdeep_enhanced, planner_pro, license_info, chat_plus, financial_impact). LLMs infer intent from verb-first naming. Vague or noun-based names force the model to read full descriptions before deciding to use a tool, increasing latency and error rate.
CRITICAL: No input or output schemas are documented in the source code. For every tool, the actual response structure is unknown. Handlers are either missing or return untyped strings/objects. LLMs cannot infer field names or types for downstream tool chaining. Agents will guess at response structure, leading to failed chains and wasted context.
Rename all tools to start with action verbs. Examples: 'mcp__gemini__analyze_codebase' ✓ (keep), 'mcp__gemini__team_orchestrator' → 'mcp__gemini__orchestrate_team_workflow', 'mcp__gemini__chat_plus' → 'mcp__gemini__start_multi_model_chat', 'mcp__gemini__thinkdeep_enhanced' → 'mcp__gemini__reason_step_by_step', 'mcp__gemini__planner_pro' → 'mcp__gemini__create_project_plan'.
Document complete input and output schemas for every tool. For each tool handler, return a structured object with typed fields. Example for analyze_codebase: { files_analyzed: number, total_lines: number, language_distribution: {[lang: string]: number}, insights: string, structure: {[path: string]: metadata} }. Make schemas machine-parseable so LLMs can plan tool chains.
Add formal enums to all parameters with constrained values. Examples: framework enum [react, vue, angular, svelte], language enum [javascript, python, java, go, rust], status enum [open, in_progress, resolved, closed]. Remove implicit enums in defaults and make them explicit in schema.
Expand parameter descriptions to 50-100 characters with format and range guidance. Examples: 'Path to codebase root (absolute or relative path, max 500 chars)', 'Team size for estimation (1-100, used for timeline projection)', 'Risk tolerance (low: <$10k impact, medium: <$100k, high: unlimited)'.
Add error handling with categorization and recovery guidance to all tools. Include try-catch blocks that classify errors as Retryable, UserFixable, or Fatal. Return structured error responses: {type: 'user_fixable', message: 'Invalid language. Must be one of: javascript, python, java', next_step: 'Re-call with valid language', available_options: [...]}'.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↓ 15 points across a rubric change (v1 → v2)
38/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
38
<=2025-11-25
v2
2026-03-09
D
53
2024-11-05+
v1
auth
50/100
Generate UI components for React, Vue, Angular, Svelte
mcp__gemini__license_inforead only50/100
Get detailed information about your current license, usage, and features
CRITICAL: WRITE tools (activate_license, start_trial) have no visible permission checks, rate limits, or audit trails. activate_license accepts a license_key parameter as input, vulnerable to brute-force attacks. start_trial accepts email and can be abused to create unlimited trial accounts. No error categorization or recovery guidance. Violates security and compliance patterns.
HIGH: Enumerated values are used as defaults (e.g., framework='react', language='javascript', model_preference='auto') but are NOT declared as formal enums in the JSON Schema. This invites hallucination, LLMs will invent values not in the implicit set. Example: generate_component defaults to 'react' but does not declare an enum of [react, vue, angular, svelte].
HIGH: No error handling or recovery guidance visible in any tool. Missing: error classification (retryable vs user-fixable vs fatal), actionable error messages, and suggestions for next steps. Example: if activate_license fails with 'Invalid key', does the agent retry? Ask the user? Stop? Unknown.
MEDIUM: Most tool handlers are not visible in the provided source code. Only analyze_codebase and fix_tools.cjs snippet are shown. For the remaining 13 tools, tool definitions exist (object with description, parameters, handler key) but the handler implementations are missing or opaque. This violates the 'tool definition must be visible and verifiable' principle. Per HARD SCORING RULES, if tool definitions are inferred rather than directly visible, cap those tools at 50.
MEDIUM: Tool descriptions are vague or marketing-focused rather than task-focused. Examples: 'Advanced collaborative AI chat with automatic model switching' (what does 'automatic' mean? on what criteria?), 'Interactive project planning with templates' (interactive implies multi-turn, but MCP tools are stateless single-call), 'Extended AI reasoning with step validation' (what is 'extended'? how does validation work?). Descriptions should state WHAT, WHEN, and WHY, not brand messaging.
MEDIUM: No tool composition or chaining guidance. Tools are siloed. Example: analyze_codebase returns undefined output structure, so there is no way to pass its results to refactor_suggestions or generate_component. No clear dependencies or ordering. Agents cannot plan multi-step workflows.
LOW: Array parameters like 'load_scenarios' and 'metrics' in performance_predictor, and 'team_members' in team_orchestrator, lack element type definitions. Example: load_scenarios default is ['current','2x','10x'], are strings, numbers, percentages? Should be documented as array of enum or pattern.
Implement permission checks and rate limiting for WRITE tools (activate_license, start_trial). Check user/agent authorization before executing. Log all WRITE operations with timestamp, user, parameters, and result for audit compliance. Reject license keys that fail format validation before calling backend.
Add tool composition guidance. Document which tools should be called in sequence and what data flows between them. Example: 'Call analyze_codebase first to identify refactoring candidates, then pass file paths to refactor_suggestions.' Include chaining IDs in responses (e.g., codebase_analysis_id so refactor_suggestions can reference it).
Reduce the scope of individual tools, many are doing 'too much'. Examples: 'generate_component' should focus on React alone, or split into generate_react_component, generate_vue_component, etc. 'performance_predictor' combines prediction, capacity planning, and recommendations, split into predict_performance and recommend_optimization. Narrower tools are easier to test, compose, and reason about.
Document idempotency for all tools. State explicitly whether repeated calls with same input produce same output or have side effects. WRITE tools should support a dry_run parameter to preview changes before committing.
Add pagination parameters (limit, offset/cursor) to any tool that may return many results. Example: analyze_codebase currently returns all files, add limit (default 50) and offset to handle large codebases. Cap results in description: 'Returns up to 50 files per call; use offset to paginate.'
Remove marketing language from descriptions. Replace 'Advanced collaborative AI chat with automatic model switching' with 'Send a message to start a conversation; the server selects the best model based on message content (coding, analysis, creative). Supports conversation_id for multi-turn context.' Be explicit about mechanics, not benefits.
Validate all input early and return clear, actionable error messages. Example: framework='svelte' gets 'Invalid framework: svelte. Supported: react, vue, angular. Did you mean vue?' Let the LLM self-correct in one retry instead of silent failures or generic errors.
For array parameters, specify min/max items and element type. Example: 'load_scenarios (array of strings, 1-10 items, each from: current, 2x, 5x, 10x, 100x). Defines load multipliers to simulate.'
Add tool annotation hints to the server registration (readOnlyHint, destructiveHint, idempotentHint per MCP spec). This lets clients warn users before executing dangerous tools. Example: activate_license should have destructiveHint=true.
Document dependencies and prerequisites in descriptions. Example: 'Requires: valid license (call license_info first). Team size must match API rate limits.'
Remove empty or trivial defaults. Example: team_size default=1 is rarely useful. Either pick a sensible default (5) or make it required. Defaults should make the common case work without extra reasoning.
Create a tool index/discovery guide. With 15 disparate tools, agents waste tokens deciding which to call. Provide a structured map: {codebase_analysis: [analyze_codebase, refactor_suggestions], code_generation: [generate_component, generate_api], planning: [planner_pro, team_orchestrator], ...}. Help agents understand when to use which tool.