MCP server integrating GEPA (Genetic-Evolutionary Prompt Architecture) for automatic prompt optimization
GEPA MCP server has moderate definition quality with a mix of strengths and significant gaps. Tool naming follows verb_noun conventions well (optimize_*, explain_*, etc.), which is excellent. However, parameter descriptions are sparse or missing in several tools, and output schemas are not formally documented in the source code. Tool descriptions are adequate (range: 150-300 chars), above the 10-char minimum but below the ideal 50-200 char LLM-optimized range. No input validation details, constraint documentation, or error recovery guidance visible in the tool definitions. The tools are READ_ONLY by design, so no destructive operation warnings are needed. Training data expectations (JSON format) are documented in optimize_prompt but not consistently across other tools. Output is returned as JSON strings rather than structured objects, which requires LLM parsing.
Intelligently choose optimization strategy and generate training data as needed. Analyzes the prompt and context to decide between quick vs advanced optimization.
Optimize prompts based on conversation context and user feedback patterns. Uses conversation history to understand what works well and adapts the prompt optimization accordingly. This is particularly useful for mid-conversation prompt improvements.
Analyze WHY the optimization worked. Build meta-knowledge about prompt engineering principles. Makes the system educational, not just magical.
Optimize for multiple criteria simultaneously, not just keyword matching. I can evaluate prompts on dimensions I actually understand.
Optimize a prompt using GEPA (Genetic-Evolutionary Prompt Architecture). This is the core GEPA algorithm from the research paper that uses genetic-evolutionary methods to optimize prompts through multiple generations of variation and selection.
Output schemas not documented. All tools return JSON strings, but the structure and fields are not formally defined in the source code. LLMs cannot predict what fields to expect without reading implementation details.
Parameter constraints and validation rules not documented in descriptions. 'budget' parameter has no min/max bounds stated; 'optimize_for' array has no limit on size; 'domain' and 'task_type' have no enum constraints or valid value guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Optimize prompt with AI-generated training data. If generate_training=True, I (Claude) will create domain-specific training examples before running GEPA optimization.
Quick prompt improvement using GEPA principles with a single optimization cycle. Ideal for fast improvements when you don't have specific training data or need immediate results. Uses the same core GEPA algorithm but with minimal budget.
Apply successful optimization patterns from one domain to another. Learn once, apply everywhere.
JSON input parameters (training_examples, conversation_history, successful_patterns) lack format specifications. The expected JSON schema (required fields, structure, nesting depth) is only described in docstrings, not in parameter metadata. LLMs may construct malformed JSON.
Error handling responses are JSON strings with 'error' field, but no recovery guidance. E.g., 'Invalid JSON in training_examples parameter' tells the LLM what went wrong but not how to fix it or which tool to call next.
Parameter 'task_type' in quick_prompt_improve defaults to 'general' but has no enum documentation of valid options (summarization, analysis, creative, etc. mentioned in docstring but not in schema).
Tool descriptions occasionally reference internal implementation details ('GEPA algorithm', 'genetic-evolutionary methods') rather than user-centric outcomes. Descriptions should emphasize what the agent achieves, not how.