MCP server providing semantic search over OpenRewrite recipes with PostgreSQL backend and embeddings-based discovery
The OpenRewrite MCP server defines 3 tools with explicit schemas and descriptions visible in mcp-server/src/server.py. Tool naming follows verb_noun convention (test_connection, find_recipes, get_recipe) with action-oriented verbs. Descriptions are present and reasonably detailed (194 - 250 chars), within the production baseline of 34 - 392 chars. However, there are significant gaps in schema completeness, parameter descriptions, and output documentation. The find_recipes tool has a sophisticated oneOf schema for flexible input (string or array), but other tools lack proper output schema documentation. Error handling and recovery guidance are absent. Parameter validation constraints are present but descriptions could be more explicit about formats and constraints. The test_connection tool is trivial and contributes minimal value.
Find OpenRewrite recipes based on user intent. Uses semantic search to discover relevant recipes. SUPPORTS MULTI-QUERY: Pass 'intent' as a STRING for single query OR as an ARRAY of 2-5 query variations for batched search with improved recall.
Get detailed documentation for a specific OpenRewrite recipe. Returns complete information including usage instructions, examples, and configuration options.
Test the MCP server connection. Returns status and echoes an optional message.
Output schemas are not documented for any tool. LLMs cannot plan downstream calls or extract structured data without knowing what fields to expect. This violates the pattern:tool requirement and increases hallucination risk.
get_recipe has minimal description (68 chars) that lacks actionable context. Does not explain when to call it vs find_recipes, what format documentation is returned in (markdown? structured JSON?), or whether it requires exact recipe IDs. Description should be 50 - 200 chars with explicit dependency hints.
find_recipes parameter descriptions lack explicit format guidance. The 'min_score' description states '0.0 to 1.0' but does not explain what scores mean semantically (e.g., 'cosine similarity of query embedding to recipe embedding, where 1.0 is perfect match'). This forces LLMs to guess.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 55 | - | v1 |
test_connection tool is a no-op that adds no value. It only echoes a message and reports connection status, which should be implicit in successful tool invocation. Suggests a placeholder or testing artifact left in production. This wastes LLM reasoning cycles.
No error handling or recovery guidance documented. What happens if a recipe_id is malformed? What if semantic search returns zero results? The tool descriptions do not tell LLMs what to do next, violating pattern:recovery-guide.
find_recipes has a oneOf schema allowing both string and array for 'intent', with detailed description of multi-query support. However, no documentation of how results differ between single vs. multi-query (e.g., does it return a union of matches, or re-ranked results?). LLMs cannot reason about which mode to use.
No pagination or result-limiting guidance for find_recipes beyond the 'limit' parameter (max 20). Does the tool return a total_count? A next_cursor for pagination? The description does not document output structure, forcing LLMs to infer or fail on large result sets.