Spring Boot starter for MCP vector search tools with pgvector support, multi-provider embeddings (Ollama/ONNX/OpenAI), and conversation/docs/code ingestion with AST-aware chunking.
This is a Java/Spring-AI MCP server providing vector search and semantic indexing across conversations, documentation, and source code. Strong points: all 6 tools have clear action-verb names (embeddings_*) and good descriptions (average ~180 chars). Parameter schemas are present with types and descriptions. Weaknesses: no output schemas documented (critical gap), parameter constraints missing (e.g., threshold 0.0-1.0 not enforced as range), error handling guidance absent, no tool annotations (readOnlyHint/destructiveHint). The server correctly identifies 5 read-only tools and 1 write tool (embeddings_reindex) but does not formally declare this via annotations. No evidence of idempotency patterns or recovery guidance.
Starts background re-indexing. Type: 'conversation' for Claude conversations, 'docs' for markdown documentation, 'all' for both. Indexing is incremental (only new/modified files) and asynchronous: returns immediately with status 'started'. Use embeddings_stats to monitor progress.
Semantic search in the vector store with MMR reranking for diversity. Finds documents similar to the provided text, eliminating redundancy. Supports type filters (conversation, docs). Uses adaptive-k to optimize result count.
Semantic search in indexed source code (Java, Go, Python, TypeScript, C, Rust) with MMR. Returns relevant code chunks with file path, language, function/class name, and line numbers. Uses AST-aware chunking: each result is a complete function or method.
Semantic search in Claude Code conversations with MMR for diversity. Returns chunk text, session_id, turn, and source file.
Semantic search in infrastructure documentation (CLAUDE.md, README, docs/) with MMR. Returns relevant sections with file_path and heading, diversified by content.
No output schemas documented. Users cannot predict what fields embeddings_search returns (e.g., does it return 'similarity_score', 'score', 'distance'? Are results paginated?). This violates the documented output schema requirement and forces LLMs to guess field names.
Parameter constraints not enforced in descriptions. 'threshold' is described as 0.0-1.0 but no type range is declared; 'topK' has default=5, max=20 in description but these are not schema constraints (minItems/maxItems). LLMs cannot reliably enforce these limits.
No error handling guidance. Tools do not document what failures look like (e.g., vector DB down, no matching documents, invalid embedding model). LLMs cannot recover from errors without explicit error response documentation.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
Shows statistics on indexed embeddings: count by type, number of files and chunks, last update. Includes chunk_versioning with current version, stale files to migrate, and distribution by version. Also includes ongoing reindex status (if active).
No tool annotations. The server correctly identifies read-only vs. write operations (5 read, 1 write) but does not expose readOnlyHint/destructiveHint/idempotentHint in tool metadata. LLM cannot optimize planning or safety (e.g., retry behavior) without this.
No pagination or result limits documented for search tools. embeddings_search defaults to topK=5 but upper bound (max=20) is only in description text, not enforced. Large result sets could blow token budgets. No next_cursor or continuation token pattern documented.
embeddings_reindex is a write operation that starts async background work ('returns immediately with status started'). No idempotency guarantee documented, no dry-run option, no confirmation before reindex. If called twice rapidly, could cause race conditions. Missing idempotent-operation pattern.