MCP server that provides multimodal search capabilities (text-to-image, image-to-image) using ChromaDB with OpenCLIP embeddings
The server exposes 2 tools with basic schemas and descriptions. However, both tools suffer from significant definition gaps: descriptions lack depth and context for LLM decision-making, parameter constraints are absent (e.g., no bounds on top_k, no guidance on image encoding), output schemas are documented in docstrings but lack formal JSON Schema structure, and error handling is minimal. Tool naming follows verb_noun convention (search-based), which is positive, but the overall definitions fall well below production readiness for agent deployment. The server is HTTP-based (FastMCP framework), which supports modern transport, but the tool definitions themselves are underdeveloped.
Perform an image to image search using the provided ChromaDB collection.
Perform a text to image search using the provided ChromaDB collection.
Parameter descriptions lack constraints and context. 'image_query' and 'text_query' parameters have minimal guidance on format, encoding, or expected input size. 'top_k' has no bounds (min/max), no default, and no guidance on typical values.
Tool descriptions are generic and under 100 characters. They state WHAT the tool does but lack WHEN to use it (vs. the other search tool), what prerequisites exist (ChromaDB collection must be loaded), or what dependencies exist between the two tools.
No formal output schema in tool registration. Return type is documented in docstrings ('List[Dict]') but not exposed via MCP tool metadata. LLM cannot predict response structure from the registered schema, only from runtime execution or docstring parsing.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 55 | 2025-06-18+ | v2 |
Error handling is minimal. text_to_image_search_tool catches exceptions and returns an empty list; image_to_image_search_tool has no try/catch. Neither tool provides recovery guidance (e.g., 'Collection not found, try initializing ChromaDB first'). Agents cannot distinguish transient failures from logic errors.
No result limits enforced. Tools accept arbitrary top_k values (e.g., top_k=10000) which could bloat response tokens and degrade LLM reasoning. Baseline best practice is to cap results at 20-50 and offer pagination.
Parameter naming could be more explicit. 'image_query' is vague about format, is it base64-encoded PNG, JPEG, or other? The description mentions 'base64-encoded string' but this constraint should be enforced in schema with a pattern or format field.