MCP server for natural-language image search over local directories
This MCP server provides 8 image search tools with reasonable naming conventions and descriptions. Most tools follow verb_noun patterns (scan_directory, search_images, get_image_details, find_similar_images, reindex_changed_files, list_indexed_directories, remove_indexed_directory, get_statistics). Input schemas are present and properly typed with required fields and defaults. However, there are notable gaps: (1) descriptions lack WHEN/WHY guidance for tool selection, they state WHAT but not the decision context; (2) parameter descriptions are minimal (one-liner) and lack format/constraint details; (3) no documented output schemas, LLMs cannot see what fields are returned; (4) no error handling guidance in tool descriptions; (5) no examples of parameter relationships or dependencies; (6) some descriptions are generic ('includes its caption, OCR text, tags, EXIF metadata, and file information') without explaining when to use this vs search_images. The tool definitions are inferred from the @register_tools() decorator pattern, explicit registration with output schema documentation is not visible in the source excerpt, capping per-tool scores accordingly.
Find images visually similar to a given image using semantic embeddings. Returns images with similar content, style, or subject matter.
Get detailed information about a specific image, including its caption, OCR text, tags, EXIF metadata, and file information.
Get overall statistics about the image index, including total images, indexed count, images with OCR/captions/embeddings, etc.
List all directories that have been indexed for searching, along with their statistics (number of images, last scan time, etc.).
Reindex files that have been modified since last indexing. Checks file hashes to detect changes. Can target a specific directory or reindex all previously indexed directories.
Remove a directory from the indexed list. Note: this does not delete the indexed data, it just marks the directory as inactive.
No documented output schemas for any tool. LLMs cannot predict what fields are returned, forcing trial-and-error or requiring them to hallucinate field names for downstream processing.
Descriptions lack WHEN/WHY guidance. They state WHAT each tool does but not the decision context for tool selection. E.g., search_images vs get_image_details both access image data but descriptions don't explain the trade-off (summary vs. detailed analysis). LLMs struggle to pick the right tool without this context.
Parameter descriptions are minimal (one-liners) and lack format/constraint details. E.g., 'path' parameters state 'Absolute path to...' but do not clarify handling of symlinks, relative paths, path traversal risks, or whether the path must be indexed first. LLMs may pass invalid inputs.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 35 | - | v1 |
Scan a local directory for images and index them for searching. This will extract text (OCR), generate captions, tags, and embeddings for each image found. Can scan recursively through subdirectories.
Search for images using natural language queries. Searches across image captions, OCR text, and tags. Supports queries like 'find screenshots with Python code', 'images with dogs', 'receipts or invoices', etc.
No error handling guidance in tool descriptions. Descriptions do not tell LLMs what can go wrong (file not found, invalid path, analysis failure) or what to do next (retry, call a different tool, ask the user). This violates the recovery-guide pattern.
Tool definitions are inferred from code patterns (register handler, Tool() object) but no explicit registration block or tool inventory visible in the source excerpt. Per hard-scoring rules, this caps per-tool scores at 50 when tool definitions cannot be directly verified.
Descriptions use vague language like 'etc.' (get_statistics) or generic plurals ('and file information') without specifying what data is returned. LLMs cannot reliably extract field names and types from these descriptions.