A multimodal image research AI assistant with vision analysis and Wikipedia research capabilities
The server has 3 tools with clear action-verb naming (load_, get_, fetch_). All tools include descriptions and visible input schemas. However, descriptions are inconsistent in quality and token efficiency, ranging from 86 chars to 327 chars. Schema definitions are present but lack depth, no output schemas are documented, parameter descriptions are minimal, and error handling provides only basic fallback messages. The tools are well-scoped (single responsibility each), but the implementation shows gaps in LLM-optimal documentation patterns. No tool annotations (readOnlyHint, etc.) are present. The server would benefit from tightening descriptions to 50-150 char range and adding explicit output schema documentation.
Searches Wikipedia for a given query and returns the title, URL, and a brief summary for the top matching articles.
Performs a deep analysis of a Base64 encoded image and returns a detailed, descriptive paragraph about its content. If the image is of a known landmark, it will be specifically identified. This description is intended to be used as a high-quality search query for a research tool.
Loads an image from a server-accessible file path, encodes it to Base64, and determines its MIME type.
No output schemas documented. Tools return data (dicts, strings, lists) but LLM context lacks field descriptions. get_image_description returns a string; load_image_from_path returns {base64_image_string, mime_type, error}; fetch_wikipedia_info returns [{'title', 'summary', 'url', 'error'}]. Without explicit output schema definitions, LLMs cannot reliably extract and chain results.
Parameter descriptions are too brief. 'file_path' has only ~85 chars; 'mime_type' in get_image_description is minimalist (30 chars). Descriptions should explain format expectations, constraints, and examples without example values themselves. E.g., 'The absolute file path to the image (must be accessible from the server; supports JPEG, PNG, GIF, WebP formats)'.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Error handling is inconsistent and non-actionable. load_image_from_path returns {error: "..."} on failure; get_image_description returns a string 'Error analyzing image: ...'; fetch_wikipedia_info returns [{error: "..."}]. Errors do not guide recovery. E.g., 'File not found at path: X' should suggest 'Verify the path is absolute and the server process has read permissions.'
No tool annotations present. Tools are marked READ_ONLY in metadata but lack readOnlyHint in the MCP tool definition. Modern MCP implementations should declare idempotence and safety properties via tool annotation objects for LLM reasoning.
Inconsistent response shapes. load_image_from_path returns a dict with optional 'error' key; get_image_description returns a plain string; fetch_wikipedia_info returns a list of dicts with optional 'error' as first element. This forces LLMs to handle 3 different error patterns, increasing failure risk.
fetch_wikipedia_info does not document max results cap. 'num_articles' defaults to 1 and accepts any integer. Should state: 'The number of articles to return (1 - 10; defaults to 1). Requesting >10 may exceed context limits.'
load_image_from_path returns generic MIME type 'application/octet-stream' as fallback. While not wrong, it loses information. Should attempt format detection from file magic bytes or return an error guiding the user to pass MIME type explicitly.