This server defines 2 tools with reasonable naming conventions (verb_noun pattern: locate_objects, zoom_to_object) and descriptions present in docstrings. However, there are significant gaps in parameter-level documentation, output schema clarity, and error handling guidance. The tools are visible and explicitly registered via @mcp.tool() decorators in src/mcp_vision/server.py, so scoring is not capped for inference. Parameter descriptions exist but lack specificity around constraints and expected formats. Return types are documented via type hints (str, MCPImage) but output schema structure is not formalized. Error handling is minimal, failures return None or generic strings without recovery guidance.
Detect, find and/or locate objects in the image found at image_path.
Zoom into an object in the image, allowing you to analyze it more closely. Crop image to the object bounding box and return the cropped image. If many objects are present in the image, will return the 'best' one as represented by object score.
Parameter descriptions lack constraint specification. 'image_path' is described as 'path to the image' but does not specify: local filesystem paths vs URLs, file format restrictions (JPEG/PNG/etc), file size limits, or resolution limits. LLMs will pass invalid paths without guidance.
Output schema not formally documented. 'locate_objects' returns a string representation of a list of dicts, not a structured object. 'zoom_to_object' returns MCPImage (a base64-encoded image), but the structure and format are not described in the tool docstring. LLMs cannot plan downstream tool calls or parse the output reliably.
Error handling provides no recovery guidance. When 'locate_objects' finds no objects, it returns a string 'No objects were located in the image.' When 'zoom_to_object' finds no matching label, it returns None. Neither response tells the LLM what to do next: try a different label? Check the image path? The agent has no actionable next step.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Inconsistent return types. 'locate_objects' returns str; 'zoom_to_object' returns MCPImage | None. Returning strings with embedded data structures (e.g., repr of list) instead of structured objects violates pattern:response-shaper. LLMs must parse string representations, wasting tokens and introducing parsing errors.
'hf_model' parameter is optional with a default, but the default value is embedded in the description string rather than declared in the schema. If the default ever changes, the description becomes out of sync. Use JSON Schema 'default' field and reference it in the description.
'candidate_labels' parameter is required but has no constraint on the number of labels, label format, or example values. Can the LLM pass 1000 labels? Empty list? Special characters? The description should specify min/max length and valid characters to prevent invalid calls.