State-of-the-art depth estimation for portion size calculation using Depth Anything V2 model, designed for KAI Portion Agent to estimate food portions from images with Nigerian food-specific optimizations
This server exhibits moderate to significant gaps in definition quality. While tools are registered with basic descriptions and some schema information is present, there are critical deficiencies in parameter documentation, schema completeness, and error handling guidance. Most tools lack comprehensive descriptions of what they return, and several parameter descriptions are generic or incomplete. The server implements custom Pydantic models (DepthResponse, PortionEstimate, BatchPortionResponse) which provide some output schema documentation, but tool definitions themselves lack structured error guidance. Naming is generally acceptable (verb-noun pattern present), but descriptions fall short of the 50-200 character LLM-optimized sweet spot for several tools. The most significant issue is the absence of explicit tool annotations (readOnlyHint/destructiveHint) and minimal parameter constraint documentation.
Estimate portion sizes for multiple foods in one image using bounding boxes
Estimate depth map from uploaded image (File Upload)
Estimate depth map from base64 encoded image
Estimate portion size from uploaded image with optional food type and reference object
Estimate portion size from base64 encoded image with optional food type and reference object
Health check endpoint - returns status without requiring model to be loaded
Readiness check - only returns healthy when model is loaded
Missing output schema documentation in tool responses. While Pydantic models exist (DepthResponse, PortionEstimate), the MCP tool registration does not explicitly document return field types and descriptions. LLMs cannot reliably infer what fields downstream tools receive.
Parameter descriptions are incomplete or overly generic. Examples: 'Image file (JPEG, PNG)' lacks guidance on supported resolutions, max file size, or quality requirements. 'Type of reference object for scale calibration' does not enumerate valid values or explain consequences of omission.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No enum constraints on parameters with restricted vocabularies. 'food_type' and 'reference_object' accept free-form strings (e.g., 'jollof-rice', 'hand', 'plate', 'spoon'). LLMs will hallucinate unsupported values. Should declare enums: food_type: [jollof-rice, egusi-soup, ...], reference_object: [hand, plate, spoon, ...].
Tool definitions lack error recovery guidance. No description of what happens on invalid image formats, missing bboxes, or failed depth estimation. Error responses are not documented as retryable vs. fatal. Agents cannot plan recovery steps.
Tool annotations (readOnlyHint, idempotentHint) are absent. This server's tools are all safe (no state mutation), but explicit annotations enable better agent planning and safety verification. All tools should declare readOnlyHint: true.
No documented constraints on numeric parameters. 'reference_size_cm' lacks min/max bounds. What is valid: 1 cm, 100 cm, 1000 cm? 'page_size' ranges (1 - 100) pattern absent. Unbounded numbers invite LLM hallucinations of absurd values.
Duplicate tool functionality without clear distinction. 'estimate_depth' vs 'estimate_depth_base64' and 'estimate_portion' vs 'estimate_portion_base64' perform the same logic with different input formats (file upload vs base64). This wastes LLM reasoning cycles. Consider a single tool with input_method parameter or consolidate input handling.
Output response fields not aligned with parameter naming. Tools accept 'reference_size_cm' but response uses 'reference_object_detected' (boolean). What does the agent receive as the actual detected size? 'volume_ml' and 'portion_grams' are clear, but 'confidence' lacks documentation on scale (0-1? 0-100?) and interpretation.
Missing dependency documentation. 'batch_estimate_portion' requires a list of bboxes and food_types. Are these required? What happens if list lengths mismatch? Is there a maximum batch size? Should agents discover available food_type enums before calling? No guidance provided.