Persistent visual cache for LLM-driven software development. Caches screenshots using perceptual hashing, vector search, and AX trees to prevent token overhead and visual hallucination loops.
The Vision Memory MCP server presents 15 tools with domain-specific functionality for visual state management, but suffers from significant definition quality gaps that would impede reliable LLM tool selection and execution. While tool names follow verb-noun conventions and schemas are present, parameter descriptions are sparse, output structures are undocumented, and error handling guidance is absent. The server demonstrates ambition in its feature set (caching, vector search, video analysis) but lacks the rigor expected of production-grade agent tools. Average per-tool score: 42/100.
Ingest screenshot(s) (single/batch via items), lookup cache, return layout description & grounded elements
Compare visual states structurally (layout diff) or compare video recordings (video_a_id/video_b_id)
Create cryptographic, multi-modal evidence pack linking video keyframes, tasks, and visual proof
Export multimodal visual transitions and joint workflow trajectories (json, llava, qwen2_vl, joint)
Purge a specific state and vector embedding from storage for privacy
Find path between visual states using BFS navigation graph
Multiple tools conflate distinct operations (manage_snapshot, manage_visual_spec, manage_video) under a single action enum. This violates the single-responsibility principle and forces LLMs to reason about which parameters apply to which action, increasing error likelihood.
Overloaded and ambiguous parameters (compare_states: state_a_id vs video_a_id with no mutual-exclusivity guidance; recall_memory: query accepts both text and base64 with no format disambiguation). LLMs will struggle to choose the right input format.
No output schemas documented for any tool. LLMs cannot pre-plan what fields to extract or what downstream tools to call. This forces trial-and-error exploration and wastes tokens.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 50 | <=2025-11-25 | v2 |
Fetch aggregated visual context, cache hit ratios, token savings metrics, and server version info
Unified snapshot checkpoints (save, diff, export, restore)
Unified video memory operations for ingestion (ingest), semantic search (search), keyframe timelines (timeline)
Visual SDD design contract baseline registration (set), live verification (verify), and listing (list)
Predict best next UI action and target coordinates based on transition success rates & AX tree
Search visual memory by description query or base64 image query (read-only)
Save UI action execution outcomes, transitions, or log visual blockers (action_type: "blocker")
Revert accidental state or transition edge ingestions
Poll for target visual state until present or timeout occurs
Destructive operations (forget_state) lack confirmation steps or dry-run options. Agents make mistakes, irreversible ops should support MRTR (Multi Round-Trip Requests) for confirmation.
Parameter descriptions are skeletal and often jargon-heavy ('state', 'transition edge', 'AX tree', 'BFS') without domain explanation. LLMs unfamiliar with accessibility testing or visual memory architecture will misuse these tools.
Numeric parameters lack bounds or units (tolerance in manage_visual_spec; timeout_ms and poll_interval_ms in wait_for_visual_state). LLMs may pass invalid values (negative timeouts, absurdly large tolerances).
No error handling or recovery guidance in any tool description. If a tool fails, the LLM has no guidance on what to try next (retry? fallback tool? ask user?).
Tool composition is weak. No chaining IDs documented (e.g., analyze_screenshot output should return state_id for use in record_outcome or get_navigation_paths). LLMs cannot efficiently chain operations.