An MCP server to expose Memvid functionalities to AI clients. Provides video memory encoding, searching, and chat capabilities with Docker lifecycle management.
The server defines 7 tools with basic descriptions and input schemas, but falls short of production quality on multiple fronts. Tool names follow verb_noun convention (good), but descriptions are minimal (average ~50 chars, well below the 194 char baseline). Parameters lack detailed constraints and validation guidance. Output schemas are not documented. Error handling is absent from visible code. The validation_server.py file only shows tool registration signatures, not full implementations; actual error handling, response schemas, and input validation are not visible in the provided source. Per-tool scores average 42/100.
Add text chunks to the video memory encoder
Add PDF file content to the video memory encoder
Add raw text content to the video memory encoder
Build a video file from the accumulated text chunks with optional output path specification
Chat with the video memory using natural language queries
Get the current status of the memvid server
Search the video memory using a query string
Minimal tool descriptions (avg ~50 chars vs 194 baseline). Examples: 'Add text chunks to the video memory encoder' lacks context on when to use, what it modifies, and how results feed downstream tools. LLMs cannot reliably select tools with such sparse docstrings.
Output schemas not documented. Tools return structured data but the schema, required fields, and field types are not declared in visible code. LLMs cannot plan downstream calls or extract the right data without knowing response structure.
Parameters lack format constraints and validation guidance. 'chunks' is described as 'List of text chunks' with no guidance on chunk size, format, or encoding. 'pdf_path' accepts a string but does not specify path validation, allowed directories, or error conditions (file not found, unreadable, corrupted). LLMs cannot determine valid inputs.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 8 | - | v1 |
Error handling not visible in provided source code. No recovery guidance, error categorization (retryable vs fatal), or actionable error messages documented. Agents will receive raw errors with no guidance on next steps.
Destructive operations lack confirmation/dry-run. 'build_video' writes files and 'add_chunks'/'add_text' modify state, but no dry-run or confirmation step visible. Agents should be warned before irreversible side effects.
Batch operation pattern missing. 'add_chunks' accepts an array (good), but tools like 'add_text' only accept a single text string. For typical agent workflows (adding multiple texts sequentially), the server will force N separate calls instead of one batch operation, wasting tokens and latency.
search_memory returns results but no pagination parameters visible (limit is declared, but no offset/page or cursor, no total_count, no next_cursor). Without pagination metadata, agents cannot safely handle large result sets or iterate.
chat_with_memvid accepts 'history' as array but no schema for history items visible. Should define expected structure (role, content, timestamp?) so LLMs format conversation history correctly.
Tool implementation details hidden in memvid library import. Core business logic (state management, caching, encoding) is in 'from memvid import MemvidChat, MemvidEncoder, MemvidRetriever' but not exposed in the MCP server code. Difficult to verify input validation, output filtering, or error handling without seeing the full call chain.