A collection of MCP servers for indexing and searching audio, video, documents, and images using Pixeltable with semantic embeddings
Pixeltable MCP server demonstrates organized structure across 5 server variants (audio, video, image, document, base-SDK) with 22 total tools. However, critical definition quality gaps severely limit LLM usability. Most tools lack input parameter descriptions (e.g., 'random_string' in list_tables tools has empty description ''); parameter schemas are present but many descriptions are generic or missing. Tool names follow verb_noun convention well (setup_*, insert_*, query_*), but descriptions are minimal (10-50 chars) and lack context for LLM selection. No documented output schemas visible in the provided code. Error handling is absent, no recovery guidance, no error categorization, and no invalid input messages. Security concern: openai_api_key exposed as a tool parameter (setup_audio_index, setup_image_index, setup_video_index), violating secret injection pattern. Tools are well-composed (single responsibility), but lack the detail needed for reliable agentic invocation.
Add a computed column to a table in Pixeltable.
Create a named query in Pixeltable.
Create a table in Pixeltable.
Create a view based on a table in Pixeltable.
Execute a query on a table or view in Pixeltable.
Insert an audio file into the specified audio index.
Insert data into a table in Pixeltable.
Three list_tables tools (audio, document, image, video variants) have empty or missing input parameter descriptions. Tool 'list_tables' (audio) has parameter 'random_string' with zero description. This forces LLMs to guess what to pass and violates the rule that every parameter must have a non-empty description.
OpenAI API key exposed as a tool parameter in setup_audio_index, setup_image_index, and setup_video_index. Credentials must never appear as tool parameters, use server-side secret injection via environment variables. Agent traces log every parameter, leaking secrets into logs and prompt history.
No output schemas documented for any tools. LLMs cannot plan downstream tool calls or extract the right data without knowing what fields to expect in the response. Pattern requires documentation of return types and result structures.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 57 | - | v1 |
Insert a document file into the specified document index.
Insert an image file into the specified image index.
Insert a video file into the specified video index.
List all video indexes currently available.
List all audio indexes currently available.
List all document indexes currently available.
List all image indexes currently available.
Query the specified audio index with a text question.
Query the specified document index with a text question.
Query the specified image index with a text description.
Query the specified video index with a text question.
Set up an audio index with the provided name and OpenAI API key.
Set up a document index with the provided name.
Set up an image index with the provided name and OpenAI API key.
Set up a video index with the provided name and OpenAI API key.
Tool descriptions are generic and too short (10-45 characters on average). Most lack context for LLM selection or explanation of when to use each tool. E.g., 'Insert an audio file into the specified audio index' does not distinguish insert_audio from insert_document, insert_image, or insert_video.
No error handling guidance visible in source code. Tools provide no recovery hints, error categorization (retryable vs. user-fixable), or actionable error messages. When insert_audio fails, the LLM has no path forward.
Multiple tools named 'list_tables' across different servers (audio, document, image, video). LLMs cannot distinguish between them, they all have the same name. Either namespace them (e.g., list_audio_tables, list_document_tables) or consolidate into a single parameterized tool.
Parameter 'random_string' in list_tables (audio) and 'columns' in create_table lack clear descriptions of expected format or constraints. 'columns' is described as 'A dictionary of column names and types' but does not specify valid type values or structure. LLMs cannot validate their input.
No pagination support visible in tools that return lists (list_tables, query_audio, query_document, query_image, query_video). Large result sets will exhaust context windows. Tools should accept limit/offset and return total_count or next_cursor.
Parameter 'expression' in add_computed_column and 'filter_expr' in create_view require Python expressions. Descriptions state 'refer to other columns using notation table.column_name' but do not specify what Python functions are available, syntax constraints, or how to escape special characters. LLMs will guess and fail.
No indication of which tools are idempotent vs. destructive. WRITE tools (create_table, insert_data, delete operations if they exist) should declare whether repeated calls with the same input produce the same result. Agents rely on this to retry safely.