Media Knowledge Base MCP - Semantic search over video transcripts and other media content. Provides tools for ingesting YouTube videos, managing a searchable knowledge base, and performing hybrid semantic/keyword search.
Server has 4 well-named tools with generally solid descriptions and explicit schemas. Naming follows verb_noun convention (process_video, manage_source, explore_library, search). Descriptions are LLM-optimized (100-400 chars range). All parameters have type declarations and descriptions. However, output schemas lack formal documentation (return types inferred from code comments rather than explicit schema definitions). Error handling is basic (try/catch with error strings, no recovery guidance). Some parameter relationships undocumented (e.g., source_id required only when view='source'). No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Overall solid foundational quality with room for schema formalization and error enrichment.
Explore and browse the knowledge base metadata. This is the universal "lookup" tool for browsing sources, viewing details, listing tags, or checking statistics. Use different 'view' modes for different types of lookups.
Update a source's tags and/or summary. Returns the updated source. This is the universal "edit" tool - modify tags, update summaries, or both in a single call. The updated source is returned so you can confirm changes without a separate lookup.
Process a YouTube video and add it to the knowledge base. Extracts transcript, generates embeddings, and stores for semantic search. Optionally applies tags and summary in one atomic operation. Safe to call multiple times - existing videos are skipped.
Search the knowledge base using hybrid semantic and keyword search. Uses a two-stage retrieval pipeline: 1. Hybrid search combining vector similarity and full-text search 2. Cross-encoder reranking for improved relevance 3. HyDE query transformation for semantic bridging (if enabled) Results include relevance scores and YouTube timestamp links for navigation.
Output schemas not formally documented in code. Return types inferred from ProcessResult, Source, SearchResults, LibraryStats classes but no explicit JSON Schema declarations visible in tool registrations. explore_library returns Union[List[Source], Source, List[str], LibraryStats], LLMs need to know the exact shape of each variant.
explore_library has conditional required parameters (source_id required only when view='source') documented in description text but not enforced in schema. Parameter dependencies not formalized.
Error handling returns ProcessResult(success=False, error=str(e)) with unstructured error text. No recovery guidance, categorization, or actionable next steps. Example: a network error and a 'video not found' error both become opaque strings.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 46 | - | v1 |
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). explore_library and search are READ_ONLY but marked only in comments; process_video and manage_source are WRITE but not annotated. LLMs cannot detect safety properties from schema alone.
manage_source description says 'Safe to call multiple times' (idempotent) for process_video but manage_source itself is NOT idempotent (adding same tags twice, or adding then removing, have side effects). process_video is idempotent ('Safe to call multiple times - existing videos are skipped'), this distinction should be encoded in annotations.
process_video accepts URL in 'any format' but no validation or error message documented for invalid URLs. manage_source references 'source_id (e.g., "dQw4w9WgXcQ")' but no format spec provided (is it YouTube video ID? UUID? How long?).
explore_library filter_tags parameter says 'at least one matching tag' but unclear if this is AND or OR logic across multiple tags. E.g., filter_tags=['python', 'ml'], returns sources with python OR ml, or both?