Model Context Protocol server for Opik (Comet's LLM observability platform).
Server provides 3 tools with complete JSON schemas and reasonable descriptions. However, descriptions lack LLM-optimization guidance (they are functional but generic), parameters lack depth (missing format constraints, ranges, and examples), error handling is not documented, and composition patterns are weak. The tools are read-only discovery/fetch operations, which limits their usefulness for agents requiring state modification. Schema completeness is good, all tools have proper input definitions, but descriptions fall short of the 50-200 character LLM-optimized baseline for A+ tools. Parameter descriptions are present but generic ('Filter by...', 'Optional'). No enum constraints visible, no recovery guidance in docs, and no pagination metadata (e.g., total_count, next_cursor) documented in output schemas.
Paginated discovery and search of Opik entities. Supports filtering, sorting, and searching across traces, spans, threads, experiments, datasets, prompts, and other Opik entity types. Returns pipe-delimited table output with configurable columns.
Fetches any Opik entity by id or name. Accepts opik:// URIs and web links. Returns entity as JSON with optional field projection. Supports time windows for windowed entity types.
Lookup of a write operation's JSON Schema or list reference documentation. Returns schema, example, oauth scope, batch support, parent_id fields, failure modes, and description.
Output schemas are not documented. Tools return data but the response structure is not specified, agents cannot reliably extract fields or plan downstream calls.
Parameters lack formal constraints. 'entity_type', 'operation', time formats, and OQL syntax are mentioned in descriptions but not formalized as enums, regex patterns, or format specs. LLMs must guess valid values.
Error handling and recovery guidance missing. No documented failure modes, no actionable error messages, no guidance for 'not found' cases or malformed input. Agents cannot self-correct.
Pagination metadata not documented. 'list' tool accepts 'page' and 'size' parameters but response schema does not specify total_count, next_cursor, or has_more, agents cannot iterate safely or know when results are truncated.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 41 | 1.26.0+ | v1 |
Tool descriptions lack LLM-optimization. Descriptions are functional but do not answer 'when should I call this vs similar tool?', 'what does it return?', or 'are there prerequisites?'. Average length 138 chars (near baseline 194) but specificity is low.
'list' tool description references 'pipe-delimited table output' but does not document this format, agents expecting JSON must parse text, increasing error rates.