MCP server for querying OpenTelemetry traces from LLM applications with support for multiple backends (Jaeger, Tempo, Traceloop)
Strong tool definitions with comprehensive parameter schemas and clear descriptions. All 11 tools have proper JSON Schema input definitions with typed parameters. Descriptions are substantive (averaging ~120 chars) and explain purpose, filters, and use cases. Tool names follow verb_noun convention consistently (search_traces, get_trace, list_services, etc.). However, output schemas are not explicitly documented in the source, limiting clarity on return structures. Error handling descriptions are minimal, tools do not explain recovery paths or what to do when queries return empty or fail. Security considerations around data access are not documented (no scope declarations). One composition concern: search_traces and search_spans are nearly identical in capability, risking LLM confusion.
Find traces with errors. Including detailed error messages, stack traces, and LLM-specific error information.
Find traces with high token usage (expensive LLM calls). Identifies resource-intensive operations by total token count.
Find traces that took longer than a threshold duration. Useful for performance debugging and identifying slow LLM operations.
Get aggregated LLM usage metrics (token counts) for a time period. Provides breakdowns by model and service.
Get statistics about model usage across traces. Includes counts, token usage, and error rates grouped by model.
Get complete trace details by trace ID. Returns all spans with attributes, including parsed Opentelemetry data for LLM operations.
Output schemas not explicitly documented. Tools return JSON strings but the structure (fields, types, nesting) of those responses is not declared in tool definitions or source comments. LLMs cannot plan downstream processing without knowing return structure.
Semantic overlap between search_traces and search_spans. Both tools accept nearly identical filter parameters (service_name, operation_name, start_time, end_time, duration filters, LLM system/model filters, has_error, tags, filters, limit). The distinction (traces vs. spans) is not explained in descriptions. LLMs will struggle to choose between them, risking wasted calls.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 73 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
List all LLM tools being used by identifying traceloop.span.kind == tool. Discovers which tools/functions LLM applications are calling, grouped by tool name with usage statistics.
List all LLM models detected in traces. Discoveries models being used by your LLM applications across all services.
List all available services in the OpenTelemetry backend.
Search for individual spans with advanced filtering. Supports detailed span-level queries with comprehensive filter operators.
Search for OpenTelemetry traces with filters. Supports both simple parameters and advanced generic filter system.
Error handling guidance missing. Tool descriptions do not explain what happens on empty results, invalid time ranges, backend connectivity failures, or malformed filters. Descriptions like 'Search for OpenTelemetry traces with filters' do not tell LLMs how to recover if the query returns nothing or times out.
No pagination guidance in descriptions. tools accept 'limit' parameters (capped at 100-1000 items) but descriptions do not explain how to iterate over large result sets, whether a next_cursor is returned, or what happens if the user needs results beyond the limit.
Parameter descriptions could be more prescriptive about constraints. E.g., 'start_time in ISO 8601 format' is mentioned, but descriptions do not clarify required vs. optional, behavior if omitted, timezone handling (assumed UTC?), or what happens if start_time > end_time.
No security/permission declarations. Tools do not state what access level is required (e.g., 'read:traces', 'read:llm-usage'). No audit trail guidance on what caller context is logged. Agents cannot reason about least-privilege configurations.
Tool descriptions for list_services and list_models are under-specified. 'List all available services' (16 chars) and 'List all LLM models' (19 chars) barely exceed the 20-char minimum. They do not explain what structure is returned, what the services/models are used for, or when to call them vs. other tools.