MCP Server for Datadog API, enabling interaction with Datadog resources
This server has 10 tools with visible schemas and descriptions. All tools have proper verb-noun naming (get_*, search_*, aggregate_*), which is excellent. Descriptions are present and reasonably detailed (ranging from ~100-250 chars), following the baseline of 194 chars average. However, several critical gaps emerge: (1) Input schemas are defined via Zod in the code but lack explicit output schema documentation, the tools return JSON.stringify(result) with no documented response shape, violating the pattern:tool-schema requirement. (2) Many parameter descriptions reference API concepts (e.g. 'groupStates to filter by monitor status') but do not document constraints, valid values, enums, or ranges, LLMs cannot infer valid inputs without explicit enums. (3) Error handling is absent from tool definitions, no guidance on what happens on failure, how to recover, or which errors are retryable. (4) No tool annotations (readOnlyHint, idempotentHint) are declared despite all tools being READ_ONLY. (5) Search and filtering tools (get-monitors, get-events, search-logs, aggregate-logs) accept free-form string parameters ('tags', 'sources', 'priority') that should be constrained with examples or patterns. (6) The aggregate-logs tool accepts deeply nested 'compute' and 'groupBy' arrays with no concrete examples or validation hints. Overall: solid foundation with good naming and presence of descriptions, but deficient in schema documentation, constraint specification, and error guidance.
Perform analytical queries and aggregations on log data. Essential for calculating metrics (count, avg, sum, etc.), grouping data by fields, and creating statistical summaries from logs. Use this when you need to analyze patterns or extract metrics from log data.
Get the complete definition of a specific Datadog dashboard by its ID. Returns all widgets, layout, and configuration details.
Retrieve a list of all dashboards from Datadog. Useful for discovering available dashboards and their IDs for further exploration.
Search for events in Datadog within a specified time range. Events include deployments, alerts, comments, and other activities. Useful for correlating system behaviors with specific events.
List incidents from Datadog's incident management system. Can filter by active/archived status and use query strings to find specific incidents. Helpful for reviewing current or past incidents.
NO OUTPUT SCHEMAS DOCUMENTED. All tools return text JSON via JSON.stringify(result) with zero documentation of response field names, types, or structure. LLMs cannot plan multi-step chains without knowing what fields are returned (e.g., does get-monitors return 'id' or 'monitor_id'? Does it include 'status' or 'state'?). This violates pattern:tool and mxe:response-field-naming.
FREE-FORM STRING PARAMETERS LACK ENUMS/CONSTRAINTS. Parameters like 'tags', 'sources', 'priority' (get-events), 'query' (get-incidents, search-logs), and deeply nested compute/groupBy specs (aggregate-logs) accept arbitrary strings with minimal guidance. LLMs cannot infer valid values, they hallucinate. 'priority' has an enum (normal, low) but most others don't. The description says 'filter by tag criteria' but doesn't show format (comma-separated? space-separated? key=value?). The 'aggregation' field in aggregate-logs lists examples in description ('count, avg, sum, etc.') but no schema enum, forcing LLMs to guess. Violates pattern:constrained-input.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Retrieve detailed metadata about a specific metric, including its type, description, unit, and other attributes. Use this to understand a metric's meaning and proper usage.
List available metrics from Datadog. Optionally use the q parameter to search for specific metrics matching a pattern. Helpful for discovering metrics to use in monitors or dashboards.
Get detailed information about a specific Datadog monitor by its ID. Use this to retrieve the complete configuration, status, and other details of a single monitor.
Fetch monitors from Datadog with optional filtering. Use groupStates to filter by monitor status (e.g., 'alert', 'warn', 'no data'), tags or monitorTags to filter by tag criteria, and limit to control result size.
Search logs in Datadog with advanced filtering options. Use filter.query for search terms (e.g., 'service:web-app status:error'), from/to for time ranges (e.g., 'now-15m', 'now'), and sort to order results. Essential for investigating application issues.
COMPLEX NESTED SCHEMAS WITHOUT EXAMPLES OR VALIDATION HINTS. The 'aggregate-logs' tool has a deeply nested structure with compute[] and groupBy[] arrays containing objects with multiple optional fields ('aggregation', 'metric', 'type', 'sort'). No concrete examples of valid payloads are provided. Parameter descriptions are minimal (e.g., 'Aggregation function (count, avg, sum, etc.)' without enums). The 'sort' object has nested 'aggregation' and 'order' fields but no guidance on valid 'order' values (asc/desc?). search-logs has a similar issue with nested 'filter' and 'page' objects. These nested structures are hard for LLMs to construct correctly. Violates review:param-validation-rules.
INCONSISTENT TIME PARAMETER FORMATS. get-events expects 'start' and 'end' as Unix timestamps (number type). search-logs and aggregate-logs expect 'from' and 'to' as free-form strings ('now-15m', 'now', ISO 8601?). No explicit format guidance, LLMs must infer whether 'now-15m' is a valid value for search-logs or if it should be a Unix timestamp. This inconsistency across tools forces LLMs to reason about format differences. Violates pattern:tool-description and mxe:natural-identifiers (use consistent, self-documenting formats).
NO ERROR HANDLING GUIDANCE. No tool definition includes guidance on error scenarios: What happens if a monitor ID is invalid? If a dashboard doesn't exist? If a log query times out? Are errors retryable? Should the agent ask the user? The code returns raw Datadog API responses via JSON.stringify() with no error wrapper. If an API call fails, the LLM receives an error message with no recovery guidance. Violates pattern:recovery-guide and pattern:error-classification.
MISSING TOOL ANNOTATIONS (READ-ONLY HINTS). All 10 tools are READ_ONLY, but the tool definitions do not declare readOnlyHint=true. The MCP spec (2026-07-28) includes tool annotations for readOnlyHint, destructiveHint, and idempotentHint. Declaring readOnlyHint=true helps agents understand these tools are safe to call without user confirmation. This is a spec alignment gap, not critical (tools will work), but represents incomplete adoption of current MCP patterns.
PAGINATION NOT CLEARLY DOCUMENTED. Several tools accept 'limit' and some accept pagination params (pageOffset, pageSize for get-incidents; cursor in search-logs). Descriptions do not explain whether results are paginated automatically, what the max limit is, whether there's a next_cursor field in responses, or how to retrieve additional pages. Users/LLMs cannot know if calling the same tool multiple times with different offsets is required or if results are capped. Violates pattern:paginated-result.
METRIC NAME FORMAT NOT SPECIFIED. get-metric-metadata expects 'metricName' (string) but provides no format guidance. Datadog metrics have naming conventions (e.g., 'system.cpu.user', 'datadog.estimated_usage.indexed_logs'). Should it be dot-separated? Case-sensitive? If an LLM passes 'System.Cpu.User' (title case), will it fail? No guidance forces LLMs to guess or make extra calls to validate. Violates review:param-validation-rules.
DASHBOARD ID FORMAT AMBIGUOUS. get-dashboard expects 'dashboardId' as a string but doesn't specify format (UUID? numeric slug? slug-like string?). The get-dashboards tool presumably returns dashboards with IDs, but without output schema, the LLM doesn't know what format to expect or pass to get-dashboard. This can cause chaining failures. Violates mxe:response-field-naming and mxe:include-chaining-ids.