New Relic MCP server for observing AI coding assistants (Claude Code, Cursor, Windsurf, Copilot, Codex, and more)
New Relic Preflight demonstrates solid tool definition quality across 23 analytics and observability tools. All tools have clear, descriptive names following verb_noun patterns (get_*, report_*, subscribe_*). Descriptions are substantive and contextual, ranging from 100-300+ characters and explaining WHEN and WHY to use each tool. Input schemas are well-formed with proper JSON Schema types and descriptions. However, there are notable gaps: (1) Most tools have empty input schemas (no parameters) despite sophisticated domain logic, limiting agent ability to filter/customize results, a significant composition issue. (2) Output schemas are not documented in the provided source, forcing LLMs to infer response structure. (3) No explicit error handling guidance or recovery patterns visible in tool definitions. (4) Some parameter descriptions lack format constraints (e.g., ISO week format in nr_observe_get_weekly_summary is documented in description text but not as a formal JSON Schema constraint). (5) Tools like nr_observe_get_session_history and nr_observe_get_trends accept optional filter parameters (developer, since, weeks), but these are well-described and properly typed, this is a bright spot. Overall, the server excels at naming and description but falls short on schema completeness and error recovery guidance.
Get current AI spend vs. configured budget caps (session, daily, weekly). Returns remaining budget, % used, and any threshold alerts fired this session.
Get the impact report for the most recent change to the project's instruction file (CLAUDE.md, or the active platform's equivalent — e.g. .cursorrules on Cursor): before/after comparison with deltas and verdict.
Get a developer's collaboration profile: specificity, autonomy, correction rate, task complexity, and classification.
Get context window efficiency metrics: unique vs. repeated file reads, repeated-read ratio, and top re-read files. A high ratio suggests the model is losing context.
Get context window tracking: per-turn token growth, category breakdown (System/Tools/User/Assistant), fill percentage, and per-tool output contribution. Shows which tools consume the most context.
Most read-only tools (15 of 23) have empty input schemas {"type":"object","properties":{}} despite operating on complex, filterable data. Tools like nr_observe_get_latency_percentiles, nr_observe_get_context_tracking, and nr_observe_get_cost_breakdown accept no parameters for filtering by tool type, time range, or category. This forces agents to retrieve entire datasets and filter client-side, wasting tokens and risking context overflow. Per pattern:tool and pattern:constrained-input, every tool should accept parameters that let agents refine results.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 65 | 2026-07-28+ | v2 |
Get a breakdown of session costs by task, model, and efficiency metrics like cost per line of code. Cost figures are Preflight's own list-price estimate unless costRateMultiplier/dataResidencyPremium are configured for an org's contracted rate — see rate_multiplier_applied in the response.
Project AI spending forward based on current session rate. Returns forecast cost for end-of-day, end-of-week, and end-of-session (8h), with a confidence note. When this session was resumed after being stale, also includes resumeContext: Claude Code's own report of how long it had been and its estimated cost to re-warm the prompt cache — context for a cost spike the rate-based forecast alone would not explain.
Get cost breakdown by outcome type (bug fix, feature, refactor, etc.) with waste ratio and ROI estimate.
Get p50/p95/p99 latency percentiles for tool calls, broken down by tool type. Use to identify which tools are slowest in the current session.
Get a data-driven model recommendation ranked by historical efficiency score, cost, and task success rate across past sessions, both overall and broken down by task outcome type (bug_fix, feature, refactor, investigation, configuration, documentation, failed_attempt).
Get per-model usage statistics: request counts, token totals (input, output, thinking, cache read, cache creation), cost, and cost per million billed tokens. Identifies the most-used model, plus any PostModelSwitch events (deliberate /model changes, automatic fallbacks, or resume) recorded this session.
Returns a narrative coaching report comparing this week's personal AI coding metrics against your historical baseline. Includes highlights, regressions, streaks, and a top recommendation. Requires at least 2 weeks of session history. Returns status: "insufficient_data" with a message when history is too sparse.
Compare AI coding assistant platforms side-by-side on a given metric: efficiency, cost, task_success, tool_calls, or error_rate. Each platform bucket is tagged with visibility_level (full-hooks, self-reported, or mcp-tools-only) — platforms differ in how much built-in tool activity Preflight can observe, so a caveat is included when compared platforms span more than one level.
Get personalized optimization recommendations covering cost, efficiency, prompt engineering, CLAUDE.md impact, and model selection.
Get a list of past sessions with summary metrics (efficiency, cost, tool calls, outcome).
Get task lifecycle metrics: completed task count, average task duration, and average tool calls per task.
Get aggregated AI coding cost and efficiency metrics for all developers in the configured team, queried via New Relic NRQL. Requires teamId to be set in config.
Get trend data for a metric over time: weekly efficiency, cost, task success rate, or tool call counts.
Get a weekly summary report with per-developer breakdown, cost, efficiency, and anti-pattern counts.
Report token usage for cost tracking. Call periodically to enable accurate cost metrics. Provide the model name and token counts from the most recent API response.
Generate the current weekly AI coding summary and POST it to the configured Slack webhook immediately.
Register a Slack webhook URL to receive weekly AI coding cost and efficiency summaries.
Remove the registered Slack webhook for weekly digests.
Output schemas are not documented in tool definitions. Descriptions state what data is returned (e.g., 'context window efficiency metrics: unique vs. repeated file reads...') but the formal JSON response schema is not visible in the MCP tool registration. Per pattern:tool, LLMs need structured output documentation to plan downstream calls and extract fields correctly.
No visible error handling guidance in tool definitions. Descriptions do not explain what happens on failure (e.g., 'Requires at least 2 weeks of session history' in nr_observe_get_personal_insights states the prerequisite but not what error is returned or how to recover). Per pattern:recovery-guide, error responses must be actionable for the LLM.
Parameter format constraints are embedded in descriptions rather than formally declared in JSON Schema. E.g., nr_observe_get_weekly_summary accepts 'week' as 'ISO week (e.g., "2026-W16")' but no pattern regex is defined. nr_observe_get_session_history accepts 'since' as 'ISO date string (e.g., "2026-04-01")' but no format constraint. Per pattern:constrained-input, format constraints should be formal (pattern, format, enum) not just prose.
Write-operation tools (nr_observe_subscribe_digest, nr_observe_send_digest) lack confirmation/dry-run patterns. Per pattern:confirmation-request, irreversible operations (sending digests, subscribing webhooks) should support a dry-run or explicit confirmation to prevent accidental state changes.