MCP server wrapping the WHOOP REST API for AI assistant access to health and fitness data
The WHOOP MCP server defines 16 tools with clear names following verb_noun convention (get_*, compare_*, analyze). Descriptions are present for all tools and range from 60-90 characters, meeting the baseline average of 194 chars but on the concise side. All tools declare input schemas with proper JSON Schema typing. However, several critical gaps reduce the score: (1) Parameter descriptions are minimal or missing detail about valid ranges, formats, and constraints; (2) No documented output schemas, LLMs cannot plan downstream calls without knowing what fields are returned; (3) Error handling is not visible in the tool definitions; (4) No tool annotations (readOnlyHint, destructiveHint) despite all tools being READ_ONLY; (5) Some parameter names lack consistency (e.g., 'id' without type suffix in get_cycle_by_id). The server follows basic composition patterns (separate tools for distinct queries, pagination support via limit+nextToken), and all tools accept natural-language time parameters ('today', 'last 7 days'), which aligns with chat-data-model principles. Tool descriptions are functional but lack the LLM-optimization guidance found in A+ baselines.
Compare health metrics between two time periods
Retrieve user's baseline metrics for recovery, sleep, and strain
Get user's body measurements (weight, body fat, muscle mass)
Retrieve a calendar view of health data for a month
Retrieve a specific physiological cycle record by ID
Fetch physiological cycle records (strain, calorie data) within a time range
No documented output schemas. LLMs cannot infer what fields are returned by any tool, preventing planning of downstream calls and field extraction.
Parameter descriptions lack constraint details. e.g., 'limit' in get_recovery_collection states '1-25' in description but lacks explanation of default behavior when omitted. 'metric' in get_trend lacks enum values or examples of valid metrics.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 70 | 2026-07-28+ | v2 |
Retrieve authenticated user's basic profile information
Fetch recovery records (HRV, resting heart rate, recovery score) within a time range
Retrieve a specific sleep record by ID
Fetch sleep records (duration, performance, efficiency) within a time range
Calculate and retrieve the user's sleep debt
Get today's recovery, sleep, and strain data
Analyze health metric trends over a specified period
Get a pre-computed summary with averages and trends for a week
Retrieve a specific workout record by ID
Fetch workout records (strain, sport type, calories) within a time range
Missing tool annotations. All 16 tools are READ_ONLY per risk assessment, but readOnlyHint annotation is absent from tool definitions. This prevents clients from optimizing caching, concurrency, and safety guarantees.
Parameter naming inconsistency in by_id tools. 'get_cycle_by_id' accepts 'id' with type 'number', while 'get_sleep_by_id' and 'get_workout_by_id' accept 'id' with type 'string'. No suffix or context distinguishes which resource type. Type mismatch risks agent errors.
Sparse descriptions on by_id tools. 'get_sleep_by_id', 'get_workout_by_id', and 'get_cycle_by_id' have descriptions under 50 characters. LLMs cannot determine when to use these vs collection endpoints without richer context.
Error handling not visible in tool definitions. No evidence of error recovery guidance, categorization (retryable vs user-fixable), or actionable error messages. Agents lack guidance on how to respond to failures.