The server provides 12 well-named tools with mostly complete input schemas and reasonable descriptions. Naming is strong (verb_noun pattern, clear intent). Descriptions are present for most tools and generally adequate (median ~120 chars, within target 10-1024 range). However, several tools lack parameter descriptions (glucose_readings array in tool #8 has no description), and output schemas are entirely undocumented, the LLM cannot know what fields to expect from tool responses. Error handling is present but generic. Tool composition is sound (each tool does one thing), but the server lacks guidance on tool chaining and dependency hints. Schema quality is acceptable (JSON Schema present, types defined for most parameters), but parameter descriptions are inconsistently applied.
Output schemas are completely undocumented. No tool specifies what fields, types, or structure the LLM should expect in responses. This prevents the LLM from planning downstream tool calls and extracting relevant data.
Parameter glucose_readings in peloton_analyze_glucose_correlation has no description. The array structure, element schema, and expected timestamp format are undocumented, forcing the LLM to guess the correct format.
Document output schemas for all 12 tools. For each tool, specify the response object structure: field names, types (string, number, object, array), and what each field contains. Example for peloton_get_profile: '{ user_id: string, name: string, created_at: ISO 8601 date, workout_count: number }'
Add description to glucose_readings parameter in peloton_analyze_glucose_correlation: 'Array of glucose readings with timestamps. Each element must be an object with { timestamp: ISO 8601 string, value_mg_dl: number }.'
Create a comprehensive error-handling guide. For each tool, document: (a) Common failure modes (invalid token, rate limit, not found), (b) HTTP status code or error code returned, (c) Actionable guidance for the LLM (e.g., 'If auth fails, call peloton_refresh_token with new credentials'). Return errors in a structured format: { error: string, code: string, suggestion: string }.
Rename peloton_refresh_token to peloton_bootstrap_oauth_tokens to align with the actual functionality (storing/bootstrapping tokens, not refreshing). Update description to clarify this is a one-time setup operation.
Clarify response_format parameter. Add to tool descriptions: 'Use response_format=json if you need to extract structured data for downstream tool calls; use markdown for human-readable summaries.' Document what the JSON structure looks like.
Add dependency hints to peloton_analyze_glucose_correlation: 'To analyze a specific workout, first call peloton_get_workouts to retrieve workout IDs, then pass the ID to this tool.' Similar hints for tools that require data from other tools.
Score history
Overall score trend
↑ 37 points across a rubric change (v1 → v2)
79/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
B
79
2026-07-28+
v2
2026-03-09
F
42
-
v1
read only
auth
source verified
80/100
Calculate muscle impact scores from recent workout history
peloton_refresh_tokenwritesource verified78/100
One-time bootstrap: store Peloton OAuth access + refresh tokens extracted from the browser (Network tab on members.onepeloton.com, auth/session or oauth/token response). After bootstrap, Auth0 refresh keeps tokens alive automatically. Access token must start with "eyJ".
Get aggregate workout statistics (duration, calories, disciplines) for a date range
No error handling documentation. Tools do not specify error conditions (API failures, invalid tokens, rate limits) or recovery guidance. If a call fails, the LLM receives no actionable error message to guide retry or fallback strategies.
peloton_refresh_token name is misleading. The description clarifies this is a 'bootstrap' operation that stores OAuth tokens, not a token refresh. Rename to peloton_bootstrap_oauth or peloton_store_oauth_tokens for clarity.
Multiple tools accept response_format parameter (markdown, json) with default='markdown'. The tool descriptions do not explain when to use which format or how the LLM should interpret the response differently. This creates confusion about tool output structure.
Tool descriptions lack dependency hints and chaining guidance. For example, peloton_analyze_glucose_correlation requires a workout_id but does not hint that peloton_get_workouts should be called first to retrieve valid IDs. This forces the LLM to infer tool ordering.
No pagination or result limiting documented for tools that return lists (peloton_get_workouts, peloton_sync_workouts). Without explicit limits and pagination cursors, large result sets could exhaust token budgets.
Destructive tools (peloton_sync_workouts with WRITE risk, peloton_refresh_token which persists credentials) lack confirmation or dry-run capability. No safeguard prevents an LLM from accidentally overwriting credentials or syncing data when not intended.
peloton_sync_workoutspeloton_refresh_token
Implement and document pagination. For peloton_get_workouts, add parameters: limit (default 20, max 100), offset or cursor (for pagination). Return response: { workouts: [...], total: number, offset: number, limit: number }. Document: 'Use offset/limit to retrieve large result sets without exhausting token budgets.'
Add confirmation step for destructive operations. For peloton_refresh_token, add optional parameter confirm=false. When confirm=false, return { action: 'bootstrap_oauth', requires_confirmation: true, summary: 'This will overwrite existing Peloton credentials. Proceed?' }. Require explicit confirm=true to execute.
Implement idempotency keys for peloton_sync_workouts (e.g., idempotency_id parameter) to ensure repeated calls do not create duplicate records. Document in description: 'Safe to retry; use the same idempotency_id to skip duplicate syncs.'
Add type hints to all parameter descriptions. E.g., 'limit (number, 1-100): Maximum workouts to fetch. Defaults to 20.' ensures LLMs understand parameter types even if schema parsing fails.
Consider batch variants or parameterization for correlated operations. If the LLM often calls peloton_get_workouts then peloton_analyze_glucose_correlation for multiple workouts, offer peloton_bulk_analyze_glucose to accept an array of workout IDs in one call.