MCP server for tracking AI coding sessions and measuring developer productivity
Server has 9 tools with complete input schemas and descriptions. Naming follows verb_noun patterns (start_, log_, end_, flag_, get_). Schemas are well-structured with type definitions, enums, and required fields. However, descriptions are generic and lack LLM-optimization guidance (e.g., 'when to call this tool instead of similar ones', 'what does it return'). Several tools lack actionable error guidance. Output schemas are not explicitly documented. Some parameter descriptions are minimal. The server implements stateless STDIO-based tracking without remote accessibility, limiting production deployment options.
Complete an AI coding session and calculate ROI metrics
Report a problematic AI interaction for quality tracking
Get list of currently active AI sessions
Retrieve analytics and reports on AI session effectiveness and ROI
Get statistics on logged AI requests
Log an AI prompt/response in the current session with effectiveness rating
Log a standalone AI request/response without a session context
Tool descriptions lack LLM-optimization guidance. Descriptions are 40-80 chars but lack WHAT-WHEN-WHY structure. For example, 'Log an AI prompt/response in the current session with effectiveness rating' doesn't explain when to use this vs log_ai_request, or what the return value is. Baselines show A+ tools average 194 chars with explicit context about selection criteria.
Output schemas not documented. Tools like get_ai_observability and get_active_sessions lack documented return types. LLMs cannot plan downstream calls or extract correct fields without knowing the response structure. Rubric requires '100% of A+ tools have documented return types'.
Error handling lacks recovery guidance. No error responses documented that tell LLMs what to do next (e.g., 'If session_id not found, call get_active_sessions() first'). Error messages and classifications absent from tool definitions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Calculate and store code quality metrics for a session
Start a new AI coding session for tracking workflow metrics. CALL THIS FIRST at the start of every coding task. Returns session_id for subsequent tool calls.
Some parameter descriptions are minimal or incomplete. E.g., outcome enum values in end_ai_session lack explanation of what each status means for ROI calculation. 'estimate_source' in start_ai_session describes enum values but not why this metadata matters. Baselines show 100% of A+ tools have detailed param descriptions.
No pagination guidance for list tools. get_active_sessions and get_ai_observability return lists but lack limit, offset, page_size, or next_cursor parameters. Large datasets could blow LLM context windows. Rubric requires 'Tools returning lists should accept page/offset and limit parameters'.
Overlapping tool responsibilities. log_ai_interaction and log_ai_request both track AI exchanges but with different schemas and scoping. LLMs must reason about which to use. Rubric pattern:tool states 'Each tool should do exactly one thing'.
Parameter 'context' in start_ai_session has default empty string, but unclear if optional params should be omitted from required array or rely on defaults. start_ai_session marks developer and project as required but also have defaults, inconsistent semantics.