The merge gate for AI-written code, with signed, replayable attestation. Works with Claude Code, Codex, Cursor, and Gemini CLI.
This server has 13 tools with significant definition quality gaps. Tool descriptions are present but often generic and lack actionable context for LLM tool selection. Input schemas are visible for all tools and include parameter types, but parameter descriptions are minimal. Output schemas are not documented anywhere in the provided source code. Tool naming follows verb_noun patterns reasonably well (cost_analyze, security_audit, obs_status), but several tool names are vague (obs_logs could be get_logs or search_logs; design_extract_tokens lacks clarity on what 'extract' means in context). Error handling is not visible in the source excerpts. The server implements READ_ONLY and WRITE risk classifications, which is good governance practice, but this does not translate to LLM-facing error recovery guidance. No evidence of parameter constraints (enums, min/max, patterns) beyond basic type declarations. No evidence of pagination support for tools that likely return large result sets (obs_logs, release_plan). Tool composition is reasonable (separate concerns), but critical chaining fields (e.g., what does cost_analyze return that cost_optimize might need?) are not evident.
Analyze project costs by scanning infrastructure files.
Find cost optimization opportunities in a project.
Extract design tokens (colors, spacing, typography) from project files.
Generate documentation from source code and comments.
Validate documentation for completeness and consistency.
Search system and application logs.
Get live system metrics: CPU, memory, disk I/O, network from /proc.
No output schemas documented anywhere in source code. Tools return results but LLMs have no way to know what fields to expect or which values can chain to downstream tools. This violates the fundamental pattern:tool-schema requirement.
Parameter descriptions are minimal or generic. 'Path to project root (default: '.')' appears for multiple tools but does not explain WHAT the tool will do with that path, WHEN to provide an absolute vs relative path, or error cases if the path is invalid or not found. Descriptions under 20 chars for several parameters.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 61 | 2026-07-28+ | v2 |
Get system health status: disk space, memory, running services, uptime.
Create a git-based release plan from commit history and tags.
Get current deployment status from file-based deploy tracker.
Conduct a security audit of a project: dependency vulnerability scan, anti-pattern detection, and secret detection.
Generate test skeletons for a project using AST parsing (Python) or regex (JS/TS).
Detect test framework and run tests in a project.
No enumeration constraints visible. 'log_type' for obs_logs accepts 'syslog', 'app', 'all' but these are documented as free-form strings in the schema description, not as a constrained enum. obs_metrics.interval and obs_metrics.samples lack min/max bounds, an LLM could pass interval=0 or samples=1000000.
No pagination support evident for tools returning potentially large result sets. obs_logs returns up to 100 entries by limit parameter, but no 'total_count' or 'next_cursor' field is documented. If an agent needs all logs, how does it discover and retrieve subsequent pages?
Tool names lack clarity in intent. 'obs_logs', 'obs_status', 'obs_metrics' use 'obs' prefix which is ambiguous, is it 'observe', 'observation', 'observatory'? 'design_extract_tokens' uses 'extract' which could mean pull from code, parse CSS, read JSON, the description clarifies it's design tokens but the name alone does not. Consider get_logs, get_system_health, get_system_metrics, extract_design_tokens_from_project.
Tools marked WRITE (test_generate, docs_generate) lack confirmation or dry-run patterns. An agent could accidentally generate boilerplate tests or docs and overwrite existing ones. No dry-run parameter, no preview mode, no confirmation guidance for destructive actions.
No error handling guidance visible. Tools do not document what errors can occur (file not found, permission denied, invalid syntax) or how LLMs should recover. Pattern:recovery-guide requires error responses to tell agents what to do next.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present in schema definitions. While risk classification (READ_ONLY, WRITE) exists in comments, these are not machine-parseable or used to guide agent behavior. The current spec (2026-07-28) expects toolAnnotations in capability negotiation.