Laboratory toolchain for code experimentation, benchmarking, and repository analysis with branch swarm strategies, refactor tournaments, release rehearsals, and policy gatekeeper evaluations
Scoring was not performed
Tool names are noun phrases (branch_swarm_lab, narrated_pr_generator) instead of action verbs (run_swarm_lab, generate_narrated_pr). LLMs rely on verb_noun naming to infer intent and distinguish tools with overlapping functionality.
No output schemas documented. The code shows tools generate markdown reports and JSON files, but source code provides no structured schema describing the fields an agent should expect (e.g., leaderboard structure, report sections, JSON field names). LLMs cannot plan downstream tool calls or extract data without knowing response shape.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 24 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 26 | - | v1 |
Input parameter descriptions lack constraint details. E.g., 'config' is described as 'Path to JSON configuration file' but does not explain required JSON schema, valid fields, or format. 'allow_dirty' and similar booleans lack WHEN or WHY guidance. Parameter descriptions should follow pattern:tool-description format: WHAT it controls, WHEN to set it, and valid constraints.
Error handling is minimal and non-actionable. Code raises generic 'RuntimeError: working tree is dirty' and 'ValueError: config must be a JSON object' with no guidance for agent recovery (e.g., 'Try running: git add .' or 'Ensure config is valid JSON'). Per pattern:recovery-guide, errors should tell the LLM what to do next.
Tools write to filesystem paths (report_path) with no documented validation or length limits. No guidance on whether paths are relative or absolute, what permissions are required, or what happens if the directory doesn't exist. Agents calling these tools with malicious paths could cause path traversal attacks.
No pagination or result limits documented. repo_digital_twin accepts max_files (default 1000) and hotspot_limit (default 20) but descriptions do not warn that returning 1000 files could exhaust context or waste tokens. Per mxe:enforce-result-limits, large result sets should be capped and pagination offered.