agit provides 20 tools with complete input schemas and descriptions for all tools and parameters. Naming follows verb_noun conventions consistently (spawn_worktree, merge_worktree, claim_task, etc.). However, descriptions are uniformly short (10-20 chars per tool), falling below the 50-200 char baseline for production tools. Parameter descriptions are present but minimal. Output schemas are inferred from code (jsonResult wrapper) but not explicitly documented in tool definitions. No error recovery guidance, no dependency hints, no enum constraints on status filters. The server follows basic definition patterns but lacks the richness and LLM-optimization needed for A-grade production tools. Most tools are READ_ONLY or WRITE, but risk categories are properly annotated.
Tool descriptions are uniformly short (10-20 characters) and lack actionable guidance for LLM selection. Baseline is 50-200 chars. Examples: 'List all registered repositories' (34 chars), 'Get detailed status for a specific repository' (45 chars), 'Create an isolated worktree for an agent' (41 chars). None answer WHAT the tool does in context, WHEN to use it over alternatives, or what prerequisites exist.
Expand tool descriptions from 10-50 chars to 80-150 chars. For each, add: (1) what the tool does, (2) when to use it vs. similar tools, (3) any prerequisites. Example: 'agit_spawn_worktree: Creates an isolated Git worktree for an AI agent to work on a task. Call this after agit_create_task to initialize a branch and working directory. Generates a branch name from the task description unless branch is provided explicitly.'
Add enum constraints to status parameters. In the input schema for agit_list_tasks and agit_list_worktrees, define status as enum: ['pending', 'claimed', 'in_progress', 'completed', 'failed'] and ['active', 'completed', 'stale', 'conflict'] respectively.
Document output schemas for all list and detail tools. Add a comment block or markdown field in tool definitions that specifies the return structure. Example for agit_list_repos: 'Returns array of objects with fields: name (string), path (string), default_branch (string), remote_url (string), worktree_count (int), task_count (int).'
Implement pagination for list tools. Add optional parameters: limit (int, 1-100, default 20), offset (int, default 0). Return responses with total_count and has_more fields. This prevents context explosion and enables large-scale agent workflows.
Add confirmation/dry-run for destructive tools. For agit_remove_worktree and agit_cleanup_worktrees, introduce an optional 'confirm' boolean parameter. If false or omitted, return a preview of what would be deleted without executing. Require explicit confirm=true to proceed.
Score history
Overall score trend
↑ 2 points across a rubric change (v1 → v2)
56/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
56
2026-07-28+
v2
2026-03-09
D
54
-
v1
source verified
68/100
Get detailed information about a specific task
agit_heartbeatwritesource verified65/100
Update agent heartbeat timestamp
agit_list_agentsread onlysource verified72/100
List all registered AI agents
agit_list_reposread onlysource verified72/100
List all registered repositories
agit_list_tasksread onlysource verified68/100
List tasks for a repository
agit_list_worktreesread onlysource verified68/100
List worktrees for a repository
agit_merge_worktreewritesource verified70/100
Merge a worktree branch into the default branch, then auto-cleanup
agit_next_taskwritesource verified75/100
Atomically claim the highest-priority pending task. Returns the claimed task or null if no pending tasks exist.
Mark a claimed task as in-progress and associate a worktree
Destructive tools (agit_remove_worktree, agit_cleanup_worktrees) lack confirmation or dry-run capability. Agents can permanently delete worktrees without a safety check. No guidance in error handling to help agents recover from mistakes.
Status filter parameters (agit_list_tasks, agit_list_worktrees) accept string values without enum constraints. Parameter descriptions say 'Filter by status (pending/claimed/in_progress/completed/failed)' and 'Filter by status (active/completed/stale/conflict)' but do not declare enums, inviting hallucinated invalid status values.
Output schemas are not explicitly documented in tool definitions. Code shows responses are JSON-wrapped via jsonResult(), but tool metadata does not declare the shape of returned objects (e.g., agit_list_repos returns array of repoItem with fields: name, path, default_branch, remote_url, worktree_count, task_count). LLMs cannot plan downstream calls without knowing what fields will be returned.
No pagination support visible in list tools (agit_list_repos, agit_list_agents, agit_list_tasks, agit_list_worktrees). Large result sets could exhaust context or cause timeouts. Best practice is to include limit, offset/cursor, and total_count parameters.
No error recovery guidance in tool definitions or code comments. Code uses apperrors.IsUserError() to distinguish user errors from internal errors, but tool descriptions do not explain what errors are retryable or how to recover (e.g., 'repo not found, call agit_list_repos to see available repos').
No tool chaining hints or dependency documentation. For example, agit_start_task requires both task_id and worktree_id, but does not hint that agit_spawn_worktree should be called first if no worktree exists. Similarly, task lifecycle tools (claim → start → complete/fail) lack ordering guidance.
Enrich error messages with recovery steps. When a repo, task, or worktree is not found, return available alternatives in the error. Example: 'Repository "myrepo" not found. Available repos: ["frontend", "backend", "infra"]. Call agit_list_repos for the full list.'
Add tool chaining hints in descriptions. For tools that depend on others, include a 'Prerequisites' section. Example for agit_start_task: 'Prerequisites: Task must be claimed (call agit_claim_task first) and a worktree must exist (call agit_spawn_worktree if needed).'
Validate and document parameter types more rigorously. Ensure all numeric parameters (priority, limits) have min/max bounds. Example: agit_create_task's priority should specify range (e.g., -10 to 10, default 0).
Consider returning task/worktree IDs from all write operations. Ensure agit_spawn_worktree returns worktree_id, agit_create_task returns task_id, etc., so agents can immediately chain subsequent calls without lookup.
Add tool annotations (readOnlyHint, destructiveHint) for spec compliance. Mark READ_ONLY tools with readOnlyHint=true, DESTRUCTIVE tools with destructiveHint=true. This helps clients filter and prioritize safely.