Context Orchestrator for Documentation — autonomous multi-project documentation agent with MCP tool support for task management, ADRs, activity tracking, and agent workflows
14 tools with explicit schemas and descriptions. Most tools follow verb_noun naming (activity_list, adr_create, agent_pick). Descriptions are present but vary in clarity and completeness, some are detailed (adr_create: 'Create a new ADR. ID auto-allocated as ADR-NNN if not provided...') while others are opaque (agent_pick: 'Atomically acquire next ready task with full context...' lacks clear WHEN-to-use guidance). However, descriptions are often technical/implementation-focused rather than LLM-optimized for agent reasoning. Output schemas are not documented in the provided code, we can infer the structure from parameter names but not verify return types or field documentation. Error handling is mentioned for adr_get ('Miss returns {found: false, ...}') and agent_get ('On miss returns structured hint...') but most tools lack explicit error guidance. No documented examples, constraints, or recovery paths for most tools.
List activity events, newest first. Filters by scope_kind, scope_id, kind, actor_kind, date range with pagination.
Attach a Mermaid diagram to an ADR. Auto-appends if position omitted.
Create a new ADR. ID auto-allocated as ADR-NNN if not provided. Status can be: proposed | accepted | superseded | deprecated | rejected.
Return an ADR with its diagrams and task links. Miss returns {found: false, requested_adr_id, hint, related_tools}.
List ADRs in a project, optionally filtered by status. Returns compact rows without diagrams/links.
Mark superseded_adr_id as replaced by superseding_adr_id. Creates DAG edge AND auto-flips old ADR's status to superseded in one transaction. Idempotent.
No documented output schemas. Return types and response field structures are not declared in tool definitions. LLMs cannot plan downstream operations without knowing what fields to extract.
Descriptions lack WHEN-to-use guidance. agent_pick describes WHAT it does ('Atomically acquire next ready task...') but not WHEN an agent should call it vs alternatives, or what prerequisites exist. LLMs struggle to select between similar tools without this guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2026-07-28+ | v2 |
Re-sync an ADR body from its markdown projection. Use when the markdown is the source and the DB row fell behind. Works on ACCEPTED and terminal ADRs. Deliberately lacks status field.
Patch ADR fields. Only supplied (non-None) fields change.
Single-call L0 bootstrap for an AI agent. Returns minimum self-describing snapshot: server version, profile, role, forbidden tools, skills, task statuses, and next_action_hint. Stays under 4KB for tight context budgets.
Guarded task completion with optional proof URL. Validates acceptance & transitions task to done + releases lock. Idempotent on success.
Opt-in deep fetch when agent_pick's card didn't include something. Fetches full_doc_body, related_task, story_full, plan_export, or file_content. On miss returns structured hint with legal_what.
Atomically acquire next ready task with full context. Composes ready_for_project + lock filter + task_checkout + context_get + skill-body inlining. Idempotent on (project, agent_id). Returns task card with context, navigation, and applicable skills.
Give up without done — release lock and return task to ready. Reason is optional. Idempotent: re-running returns the already-released task.
Dispatch progress, blocker, or approval message from agent. Routes to activity log and orchestrator queue. Kind: progress | blocker | approval. Replaces old ask_human/task_block pattern.
agent_get parameter 'ref' is described as 'Reference identifier (doc_key, task_id, story_id, plan_scope, or file path), optional' but lacks clarity on when each type is required. If ref semantics depend on 'what', this dependency must be explicit in both parameter descriptions.
Error handling patterns are present (adr_get, agent_get mention 'miss' responses) but not consistent across all tools. Most tools lack guidance on retryability, user-fixable vs fatal errors, or recovery steps.
adr_supersede and agent_report descriptions do not explicitly state these are destructive/state-changing operations. LLMs need to know which calls are safe to retry and which have irreversible consequences.
agent_capabilities has an empty input schema ({}). While this is correct for a no-argument tool, the description is vague ('Single-call L0 bootstrap for an AI agent...') and does not clearly state this is a discovery/bootstrap tool to be called once at session start.