A shared record for AI agents: decisions, completed work, and commitments, with sources. Git-native engineering memory and a self-hosted state server.
Hunch exposes 28 tools with significant definition gaps. Core tools (hunch_task, hunch_context, hunch_why, hunch_query, hunch_decision, hunch_correction, hunch_finding, hunch_bug) have descriptions but lack visible input schemas in the source code provided. Parameter descriptions exist but are minimal (10-50 chars). The nuryel_* state management tools and constitution_* tools have even sparser documentation. No tool annotations (readOnlyHint/destructiveHint) are visible. Error handling guidance is absent, tools do not explain recovery paths or categorize failures. Output schemas are not documented. Many tools accept task_id and cwd parameters but lack constraints or format guidance. The server appears to be a sophisticated engineering memory system, but the MCP interface lacks the rigor expected for production agent integration.
Record a bug: a defect found in the code, with root cause and severity.
Retrieve the lineage of a bug: when it was introduced, how it evolved, and related bugs.
Check which constraints (invariants) are in scope for a file or symbol and would block an edit.
List behavior candidates for Constitution G2: behavioral patterns ready for evaluation.
Materialize behavior for Constitution G2: convert behavioral patterns into policies.
Materialize behavior policy for Constitution G2: convert a specific behavioral pattern into a policy.
Input schemas not visible in source code. Tools declare parameters (task_id, target, query, etc.) but no explicit JSON Schema definitions are shown. Cannot verify type constraints, required fields, or format specifications.
Descriptions are sparse and lack LLM-optimized guidance. Most descriptions are 25-50 characters, below the 50-200 char baseline for A+ tools. Descriptions do not explain WHEN to use each tool vs. similar ones (e.g., hunch_why vs. hunch_context distinction is unclear).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 45 | 2025-06-18+ | v2 |
Replay behavior for Constitution G2: re-evaluate historical behavioral patterns.
List candidate policies for Constitution G2: policies ready for evaluation.
Run an operational drill for Constitution G2: test policy evaluation without committing.
Check readiness for Constitution G2 evaluation: verify that the repository meets prerequisites.
Retrieve the shadow queue for Constitution G2: pending policy evaluations.
Check readiness for Constitution G3 evaluation: verify that the repository meets prerequisites.
Retrieve engineering memory context for a file or symbol: invariants that MUST NOT break (with severity and rationale), near-invariants reached through the dependency blast radius, the decisions that shaped the code (including rejected alternatives), bug history with root causes, and direct dependents.
Record a correction: a constraint that MUST NOT be broken, with severity and rationale. Corrections block edits that would violate them.
Record a decision: a choice made, its rationale, and rejected alternatives. Decisions shape code and guide future edits.
Record a finding: an observation about the codebase, architecture, or engineering practices.
Get the direct dependents of a file or symbol: what would be affected by changes.
Full-text search over Hunch memory: decisions, bugs, constraints, findings, and other records.
Retrieve a task's contribution report: verified checks, conformance results, and delivered records.
Start or finish a task's contribution report. Start once per user task, unless the host's prompt hook already opened the task and printed its verify command — then reuse that task_id and do not start. Pass the task_id to hunch_context. Finish before your final response and include the returned concise contribution card, without asking the user; skip finish only when the prompt hook's own instruction said this host closes the task and the task used no Hunch (no hunch_* call on this task_id, no verified check, no hook context you acted on, nothing to claim), and a task you started with this tool must always be finished. Applications are explicitly agent-reported and must name an exact delivered occurrence and record hash. Completion never implies successful verification. Not for storing decisions or claiming tests passed; use the CLI task verify wrapper for observed command results.
Why a file or symbol is the way it is — decisions, invariants, bug history, and blast radius from the repo's Hunch engineering memory.
List the capabilities of the state partition: what state kinds are available and their schemas.
Capture state changes: record a state mutation with proof and delivery envelope.
Capture multiple state changes in a batch: record multiple state mutations atomically.
Read state from a partition: retrieve current state for a given facet and key.
List state records: retrieve metadata about recorded state changes.
Subscribe to state changes: receive notifications when state is updated.
Write state to a partition: update or create state for a given facet and key.
No output schemas documented. Tools return data but LLMs cannot plan downstream calls without knowing response structure. E.g., hunch_context returns 'invariants' and 'decisions' but field types and nesting are not specified.
No error handling guidance. Tools do not explain what errors are retryable, user-fixable, or fatal. No recovery paths provided. E.g., if hunch_context fails to find a target, the LLM has no guidance on next steps.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). LLMs cannot distinguish safe read-only tools from destructive ones. hunch_task with action='finish' and hunch_decision (WRITE) should be marked destructive; hunch_context (READ_ONLY) should be marked safe.