Persistent memory and background intelligence for Claude Code — captures facts, verifies staleness, KingBee daemon
keephive implements 8 memory-management tools with clear, verb-based names and generally descriptive docstrings. Most tools have well-defined input schemas with parameter types and descriptions. However, there are significant gaps: (1) output schemas are not formally documented for any tool, the code does not show return type annotations or structured response definitions; (2) several tools lack error handling guidance and recovery instructions; (3) some descriptions are generic and do not clearly explain WHEN to use each tool vs. alternatives; (4) the tool set shows composition issues, hive_remember and hive_recall could benefit from richer parameter constraints (e.g., enums for category prefixes). The STDIO transport is a hard blocker for protocol readiness (capped at 50), but the tool definitions themselves are above-average for a community server.
Quality audit: three perspectives (Vault, Cleaner, Strategist), LLM synthesis, and score. Uses 3 parallel LLM calls for perspectives + 1 synthesis call. Falls back to metrics-only when HIVE_SKIP_LLM is set.
List knowledge guides, or view one by name (prefix matching).
View daily log. Defaults to today. Pass YYYY-MM-DD, 'yesterday', or N (days ago). Entries from PreCompact may include [project:name] tags for cross-project attribution.
Search all memory tiers: working memory, knowledge, daily logs (30d), archive.
Save an insight to today's daily log. Prefix with: FACT, DECISION, CORRECTION, TODO, or INSIGHT.
Status overview: facts, stale warnings, TODOs, recent entries.
Output schemas are not documented for any tool. Tools return responses (status, audit results, log entries, TODOs, etc.), but there is no formal documentation of the response structure, fields, types, or constraints. LLMs cannot predict what data is available and must infer it from trial-and-error.
Several tools accept free-form string inputs (pattern in hive_todo_done, query in hive_recall, day in hive_log) without formally constraining the format. This invites hallucinated invalid inputs and requires error recovery. hive_todo_done should constrain 'pattern' (regex? substring? prefix?). hive_log should accept enum values for common cases (today, yesterday, N days) rather than relying on narrative parsing.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
List all open TODOs with ages.
Mark the first open TODO matching pattern as done.
Error handling and recovery guidance are absent. Tools like hive_audit (parallel LLM calls), hive_todo_done (pattern matching), and hive_recall (search) can fail or return empty results, but the code does not show how these are communicated to the LLM or what next step is suggested. Example: 'No TODOs match pattern "release". Available TODOs: [list]' would enable self-correction; a bare empty list does not.
hive_remember accepts prefixes (FACT, DECISION, CORRECTION, TODO, INSIGHT) as narrative text in the description but does not enforce them as an enum in the schema. This allows LLMs to pass arbitrary prefixes like 'NOTE' or 'RANDOM', which are silently ignored or cause parsing errors. The input schema should include an enum constraint or a separate 'category' parameter with enum values.
Tool descriptions lack clarity on WHEN to use each tool vs. similar alternatives. hive_recall, hive_knowledge, and hive_log all retrieve memory, but descriptions do not explain the distinction, does recall search across tiers while log shows a single date? Is knowledge a separate index? LLMs waste reasoning cycles deciding which to call.
hive_audit documents that it uses 3 parallel LLM calls + 1 synthesis call, and mentions HIVE_SKIP_LLM fallback. However, there is no documented timeout, retry behavior, or failure mode. If an LLM call times out or fails, does the tool fail entirely or return partial results? What should the agent do in that case?