Multi-agent shared brain across Claude Code/Desktop, Codex, Antigravity CLI, Copilot, and VS Code. Cross-session memory, self-improving skill loops, inter-agent signaling — one local MCP server.
Thread-keeper defines 11 tools with explicit schemas and descriptions. Most tools have clear names (evolve_apply, db_compact) and detailed descriptions (100-300 chars). However, several critical gaps reduce the score: (1) Parameter descriptions are sparse or missing context. For example, evolve_apply_curator_report's 'report_path' lacks guidance on what happens when empty. (2) Output schemas are not documented, tools return text/structured content but the response structure is not formally specified. (3) Error handling is minimal, no recovery guidance or actionable error messages visible. (4) Tool annotations (readOnly, destructive, idempotent) are declared in the FEATURES list but not visible in the actual tool definitions in the source code provided. (5) Some tools combine multiple concerns (e.g., evolve_apply_conflicted_pr both resolves conflicts AND lands the PR), violating single-responsibility. Baseline: 194 chars avg description (these are 150-250), 72 chars avg param description (these are 30-80). 100% of A+ tools have param descriptions; this server has ~60%.
Shrink the DB file: VACUUM + mandatory dialog_fts rebuild. Run in a quiet window — VACUUM needs an exclusive lock and copies the whole file (minutes on a multi-GB DB); concurrent FTS searches during the vacuum→rebuild gap may map to wrong rows until the rebuild commits. Fails soft (with a retry hint) when the DB is busy.
Remove base-table embedding BLOBs already represented in sqlite-vec. The operation is coverage-gated: rows without a confirmed vec0 mirror keep their BLOB fallback. Defaults to a report-only dry run. Run ``db_compact`` afterwards to return the newly freed pages to the filesystem.
Implement a PROMOTED + not-yet-applied format-evolution suggestion. Spawns an `evolve_applier` child that: edits render_brief() in threadkeeper/brief.py to make the change; adds/extends a GOLDEN test asserting the new behavior appears AND the existing brief still renders; runs the FULL suite (`.venv/bin/python -m pytest -q`) until green; then opens a PULL REQUEST on a feature branch via `gh` — it NEVER pushes or commits to main (a human reviews + merges).
Repair an already-open applier PR that currently has merge conflicts. With `pr_number=0`, picks the oldest open same-repo applier PR (`roadmap/…` or `evolve/…` head branch) whose GitHub merge state is conflicted. With a number, validates that specific PR is open, applier-owned, and conflicted. The child resolves conflicts, runs the suite, and pushes the SAME PR branch; it then lands that PR into main via GitHub's protected merge flow and does not open a new PR or mark a roadmap issue applied.
Output schemas not documented. Tools return structured content but response field structure is not formally specified in JSON Schema. LLMs cannot plan downstream calls without knowing what fields to expect.
Parameter descriptions lack actionable detail. 'report_path' in evolve_apply_curator_report says 'Path to a specific Curator report, or empty string to pick the latest' but does not explain the selection logic or error cases. Descriptions should state WHAT, WHEN, and WHAT HAPPENS.
Tool annotations (readOnly, destructive, idempotent) declared in FEATURES but not visible in source code tool definitions. Cannot verify they are actually registered with the MCP server.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 63 | 2025-06-18+ | v2 |
Apply a Curator advisory report using the existing evolve_applier role. With no `report_path`, picks the latest complete `REPORT-*.md` in `THREADKEEPER_CURATOR_REPORTS_DIR` that has not already been marked applied. Single-flight: refuses while any evolve_applier child is in flight. The child may patch/delete memory through curated MCP tools, but does not edit code, use git, or open a PR.
Implement one open GitHub issue through the evolve_applier role. With `issue_number=0`, picks the next open issue: `roadmap`-labeled issues first, then FIFO by issue number. The child implements exactly one issue, runs the suite, opens a PR with `Closes #N`, then calls evolve_mark_roadmap_issue_applied(issue_number, pr_url).
Show evolve-applier config + curator/evolve queues + running applier + the last 5 apply/recovery passes.
Mark a format-evolution suggestion as APPLIED — called by the evolve_applier child after it has opened the PR. Sets applied=1 (so the suggestion drops out of the brief / evolve_review) and records the PR url. `pr_url` is REQUIRED and must be non-empty: this is the PR gate — never mark a suggestion applied without a real pull request. A human still reviews + merges the PR.
Mark a Curator report as processed by the evolve_applier child. The report must live under `THREADKEEPER_CURATOR_REPORTS_DIR`, match `REPORT-*.md`, contain `CURATOR_PASS_COMPLETE`, and still match the parent-verified content hash. The hash and applied event prevent a swapped report or replay from being accepted.
Mark a roadmap issue as handed off — called by evolve_applier only after it has opened a real pull request for that issue.
Free the default managed checkout's `.venv` after explicit confirmation. Pass `confirm=True` to delete only that auto-managed virtualenv. Its clone remains intact and the next Evolve pass rebuilds the environment. This refuses `THREADKEEPER_EVOLVE_REPO_ROOT` and auto-clone-disabled setups, so it never prunes an operator-selected checkout.
No error handling guidance. Tools like evolve_apply_conflicted_pr and db_compact mention failure modes ('Fails soft when DB is busy') but do not provide recovery steps or actionable error messages for LLMs.
Tool composition: evolve_apply_conflicted_pr combines conflict resolution, test execution, and PR landing, three distinct concerns. Should split into resolve_pr_conflicts + land_pr to enable independent composition.