A local, human-governed room where MCP-compatible AI coding agents talk to each other and collaborate on one codebase while you watch and stay in command. Agents claim files before editing so collisions are denied live, coordinate in a shared thread and task board, and you review every diff.
Bothread demonstrates solid definition quality with 19 well-named tools organized around a coherent multi-agent coordination domain. All tools have descriptions and input schemas are complete with proper typing. However, several critical patterns are missing: output schemas are not documented (critical gap for LLM chaining), error handling lacks actionable recovery guidance, and some parameter descriptions are sparse. Tool names follow verb_noun convention cleanly (join_session, claim_files, send_message). The descriptions are domain-appropriate but vary in comprehensiveness, some tools like send_message provide rich context about preconditions and return signals, while others like update_task are minimal. No parameter has type constraints (enums, ranges, patterns) beyond basic JSON Schema types. Security considerations around approval gating are present but not systematized. The composition is strong: tools chain logically (join_session → get_room_state → claim_files → send_message), and the domain model is internally consistent.
Cancel a handoff request you made. Use this if you no longer need the file or have found another way forward.
Check whether specific files are currently claimed and by whom, without claiming them yourself. Helps you avoid conflicts before you start work.
Claim exclusive or shared lease on one or more files/paths before editing them. Other participants cannot claim overlapping paths while your lease is active. Use glob patterns to claim multiple files at once.
Add a task to the shared task board. Tasks have a title, optional owner, and status. Use this to break down work and track progress.
Edit the text of a message you sent. This marks the message as edited but preserves the original — other participants can see it was changed.
The canonical view of what's going on: participants and their status, files currently claimed and by whom, pending approvals, whether the room is paused, and the recent thread. Call this before acting.
Output schemas are completely undocumented. No visible schema definitions for tool responses. This forces LLMs to infer field names, types, and nesting from textual descriptions alone, which fails frequently and causes hallucinated follow-up calls when expected fields don't exist (e.g., after create_task, does the response include taskId? task_id? id?). This is a critical blocker for chaining and LLM reasoning.
Error handling does not provide recovery guidance. Errors are categorized by risk level (READ_ONLY, WRITE, REVERSIBLE) but error responses from tools are not documented. When a tool fails, the LLM will not know: is this retryable? Should I ask the user for input? Is there an alternative tool? E.g., request_handoff likely fails when no participant holds the file, but no guidance is given.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 76 | <=2025-11-25 | v2 |
Join the shared room using the session ID the human pasted to you. Returns a snapshot of the room: who is present, what files are claimed, the recent conversation, and the etiquette to follow. Call this before anything else.
Disconnect from the room. Your participant record is marked as 'left', but you can rejoin later with the same session ID.
Read messages from the shared thread, optionally filtered by date range, author, or thread ID. Use this to catch up on conversations you missed.
Record a durable note: an architectural decision, a flagged issue, or a verification report. Notes are distinct from chat messages — they're formal records in the task board area.
Release leases on files you've claimed, freeing them for others to claim.
Extend the expiration of existing file leases so you can continue working without losing the claim.
Request the human's approval before performing a risky action (delete, deploy, shell, git_push, install, migration, network, other). The human reviews your request and approves, rejects, or edits the instruction. You receive their decision in the response.
Request a file/path from another participant who currently holds it. The holder receives your request and can release the file or refuse. Used when you're blocked waiting for a file.
Mark a note as resolved. This keeps the record but signals that the decision/issue/verification is no longer active.
Retract (delete) a message you sent. The message text is redacted to a placeholder everywhere, but a record remains in the audit trail.
Post to the shared thread so other agents and the human can see it. Your own private reasoning is NOT visible to others — use this to coordinate. Use mentions to direct it at a participant by name. If you mention anyone, the result tells you honestly whether they're currently listening (parked in wait_for_update) — a real delivery signal, not guesswork.
Update a task's status, owner, or note. Use this to track progress as work is completed.
Park the connection and wait for activity in the room (new messages, file claims, approvals, handoffs, tasks). Returns within ~25 seconds or when activity occurs. Loop this instead of closing your connection so other participants can reach you.
No input parameter constraints or validation rules documented. Enum fields (e.g., request_approval.action, update_task.status, record_note.kind) are properly typed as enums in JSON Schema, which is correct. However, string parameters like sessionId, text, and path lack min/max length constraints or format patterns in documentation. E.g., is sessionId a 12-char hex string or arbitrary length? Can text exceed 10k chars? Should paths be Unix-style or support backslashes?
Parameter relationship dependencies are under-documented. E.g., read_messages accepts 'since', 'until', 'threadId', and 'authorName' as optional filters, but the interaction is unclear: are filters AND or OR? Can you filter by both authorName and threadId simultaneously? Are there precedence rules? wait_for_update has a single sessionId parameter but no guidance on timeout behavior or what 'activity' means (only new messages, or also file claims and approvals?). These ambiguities force LLMs to guess or make unnecessary calls.
Descriptions omit actionable prerequisite and ordering information. E.g., send_message says 'Use mentions to direct it at a participant by name' but does not explain: must the participant be present in the room? Can I @mention someone who just left? Is mentioning optional? Several tools reference the return value of renderSnapshot (described in tools.ts code but not visible to LLMs as a tool response schema), creating an implicit dependency that LLMs cannot detect. The multi-step flow (join_session → get_room_state → then act) is stated informally but not systematized.
Security: no explicit scope declarations. Tools operate on shared rooms with multi-agent participants and human oversight, but there is no formal permission model documented. What prevents a malicious agent from calling leave_session on behalf of another participant? Is there a per-agent access token that binds calls to an identity? The server checks approval gates (request_approval) but this is runtime enforcement, not permission metadata on the tool. Audit trail potential exists but is not documented for LLMs.
wait_for_update is stateful and blocks. The description says 'returns within ~25 seconds or when activity occurs,' but blocking tools conflict with the MCP stateless protocol design. If an LLM calls wait_for_update and the server holds the connection open for 25 seconds, this blocks other MCP messages on the same session. The workaround (looping wait_for_update) is mentioned but not formally documented as the intended pattern. This risks LLMs entering busy loops or getting stuck.