MCP server providing cognitive enhancement tools for LLM agents: structured reasoning, mental models, and literate programming notebooks
Four tools with highly specialized, experimental domains (literate programming, notebook execution, peer gradation, JavaScript-based execution). Tool definitions are present with descriptions and input schemas, but schemas are inconsistently detailed, parameter descriptions vary in clarity, and output schema documentation is sparse or absent. The server is designed for cognitive enhancement workflows rather than conventional CRUD operations, which affects composability and discoverability. Most tools accept code/configuration strings as input, making validation and error guidance difficult. Schema quality ranges from reasonable (thoughtbox_search) to minimal (thoughtbox_notebook operation enum lacks per-operation parameter guidance).
Run JavaScript using the `tb` SDK to chain Thoughtbox operations in a single call. **One state-mutating operation per call.** Submit only one `tb.thought()`, `tb.ulysses()`, `tb.theseus()`, hub-mutating call (`tb.hub.register()`, `tb.hub.createWorkspace()`, `tb.hub.createProblem()`, `tb.hub.mergeProposal()`, etc.), claims-mutating call (`tb.claims.assert()`, `tb.claims.invalidate()`, `tb.claims.supersede()`, etc.), or merge-mutating call (`tb.merge.request()`) per `thoughtbox_execute` invocation. Each response contains guidance (patterns, session state, protocol state) that should inform your next operation. Batching multiple state-mutating calls bypasses this feedback loop and produces lower-quality reasoning. Read-only operations (`tb.session.*`, `tb.knowledge.*`, `tb.observability()`, `tb.branch.*`, `tb.hub.whoami()`, `tb.hub.listWorkspaces()`, `tb.hub.readChannel()`, `tb.claims.query()`, `tb.claims.affected()`, `tb.merge.status()`, `tb.merge.list()`, `tb.merge.claimDiff()`, etc.) and session variables (`tb.vars.*` — store intermediate values across execute calls within this MCP session) may be freely chained. Example: ```js async () => { const sessions = await tb.session.list(); await tb.thought({ thought: "Analyzing prior sessions", thoughtType: "reasoning", nextThoughtNeeded: true, }); return sessions; } ``` Use `console.log()` for debugging — output captured in response logs. All tb methods return their result directly (already parsed from the tool response).
Notebook toolhost for literate programming with JavaScript/TypeScript. Create, manage, and execute interactive notebooks with markdown documentation and executable code cells.
Brokered MCP peer notebook surface. Seeds text artifacts, invokes peers on the development-only local-process runtime (the builtin claim-extractor plus graduated notebook peers such as contradiction-scan, whose own code cells execute from the graduation snapshot), manages the draft-to-active manifest lifecycle, graduates notebooks into draft peer manifests, and reads invocations, traces, and artifacts.
Missing output schemas and response documentation. Tools return results without documented structure, LLM must infer field names and types from examples or context.
Operation-enum parameters lack descriptions and per-operation parameter guidance. 'operation' enum in thoughtbox_notebook (20 values) and thoughtbox_peer_notebook (10 values) have no enum descriptions, forcing LLM to decode which operation matches user intent.
Sparse parameter descriptions across all tools. Parameters like 'until' (thoughtbox_notebook), 'args' (thoughtbox_peer_notebook), 'manifest' lack descriptions or type guidance. LLM cannot determine valid values or format.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 13 | - | v1 |
Discover Thoughtbox operations, prompts, and resources by writing JavaScript that queries the catalog. The `catalog` object is available in scope: interface SearchCatalog { publicTools: Array<{ name: string; description: string; operations?: string[] }>; operations: Record<string, Record<string, { title: string; description: string; category: string; inputSchema?: object; }>>; prompts: Array<{ name: string; description: string; args: string[] }>; resources: Array<{ name: string; uri: string; description: string; mimeType: string }>; resourceTemplates: Array<{ name: string; uriTemplate: string; description: string; mimeType: string }>; } Modules in catalog.operations: session, thought, knowledge, notebook, theseus, ulysses, observability, branch, hub, claims, runbook, merge, vars Public MCP tools in catalog.publicTools: thoughtbox_search, thoughtbox_execute, thoughtbox_peer_notebook Examples: - List all modules: `async () => Object.keys(catalog.operations)` - List public tools: `async () => catalog.publicTools` - Find session operations: `async () => catalog.operations.session` - Search by keyword: `async () => { const q = "entity"; return Object.entries(catalog.operations).flatMap(([mod, ops]) => Object.entries(ops).filter(([_, op]) => op.description.toLowerCase().includes(q)).map(([name, op]) => ({ module: mod, name, title: op.title }))) }` - Find prompts: `async () => catalog.prompts.filter(p => p.name.includes('interleaved'))` - List resources: `async () => catalog.resources.map(r => ({ name: r.name, uri: r.uri }))`
No error recovery guidance. Tools do not document error conditions, retry guidance, or actionable error messages. LLM cannot self-correct on failures.
Code/configuration parameters lack validation constraints. 'code' parameters in thoughtbox_execute and thoughtbox_search accept arbitrary JavaScript, no pattern, length limit, or syntax guidance for LLM.
Inconsistent naming clarity. 'thoughtbox_notebook' and 'thoughtbox_peer_notebook' are domain-specific; 'thoughtbox_execute' is generic. No clear naming pattern to disambiguate when to use each.
No idempotency guidance. Tools like thoughtbox_notebook (with state-mutating operations) do not declare idempotency, LLM cannot safely retry on transient failures.
Missing permission/scope documentation. Tools accept opaque IDs (notebookId, peerId, invocationId) with no indication of access control, ownership validation, or required permissions.