A task management and project planning system with MCP server support for task creation, planning, documentation, and activity tracking
The Backlog MCP server presents a moderate implementation with reasonable tool coverage (16 tools) but significant gaps in parameter documentation, schema completeness, and error handling. Tool naming follows verb-noun conventions well (project_list, task_create, task_update, etc.), but descriptions are inconsistent, some tools lack operational context. Parameter schemas are partially visible in the source (e.g., task_create shows explicit JSON schema with types and descriptions), but many parameters lack detailed constraints (e.g., 'type' accepts 'task|bug|issue|...' but no enum is enforced in schema). Output schemas are not documented, callers don't know what fields to expect. Error handling is basic (generic -32603 errors) with no recovery guidance. The server handles 16 distinct operations across projects, tasks, plans, comments, memories, and documents, but lacks LLM-friendly prompting for parameter selection.
Add a comment to a task
Add a document to a project
List documents for a project
Get details of a specific document
Update an existing document
Add a memory/note to a project
List memories for a project
Add a plan/outline to a task
Output schemas are completely undocumented. Tools return JSON via contentText() wrapper (e.g., task_create returns toJSON(t), project_list returns toJSON(projects)), but callers cannot know what fields 'task' or 'project' objects contain. LLMs cannot plan downstream tool calls without knowing return structure.
Error responses are generic. All tool errors return -32603 'Internal server error' with only err.Error() message, no recovery guidance, no classification (retryable vs user-fixable vs fatal), no suggestion of what to try next. Example from source: 's.send(message{JSONRPC: "2.0", ID: msg.ID, Error: &rpcError{Code: -32603, Message: err.Error()}})'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
Get version history of a plan
Update an existing plan
List all projects
Create a new task
List tasks with optional filtering
Move a task to a different status
Get details of a specific task
Update an existing task
Parameter descriptions lack constraint details. 'type' parameter accepts 'task|bug|issue|improvement|feature|vulnerability|chore|spike|bucket-list' but no enum constraint is visible in schema. 'priority' accepts '1-5 (1=highest)' but no min/max is enforced. 'status' accepts 'todo|doing|done' but no enum schema exists. LLMs cannot reliably pick valid values.
No idempotency hints or confirmation mechanisms for destructive operations. task_update, task_move, comment_add, plan_update, doc_update all modify or create state with no dry-run option, no idempotent hints in tool annotations, no confirmation step. Agents may accidentally duplicate entries or overwrite data.
Tool descriptions are generic and lack operational context. Example: 'task_move' = 'Move a task to a different status' (44 chars, below rubric baseline of 50-200). No explanation of when to use it vs task_update with status parameter, no prerequisites, no side effects documented. Agents may not understand when to invoke.
task_list returns a flat list with limit=50 hardcoded. No pagination parameters (page, offset, cursor), no total count in response documented, no next_cursor. Large backlogs (>50 tasks) will be truncated with no way to fetch the rest.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) in tool definitions. The source shows tools() function must build tool list, but no hint metadata visible. LLMs cannot infer which tools are safe to call speculatively.