MCP server for managing and orchestrating LLM agents as task runners
TaskDriver MCP has 23 tools with generally present descriptions and input schemas, but quality is inconsistent. Most tools have basic verb_noun naming (create_project, list_tasks, get_task) following pattern conventions. However, descriptions vary significantly in quality, some are specific and actionable (e.g., get_project: 'Always call this first to obtain project instructions'), while others are generic (e.g., health_check: 'Check if the TaskDriver server is healthy and responsive'). Input schemas are present for all tools with typed parameters, but many lack depth in parameter descriptions. Output schemas are not documented in the provided source. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are visible. Error handling descriptions are absent, tools do not explain how to handle failures or what to do on conflict. Pagination support is present (limit/offset params on list tools), which is good, but total counts and next_cursor patterns are not documented. The design shows thoughtful composition (separate create_task vs create_tasks_bulk, distinct get_next_task vs peek_next_task), and naming clarity is generally strong (get_next_task explains agent reconnection in description). However, descriptions fail to guide LLMs on sequencing, for instance, get_project_stats lacks context on when to call it vs other read operations. No sensitive parameter handling is evident in the tool definitions themselves, though the server appears to support API-based auth (not shown in tool params, which is correct).
Clean up expired task leases and reassign tasks
Mark a task as completed with results and optional structured outputs
Create a new project
Create a single task from a task type
Create a new task type with template and variables
Create multiple tasks in bulk from a list
Extend the lease on a task the agent is currently working on
Output schemas are not documented in tool definitions. LLMs cannot plan downstream calls or extract required data without knowing what fields will be returned. This violates pattern:tool and pattern:response-shaper.
Generic descriptions lack LLM-optimized guidance on WHEN to use each tool. E.g., get_project_stats says 'Get project statistics including task counts by status' but does not explain whether to call it for monitoring, before creating new tasks, or after completion. health_check offers no context for its purpose. Descriptions should be 50-200 chars and include discovery intent.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Extend the lease on an active task (alternative to extend_lease with different parameter order)
Mark a task as failed with error details
Get statistics about active leases in a project
Get the next available task from the project queue. If agentName is provided and has an existing task lease, that task is resumed. Otherwise assigns a new task. Agent names are only used for reconnection after disconnects.
Get detailed project information including instructions and configuration. Always call this first to obtain project instructions and context that agents need to understand their role and objectives.
Get project statistics including task counts by status
Get detailed task information
Get detailed task type information
Check if the TaskDriver server is healthy and responsive
List agents currently working on tasks (agents with active task leases). This is for monitoring purposes - agents are ephemeral and only appear here when actively working.
List all projects with pagination
List all task types for a project
List tasks with filtering and pagination
Peek at the next available task without acquiring a lease (read-only inspection)
Update project details
Update task type configuration
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. The MCP protocol allows servers to mark tools with execution semantics, e.g., fail_task and complete_task should have destructiveHint=true (idempotent but have side effects), list_* tools should have readOnlyHint=true, and health_check should be marked idempotent. This helps clients and LLMs reason about retry safety and precedence.
Error handling guidance is missing. Tools do not document what errors they return, when failures are retryable, or what the LLM should do next. E.g., get_next_task might fail if no tasks are queued, the description should say 'Returns null or error if queue is empty' and suggest fallback actions. create_tasks_bulk with continueOnError=true should document the per-item error structure.
extend_lease and extend_task_lease appear to be redundant tools with different parameter orders. This violates pattern:tool composition, having two tools that do the same thing confuses LLMs and wastes reasoning. Consolidate into a single canonical tool or clearly document the distinction.
Pagination is present (limit/offset) on list tools, but total counts and next_cursor patterns are not documented. LLMs cannot determine if there are more results or how many total items exist. Responses should include total_count, has_next, or next_offset to enable proper iteration.
Parameter descriptions are often too brief. E.g., config param in create_project says 'Project configuration including defaultMaxRetries and defaultLeaseDurationMinutes' but does not describe the expected JSON structure, valid ranges for these fields, or required vs optional keys. Agents need more detail to construct valid payloads.
duplicateHandling parameter in create_task_type describes valid values ('allow, skip, or error') as text rather than an enum. Enums are machine-parseable and prevent hallucinated values. Define it as enum: ['allow', 'skip', 'error'].
status parameter in list_tasks describes valid values as text ('queued, running, completed, failed') rather than enum. Use enum constraint instead.