Persistent task queues and goal decomposition for cross-session AGI autonomy. Provides persistent task queues that survive sessions, AI-powered goal decomposition, task scheduling and monitoring, progress tracking and recovery, and Relay Race Protocol for 48+ agent pipelines (God Agent integration).
The server provides 17 tools with basic schemas and descriptions, but suffers from significant gaps in definition quality. Most tool descriptions are present but brief (10-50 chars), parameter descriptions lack detail and validation constraints, and output schemas are entirely undocumented. The codebase shows sqlite3 database operations and task/goal management, but the MCP tool registration is minimal. Parameter types are declared in JSON Schema format, but validation rules, enums, and format constraints are absent. Error handling is basic with no recovery guidance. Naming is generally clear (verb_noun pattern), but several tools lack actionable descriptions that would guide LLM selection and usage.
Create a new goal with name, description, and optional metadata.
Create a new relay race pipeline.
Create a new task associated with a goal.
Decompose a goal into subtasks using AI analysis.
Execute a complete relay race pipeline with sequential agent execution.
Get circuit breaker state for an agent.
Get Ember's quality feedback and conscience checking results.
Get goal by ID.
Output schemas completely undocumented. No tool describes what fields are returned or what structure the LLM should expect. This forces LLMs to guess what data they receive and breaks tool chaining when downstream tools need IDs from upstream responses.
Parameter descriptions are generic and lack validation constraints. E.g., 'status' parameter in update_goal_status has description 'New status value' but no enum of allowed values (active, completed, etc.). 'priority' in create_task lacks min/max bounds (1-10? 1-100?). This invites LLM-generated invalid values.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Get next task from queue (highest priority, no unmet dependencies).
Get a relay race pipeline by ID.
Get task by ID.
List all goals, optionally filtered by status.
List tasks with optional filters by goal, status, and limit.
Record a failure for circuit breaker tracking.
Record a successful execution for circuit breaker recovery.
Update goal status.
Update task status and optionally store result or error.
Tool descriptions are too brief (10-40 chars for most tools). Rubric baseline is 194 chars; these provide minimal context for LLM selection. E.g., 'Get goal by ID' does not explain when to use it vs list_goals, what context it provides, or how its output feeds downstream tools.
No error handling or recovery guidance documented. Tools have no descriptions of what errors are possible, whether they are retryable, or what the LLM should do if a call fails (e.g., if decompose_goal fails due to AI unavailability, should it retry? ask user?). This violates the recovery-guide pattern.
Parameter descriptions lack type hints and format guidance. E.g., 'dependencies' in create_task says 'Optional list of task IDs' but no guidance on whether empty list is allowed, max length, or whether circular dependencies are rejected. Format constraints must be explicit.
No pagination guidance in list tools. list_goals and list_tasks accept optional 'limit' but no 'offset' or 'cursor'. If there are 1000 goals, how does the agent fetch the next batch? No guidance on default limit or max results.
Tool composition assumes IDs are always available. E.g., update_task_status requires task_id but does not document whether LLMs should use get_next_task to fetch it or call list_tasks first. No guidance on expected ID passing between tools.