Professional MCP supervisor server with clean architecture for job and task management with LLM-powered intelligence
SupervisorMCP provides 8 well-named tools with mostly complete input schemas and descriptions. Naming follows verb_noun convention consistently (start_job, update_task, complete_task, report_problem, get_all_jobs, get_job_tasks, prune_job, get_all_problems). All tools have descriptions exceeding 20 characters, and all parameters have type definitions and descriptions. However, there are notable gaps: (1) output schemas are NOT documented, callers must infer structure from code inspection; (2) no tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk categories; (3) parameter enums are declared in descriptions rather than as formal JSON Schema enums (priority, status); (4) error handling lacks recovery guidance, functions return generic {'error': '...'} without suggesting next steps; (5) no pagination support in list operations (get_all_jobs, get_all_problems) despite potentially large result sets; (6) destructive tool (prune_job) lacks confirmation/dry-run pattern. Per-tool scores average to 62; this reflects solid naming and basic schema coverage but missing output documentation, structured error responses, and advanced patterns.
Mark a task as completed and get next task recommendations.
Get comprehensive list of all jobs with their current status.
Get all stored problems and their solutions.
Get detailed task information for a specific job.
Delete a job and all its associated tasks.
Report a problem and receive intelligent troubleshooting advice.
Start a new job with intelligent task breakdown.
Output schemas are not documented. Code returns dictionaries with unspecified fields (e.g., start_job returns {'job_id', 'title', 'description', 'tasks', ...} inferred from supervisor_service). LLMs cannot plan downstream tool calls or validate field expectations without explicit response schemas. This is a critical gap affecting all 8 tools.
No tool annotations despite clear risk metadata. Tools define risk categories (WRITE, DESTRUCTIVE, READ_ONLY) in comments but do not register readOnlyHint, destructiveHint, or idempotentHint in MCP schema. This prevents clients from applying appropriate safety gates (e.g., confirmation before destructive ops).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Update task progress with intelligent feedback.
Parameter enums declared in descriptions, not as JSON Schema. 'priority' accepts 'low, medium, high, critical' and 'status' accepts 'pending, in_progress, completed, failed', these should be formal JSON Schema enum constraints, not prose. LLMs cannot reliably parse prose enums and may hallucinate invalid values.
Error handling lacks recovery guidance. Functions return {'error': 'Failed to retrieve jobs: ...'} with no suggestion of next steps. Example: get_job_tasks returns {'error': 'Job not found'} but does not suggest calling get_all_jobs() to discover valid job IDs. This forces agents to guess recovery strategies.
No pagination support in list operations. get_all_jobs and get_all_problems return all results without limit or offset parameters. If 1000+ jobs exist, the response will bloat context, waste tokens, and degrade LLM reasoning. Should include limit, offset/cursor, and total_count.
Destructive tool prune_job lacks confirmation or dry-run pattern. No explicit error handling, idempotence guarantee, or confirmation step before deletion. Agents could accidentally delete critical jobs. Should implement confirmation_request pattern or at least return detailed warning.
get_all_jobs and get_job_tasks return verbose task summaries without limiting result size. Example: get_job_tasks iterates all tasks without returning a tasks_summary count or paginated subset. Large jobs with hundreds of tasks will produce bloated responses.