A comprehensive agentic workflow system with MCP integration for building and running automated coding workflows. Supports multiple LLM providers, browser automation, workspace management, and background task execution.
The server defines 3 tools with reasonable action-verb naming and detailed descriptions. However, there are significant gaps in schema completeness and parameter validation. All tools have descriptions exceeding the minimum threshold, and naming follows verb_noun conventions. The main weakness is that input schemas lack comprehensive parameter type definitions and constraints, only `capture_context` has detailed parameter descriptions with types; `get_activity_status` has an empty properties object; `trigger_and_auto_notify` has minimal parameter guidance. Error handling is not visible in the tool definitions, and the descriptions, while present, do not consistently guide LLMs on when to use each tool vs alternatives. Output schemas are not documented. This puts the server in the C-to-low-B range.
Capture durable user-supplied runtime business context for this workflow. Use only after the user confirms the item should be remembered across runs, and only when the Workflow Profile allows business-context accumulation. Writes to knowledgebase/context/context.md. Keep the capture concise and authoritative. Use for persistent rules, preferences, constraints, assumptions, examples, ICP filters, approval rules, brand voice, or domain context that workflow steps must respect. Do not use for one-off instructions, general chat memory, workflow-discovered facts that belong in knowledgebase/notes, or execution recipes that belong in learnings.
Return a JSON snapshot of currently running workflow executions and workflow schedules. Use this when the user asks what workflows, background runs, or cron jobs are running right now.
Run plain Python trigger code asynchronously and return immediately. When it exits, fails, or reaches its timeout, the platform resumes this same chat through [AUTO-NOTIFICATION]. Print the information the resumed agent should receive. Use for a timer or bounded polling condition. This does not create workflow scripted steps, survive a server restart, or provide an incoming webhook listener.
get_activity_status has no input parameters defined (empty properties object), yet the description implies it should accept optional filters. Schema score is 15 due to lack of any parameter definitions.
trigger_and_auto_notify description does not clarify idempotency, side effects, or retry behavior. The description mentions 'asynchronously' and 'auto-notification' but does not explain what happens if the agent retries the same trigger or what guarantees exist about execution (at-most-once vs at-least-once). Critical for an async execution tool.
No output schemas documented for any tool. LLMs cannot plan downstream tool chains or extract required fields (e.g., workflow IDs from get_activity_status, or confirmation IDs from trigger_and_auto_notify) without documented return types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 67 | 2026-07-28+ | v2 |
capture_context 'section' parameter has no constraint on allowed values. The description says 'Defaults to General when empty' but does not specify what other section names are valid, making it an unbounded free-form string that could produce invalid markdown section headers.
trigger_and_auto_notify 'timeout_seconds' parameter has no guidance on what happens if the timeout is reached. Does the async task continue running? Is it killed? Do the results still come back? The description lacks this critical dependency and failure mode information.
No error handling patterns visible. Tool descriptions do not guide LLMs on recovery steps (e.g., if capture_context fails due to invalid section, what should the LLM do? If trigger_and_auto_notify times out, should it retry?). No categorization of errors as retryable vs fatal.