Build LLM applications with ease. A framework for wiring LLM components, agents, jobs, and services with support for multiple agent backends (AgentScope, Claude Code SDK) and MCP server integration.
FlowLLM provides 5 tools with explicit schemas and descriptions, but significant gaps exist in parameter documentation, error guidance, and composability. Tool naming is reasonable (verb-based) but descriptions are inconsistent in depth. The exec_command tool is particularly concerning, it accepts arbitrary shell commands and Python code with minimal constraints, creating a high-risk attack surface. Parameter descriptions exist but lack actionable guidance (constraints, format specs, valid ranges). Error handling is minimal, most tools lack recovery guidance for common failure cases. The server uses HTTP transport (fastmcp/FastAPI) which is current, but lacks several production-grade patterns like structured error responses, input validation hints, and confirmation steps for destructive operations.
Execute a command or uploaded Python code and stream output.
Terminate all currently running tasks.
Terminate a running task by id.
List tracked tasks, optionally filtered by status.
Return server health and process metadata.
exec_command accepts arbitrary shell commands and Python code with no input sanitization or prompt-injection defense. Accepts both 'command' and 'python_code' simultaneously but schema does not document mutual exclusivity or validation rules.
exec_command lacks actionable error descriptions. No recovery guidance if execution times out, fails, or produces no output. LLM cannot distinguish between 'command not found', 'syntax error', 'timeout', or 'permission denied', all return raw subprocess output.
kill_task and kill_all_tasks lack confirmation/dry-run pattern. These are irreversible operations but descriptions do not warn of side effects or suggest verification steps.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 31 | - | v1 |
exec_command timeout parameter (default 3600s) is unbounded in description. No guidance on valid range, consequences of exceeding max timeout, or what happens if a task exceeds the timeout mid-execution.
list_tasks 'tail' parameter defaults to 30 but lacks minimum/maximum documentation. No explanation of what 'output_tail' format is (list of strings? single string with newlines?).
No documented output schemas for any tool. LLMs cannot plan downstream calls or extract required fields. Example: kill_task returns {task_id, status, message} but this is not formally documented in a schema.
exec_command description (66 chars) is below the 10 - 1024 recommended range in terms of completeness. Does not explain when to use it vs raw shell commands, what happens if both command and python_code are provided, or how streaming output works.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. LLMs cannot infer which tools are safe to retry, which are read-only, or which have side effects without explicit hints.