AI organism evolution and parallel task execution - Speed optimized tool-enabled agents. Bug colonies with specialized roles for parallel task execution via Ollama and MCP tools.
Agent Farm has 14 tools with basic schemas and descriptions, but falls significantly short of production quality. All tools have input schemas with proper JSON Schema types and descriptions, which is good. However, descriptions are uniformly brief and lack the LLM-optimization guidance required for reliable tool selection. Most descriptions are 30-50 characters, below the 194-char baseline for production tools. Parameter descriptions are sparse and lack constraints, ranges, examples of valid values, or guidance on when to use each tool vs. alternatives. No output schemas are documented. Error handling is minimal, tools return error dicts but provide no recovery guidance. No tool has destructive/readOnly/idempotent hints. The tools themselves are competently implemented (path safety, command validation, timeouts) but the MCP interface is bare-bones. This is typical of a functional but unpolished STDIO server.
analyze_code(code) - Basic code analysis
check_service(name) - Check if a systemd service is running
disk_usage(path) - Get disk usage for a path
exec_cmd(cmd) - Execute a shell command (with safety checks)
file_exists(path) - Check if file exists
http_get(url) - HTTP GET request
http_post(url, body) - HTTP POST request
All tool descriptions are 30-50 characters, well below the 194-char production baseline. Descriptions lack WHEN to use, WHAT it returns, or how to distinguish from similar tools. E.g., 'http_get(url) - HTTP GET request' gives the LLM almost no selection context.
No output schemas documented for any tool. LLMs cannot infer result structure; they must guess what fields to expect (e.g., does http_get return 'body' or 'content' or 'text'?). This forces exploratory calls and wastes context.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 33 | - | v1 |
kmkb_ask(question) - Ask KMKB a question
kmkb_search(query) - Search KMKB knowledge base
list_dir(path) - List files in a directory
process_list(sort_by) - List top processes
read_file(path) - Read contents of a file
system_status() - Get basic system status
write_file(path, content) - Write content to a file. IMPORTANT: Quote content with spaces/commas.
exec_cmd accepts arbitrary shell input with only blocklist validation (no allowlist). LLMs are prone to prompt injection and can bypass blocklists. Parameter description does not mention timeout (60s), output truncation (10KB), or safety checks, the agent has no way to know this tool has hard limits.
process_list 'sort_by' parameter accepts only 'cpu' or 'memory' but has no enum constraint in schema. Description says 'Sort by "cpu" or "memory"' but LLM will not necessarily honor this; a formal enum is needed.
Error responses are minimal. E.g., exec_cmd returns {'error': 'Blocked dangerous command pattern'}, no recovery guidance. Should say: 'Command contains blocked pattern (e.g., rm -rf). Consider using a safer alternative or breaking the task into steps.'
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). LLM cannot tell which tools are safe to retry, which are idempotent, or which cause side effects. This is critical for exec_cmd (destructive) vs read_file (safe).
write_file description warns: 'IMPORTANT: Quote content with spaces/commas.' This is a red flag, the parameter parsing should handle this, not require LLM compliance. The fact that it needs this warning suggests the tool interface is fragile.
Role-based tool access (ROLE_TOOLS dict) is implemented in Python but NOT exposed via MCP. The agent/client has no way to know which tools it can access without attempting calls. Permission gates should be reflected in tool availability or error messages.