An autonomous agent that can execute tools, manage skills, interact with MCP servers, and handle multi-channel communication (Slack, Telegram, WhatsApp)
Auton exposes 9 tools with varying quality. All tools have descriptions and declared input schemas visible in src/agent.js. However, schemas lack proper JSON Schema formalization (no explicit type declarations in structured format), descriptions are often generic and miss context about when to use each tool, and output schemas are completely undocumented. The server combines multiple responsibilities in single tools (e.g., 'skill' handles list/activate/read/install), lacks error handling guidance, and provides no recovery paths for failures. Security is a critical concern: bash tool is unrestricted and allows arbitrary command execution; file operations (read, write, append) have minimal path validation; memory/goal/schedule/config tools lack permission gates. Naming is sometimes ambiguous (e.g., 'skill' action parameter with enums vs 'action' in goal/schedule/memory, inconsistent patterns). Parameter descriptions exist but often lack format specifications, ranges, or validation guidance. No tool documents its output schema or provides pagination/limit guidance. This is a STDIO-only server capped at 50 maximum protocol readiness.
Append content to a file. Creates file if it doesn't exist.
Execute shell commands with optional timeout and environment variables. Returns stdout, stderr, and exit code.
Get or update agent configuration. Supports reading current config, updating LLM settings, agent parameters, or channel credentials.
Manage agent goals: list current goals, add a new goal, or mark a goal as complete.
Store or retrieve agent memory for a user. Supports set (store memory), get (retrieve memory), and clear (delete memory) operations.
Read file contents from the filesystem. Supports text and JSON files.
Manage scheduled tasks: list scheduled tasks, add a new task with cron schedule, or remove a scheduled task.
bash tool permits unrestricted shell command execution with no sandboxing, input validation, or permission checks. Any LLM call to bash can execute arbitrary commands, rm -rf /, curl exfiltrate_secrets, etc. Critical security vulnerability.
File operations (read, write, append) lack thorough path traversal validation. 'read' accepts arbitrary paths; only write/append call resolve(). An LLM could read /etc/passwd or ~/.ssh/id_rsa without restriction.
No output schemas documented for any tool. LLMs cannot infer what fields to expect (e.g., does 'skill list' return {name, description, path} or {name, desc, dir}?). Blocks downstream tool chaining and forces LLMs to guess.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Manage agent skills: list available skills, activate a specific skill to get full instructions and resources, read skill files, or install new skills from git/local sources
Write content to a file. Creates file if it doesn't exist, overwrites if it does.
Input schemas lack proper JSON Schema formalization. Parameters are described in text but not declared with 'type', 'required', 'enum', 'minLength', 'maxLength', etc. Cannot be validated by JSON Schema validators or parsed by strict clients.
bash tool timeout defaults to 30000ms but is configurable, acceptable. However, description does not mention timeout behavior, failure modes, or that this tool can hang indefinitely if timeout is set to 0 or very high. Incomplete error guidance.
Memory, goal, schedule, config tools lack permission checks. Any LLM can set/clear memory for any user ID, modify agent goals, schedule tasks, or change LLM configuration (including API_KEY). No user/agent identity validation.
No audit logging in tool implementations. No trace of who called what tool, when, with what parameters, or what the result was. Compliance and incident response will fail.
skill tool combines four responsibilities (list, activate, read, install) into one tool. Ambiguous 'action' enum requires LLM to reason about which action to take, and different actions have incompatible required parameters. Should split into list_skills, activate_skill, read_skill_file, install_skill.
Descriptions are generic and lack context about when to use each tool. E.g., bash: 'Execute shell commands...' does not explain risks, when to prefer it over other tools, or common failure modes. No pattern matching against typical agent workflows.
No error handling in tool implementations. Functions return plain strings like 'ERROR: Path escapes skill directory' with no structured error codes, recovery hints, or retry guidance. LLMs cannot parse or reason about failures.
bash tool description mentions 'timeout in milliseconds (default 30000)' but does not state timeout behavior or what happens on timeout (kills process? returns partial output? exception?). Incomplete parameter documentation.
config tool permits setting 'llm.apiKey' via LLM-accessible parameter. Credentials should NEVER be exposed as tool parameters, they should be injected server-side via environment or vault. This invites secrets into logs and traces.
File operations lack idempotency guarantees. write/append are not idempotent (repeated calls append/overwrite data). No confirmation or dry-run pattern for destructive operations.
No pagination or limit enforcement. memory.get, goal.list, schedule.list, bash output could return massive results (bash stdout unbounded, memory/goal/schedule state unbounded). No cap stated; LLM context could be exhausted.