Composable orchestration platform for conversational AI with voice agent, backend API, and frontend. Supports STT, LLM, and TTS integration with LiveKit for real-time communication.
Arkenos exposes 32 tools across HTTP REST endpoints and agent runtime modules. Critical issues pervade the suite: (1) Most tool descriptions are minimal action statements lacking WHEN/WHY context or return value documentation required for LLM planning. (2) Input schemas present but lack rigorous type constraints, many accept free-form 'object' or 'string' params where enums or strict formats should apply. (3) Parameters largely undescribed, e.g., POST /api/agents/{agent_id}/containers/build accepts no documented input beyond agent_id, leaving LLMs guessing about required config. (4) Output schemas entirely absent from visible source, no documentation of what create_agent, build_container, or synthesize_speech actually return, forcing LLMs to reason blindly. (5) Error handling non-existent, no guidance on retryability, recovery paths, or actionable error messages. (7) Security tool (PUT /api/settings) allows raw API key/secret injection without encryption, audit logging, or permission gating. Definition quality is below community median.
Tools (13)
download_agent_fileswriteauthsource verified
Download all files for the given agent from MinIO into /app/workspace/.
fetch_agent_configread onlyauthsource verified
Fetch agent configuration from the backend API.
fetch_and_inject_keysread onlyauthsource verified
Fetch API keys from backend (async). Safe to call inside the entrypoint.
Fetch API keys from backend (synchronous). Only safe at module level / startup.
install_extra_requirementswritesource verified
If the workspace contains a requirements.txt, pip-install it.
load_custom_agentwritesource verified
Import workspace/agent.py and return the AgentServer instance. Looks for a module-level `server` variable first (preferred pattern). Falls back to calling `create_agent()` if `server` is not found.
Output schemas entirely absent from source code. No visible documentation of return types for any tool (create_agent returns?, build_container returns?, synthesize_speech returns?). Violates pattern:tool and mxe:response-field-naming. LLMs must reason blindly about available fields.
POST /api/agentsGET /api/agentsGET /api/agents/{agent_id}POST /api/sessionsPOST /api/callsPOST /api/agents/{agent_id}/containers/buildresemble_synthesizelog_stt_usagelog_llm_usagelog_tts_usage
Tool descriptions are minimal action statements (10 - 50 chars) lacking WHEN to use, prerequisites, or return value documentation. Examples: 'Create a new agent' (19 chars), 'Update an existing agent' (24 chars), 'Stop a running agent container' (30 chars). No WHEN/WHY rationale for LLM planning.
POST /api/agentsPUT /api/agents/{agent_id}POST /api/sessionsPOST /api/telephony/initiatePOST /api/callsPOST /api/agents/{agent_id}/files
Recommendations
Document output schemas for all 32 tools. For each endpoint, explicitly specify: response status codes, returned object structure, field types, and which fields are safe to chain to downstream tools. Example: 'create_agent returns {agent_id, name, description, agent_mode, config, created_at, status}. Use agent_id for all subsequent agent-specific operations.'
Expand tool descriptions from 1 - 3 lines to 50 - 200 chars (per A+ baseline of 194 chars avg). Include: WHAT does this do? WHEN should the LLM call it (vs. similar tools)? WHAT are the prerequisites or dependencies? WHAT does it return and how should it be used? Example for POST /api/agents: 'Create a new voice agent. Call this first to set up an agent with STT, LLM, and TTS configuration. Returns agent_id for use in create_session, build_container, and other agent-specific operations. agent_mode controls whether the agent runs in sandbox (STANDARD) or custom code (CUSTOM).'
Add type constraints and enums to all parameters. Replace 'agent_mode': {'type': 'string', 'description': 'STANDARD or CUSTOM'} with explicit enum: {'type': 'string', 'enum': ['STANDARD', 'CUSTOM'], 'description': 'STANDARD: sandbox mode, CUSTOM: custom Python code'}. For 'config', define the nested schema or provide a link to schema docs.
Add descriptive parameter documentation for all inputs. Every parameter needs: type, constraints (min/max length, regex, enum), and LLM-actionable guidance. Example: 'limit (integer, 1 - 100): Maximum results per page. Defaults to 20. Larger values risk context window exhaustion; use pagination for large result sets.' Replace generic 'object' with concrete nested schemas or point to OpenAPI/JSON Schema reference.
Score history
Overall score trend
↓ 33 points across a rubric change (v1 → v2)
12/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
12
<=2025-11-25
v2
2026-03-09
F
45
-
v1
log_stt_usagewriteauthsource verified
Fire-and-forget: log STT usage. Quantity = minutes of audio.
log_tts_usagewriteauthsource verified
Fire-and-forget: log TTS usage. Quantity = character count.
lookup_agent_by_phoneread onlyauthsource verified
Look up an agent by its assigned Twilio phone number.
Try to resolve agent ID via SIP phone number lookup, with retries.
setup_python_pathwritesource verified
Add workspace and lib dirs to sys.path AND PYTHONPATH. PYTHONPATH must be set so that child processes spawned by the LiveKit agents forkserver can re-import the custom agent module by name.
Free-form 'object' type parameters without enumeration or nested schema. 'config' (POST /api/agents), 'usage_data' (POST /api/usage/events), and unnamed payload in container build all accept arbitrary JSON. Violates pattern:constrained-input. LLMs cannot validate inputs and will hallucinate fields.
POST /api/agentsPUT /api/agents/{agent_id}POST /api/usage/eventsPOST /api/agents/{agent_id}/containers/build
No error handling guidance. Tools lack actionable error messages, recovery steps, or retryability hints. Violates pattern:recovery-guide. Example: 'delete agent' has no docs on what happens if agent is in-use, whether deletion is idempotent, or what error codes may occur.
PUT /api/settings accepts raw API keys, secrets, and auth tokens as parameters without encryption, masking, or server-side injection. Violates pattern:secret-injection. Secrets will appear in agent call logs, request traces, and potentially user-facing responses.
Six agent runtime tools (resemble_synthesize, log_*_usage) appear as module functions without explicit MCP tool registration visible in pyproject.toml or agent/__init__.py. Schema and description are inferred from docstrings only.
No permission gates on destructive operations. DELETE /api/agents/{agent_id}, DELETE /api/agents/{agent_id}/files/{file_id}, and POST /api/agents/{agent_id}/containers/build have no visible permission checks or confirmation workflow. Violates pattern:permission-gate and pattern:confirmation-request.
No pagination parameters on list endpoints. GET /api/agents, GET /api/sessions, GET /api/usage/events, and GET /api/agents/{agent_id}/files lack limit, offset, page_size, or cursor parameters. Violates pattern:paginated-result. Returning unbounded lists risks context window exhaustion.
GET /api/agentsGET /api/sessionsGET /api/usage/eventsGET /api/agents/{agent_id}/filesGET /api/agents/{agent_id}/coding-agent/conversations
Date/time parameters lack format specification. 'start_date' and 'end_date' (GET /api/usage/events) are strings with no ISO 8601 notation, Unix timestamp guidance, or example. LLMs frequently miscalculate date formatting, producing invalid inputs.
GET /api/usage/events
Implement error handling guidance. For each tool, document: (1) Possible error codes (400, 404, 409, 500, timeout), (2) Whether the error is retryable, (3) What the LLM should do next. Example for DELETE /api/agents: '404 Not Found (non-retryable): agent does not exist; call GET /api/agents to list available agents. 409 Conflict (retryable): agent has active sessions; wait 5 sec and retry, or call GET /api/sessions to check active sessions. 500 Internal Server Error (retryable): transient error; retry up to 3 times with exponential backoff.'
Move API keys, secrets, and credentials out of tool parameters. Implement server-side secret injection via environment variables or a secure vault. Secrets in parameters are logged, traced, and risk exposure. Tool definition for PUT /api/settings should NOT accept raw livekit_api_key; instead, accept a 'secret_reference' string and resolve it server-side.
Add permission gates and audit logging to destructive operations (DELETE, PUT settings). Before DELETE /api/agents/{agent_id}, verify the calling user owns the agent and log the deletion attempt with user ID, timestamp, and agent_id. Return 403 Forbidden if the user lacks permission.
Support human-readable identifiers alongside UUIDs. For agent_id, accept both UUID and agent_name. Implement a resolution function server-side. Example: POST /api/sessions accepts {'agent_id_or_name': 'my-voice-agent'} and resolves to the UUID. This matches mxe:chat-data-model, users say 'use my sales agent', not 'use 550e8400-e29b-41d4-a716-446655440000'.
Add pagination parameters (limit, offset, cursor, page_size) to all list endpoints. GET /api/agents should return {'agents': [...], 'total': N, 'page': 1, 'limit': 20, 'next_cursor': '...'} to support large result sets. Document the default and max limits (e.g., default=20, max=100).
Standardize date/time formats to ISO 8601 (e.g., 2024-01-15T10:30:00Z) and document in parameter descriptions. Replace ambiguous 'start_date' with 'start_date (ISO 8601 string, e.g., 2024-01-15T00:00:00Z): earliest event to include. If omitted, defaults to 30 days ago.'
Verify and explicitly register agent runtime tools (resemble_synthesize, log_*_usage) in MCP tool manifests. If these are Python module functions, convert them to explicit tool definitions with JSON Schema inputs, typed outputs, and server-side registration. Currently they appear inferred, risking ambiguity.
Add optional confirmation/dry-run support for irreversible operations. POST /api/agents/{agent_id}/containers/build could accept an optional 'dry_run': true parameter. Returns a preview of what would be built (Dockerfile analysis, resource usage) without committing resources. Reduces accidental destructive calls.
Batch operations for tools called in loops. Instead of N separate POST /api/usage/events calls, support batch_log_usage accepting an array of usage events. Reduces token overhead and latency for agents logging multiple STT/LLM/TTS events per session.