AgentOS agent runtime: ingestion, agentic RAG and Deep Agents execution with MCP server integration
AgentOS presents a well-structured set of 5 tools with clear naming conventions and comprehensive descriptions. All tools follow verb_noun naming patterns (list_tools, run_mission, resume_mission, cancel_mission, ingest_document), which is excellent for LLM parsing. Descriptions are detailed and explain WHAT the tool does, WHEN to use it, and important operational context (e.g., background task execution, idempotency semantics). However, there are significant gaps in parameter validation, output schema documentation, and error handling guidance. Input schemas are present and properly typed, but parameter descriptions lack detail about formats, ranges, and constraints. No output schemas are documented in the source code. Error handling is minimal, no guidance on recovery paths or error categorization.
Cancel a running mission. Cancelling a mission that is not running is not an error: the control plane is the authority on status, and this is idempotent.
Parse, chunk, embed and index one uploaded document version. Ingestion runs as a background task for the same reason a mission does: a large scanned PDF can take minutes through OCR, and the control plane must not hold a connection open for it. Progress is written to the document version row, which is what the knowledge base UI reads.
Report the tools and specialists registered for a workspace. The control plane renders this directly, so the product's tool inventory is what the runtime would actually give an agent — including MCP tools discovered from that workspace's configured servers.
Resume a mission with an approval decision. Handles approval, rejection, or editing of mission arguments.
Execute a mission. Missions run as background tasks rather than inside the request. A mission takes minutes; holding an HTTP connection open for it would tie the control plane's liveness to the agent's, and any proxy in between would time it out. The caller gets an immediate acknowledgement and follows progress through the event stream.
No output schemas documented. Tools define inputs correctly, but return types are not specified in any source code visible. LLMs cannot infer what fields to expect, breaking tool-chaining and forcing exploratory calls.
Parameter descriptions lack format specifications and constraints. 'workspace_id' and 'mission_id' are typed as strings but no description explains UUID format, length, or validation rules. LLMs may pass invalid identifiers.
No error handling guidance provided. Tools are marked as WRITE/READ_ONLY but no descriptions explain what errors can occur, whether they are retryable, or what recovery steps the LLM should take.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 64 | 2026-07-28+ | v2 |
resume_mission requires 'edited_args' conditionally (only if decision='edit'), but the schema does not mark it as conditionally required. No parameter description explains the dependency between 'decision' and 'edited_args'.
No pagination parameters documented for list_tools. If the tools list grows large, response could exceed context limits. No total_count, next_cursor, or limit/offset parameters are visible.
Parameter 'file_path' in ingest_document lacks description explaining path format, object storage location, or what happens if the file does not exist.