Backend service for Seraph agent platform with native tools for browser automation, code execution, filesystem operations, goals management, and delegation to specialist agents
The Seraph Backend MCP Server defines 14 tools with generally adequate naming and descriptions, but significant gaps in schema completeness, parameter documentation, and error handling guidance. Most tools start with clear action verbs (browse_, delegate_, create_, update_, get_), and descriptions are present and reasonably detailed (most 100-250 chars, within baseline range of 34-392). However, parameter schemas show inconsistent typing and description coverage. Many parameters lack explicit type constraints (enums, min/max bounds), and error handling is minimal, no guidance on retryability, recovery steps, or actionable error messages. Output schemas are not documented anywhere in the source provided. The tools cover domain-appropriate operations (browser automation, code execution, file I/O, goal management, task delegation), but the execution context and guardrails remain opaque (e.g., what happens inside execute_code sandbox? What prevents path traversal in read_file/write_file?).
Apply a workspace patch operation.
Browse a webpage and extract its content. Use this tool to visit web pages, read articles, check documentation, or gather information from any URL. This uses a real browser so it works with JavaScript-rendered pages.
Manage structured browser sessions and page references. Args: action: One of providers, open, list, read, snapshot, or close. url: URL for open. session_id: Session id for read, snapshot, or close. ref: Snapshot reference for read. provider: Optional browser provider name. capture: extract, html, or screenshot.
Request a missing input from the user before continuing. Use this when a required parameter or decision is missing and guessing would be risky or would degrade the output. Ask one concise blocking question.
Create a new goal in the user's goal hierarchy. Use this when the user mentions a goal, objective, or task they want to achieve. Decompose large goals into smaller sub-goals by setting parent_id.
Multiple tools lack enum constraints for categorical parameters. 'browser_session.action', 'delegate_task.specialist', 'execute_code.language', 'create_goal.level', 'create_goal.domain', 'update_goal.status', 'get_goals.level', 'get_goals.domain', 'get_goals.status' accept free-form strings when they should enforce specific values. This invites hallucinated or invalid values from LLMs.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract specific fields if they don't know what the response structure is. This is a pervasive gap affecting all 14 tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | 1.7.0+ | v1 |
Delegate a task to a specialist agent.
Execute bounded code inside Seraph's sandbox and return the output. This is the preferred Hermes-parity runtime surface for sandboxed code execution. It keeps the current Python-only sandbox contract for v1 while leaving broader shell/process work to the dedicated runtime slice.
Fill a template file with variable substitution.
Get a summary of goal progress across all life domains.
Get the user's goals, optionally filtered.
Preview a workspace patch operation before applying it.
Read a file from the workspace.
Update a goal's status or title.
Write or create a file in the workspace.
WRITE tools (delegate_task, execute_code, write_file, apply_workspace_patch, fill_template, create_goal, update_goal) lack error handling guidance, recovery hints, and confirmation/dry-run patterns. Agents are blind to failure modes and have no guidance on retryability or user corrections.
Critical security gaps in file I/O tools (read_file, write_file, preview_workspace_patch, apply_workspace_patch, fill_template). No documentation of path traversal prevention, file size limits, or workspace boundaries. execute_code sandbox constraints (network, file system, time, memory) are completely undocumented.
Descriptions are present but often terse and lack actionable context. 'delegate_task': 'Delegate a task to a specialist agent.' does not explain which specialists exist, when to use vs other tools, or what the response contains. 'read_file' does not explain workspace location or size limits. Most descriptions are 50-100 chars when 150-250 is more useful for LLM decision-making.
No pagination or result limits documented for 'get_goals' and 'get_goal_progress'. If these return thousands of goals, context window degrades and LLM reasoning suffers. Missing: limit/offset/page params, total count in response, next_cursor guidance.
Tool dependencies and chaining not documented. After 'create_goal' returns, what field does 'update_goal' expect as parent_id? Does 'delegate_task' return a task_id that can be tracked? Does 'execute_code' return stdout/stderr separately? Undocumented chains force discovery detours.
Date handling in 'create_goal' and 'update_goal' mentions ISO format but no validation or error guidance. LLMs frequently miscalculate dates, explicit format constraints and examples of valid formats would reduce errors.