A personal AI assistant that remembers, learns from experience, and never makes the same mistake twice. Includes tool management, MCP server integration, and multi-channel messaging support.
This server has severe quality gaps across all dimensions. While 23 tools are defined, most lack adequate parameter descriptions, none have documented output schemas, and descriptions are minimal (10-50 chars) rather than LLM-optimized (50-200 chars baseline). Tools like 'spawn', 'agent_browser', 'coding_agent' have dangerously vague interfaces with single generic parameters ('label', 'action', 'agent') that do not clarify constraints, enums, or expected formats. Critical operational tools ('exec', 'forget') lack confirmation patterns and error recovery guidance. The server exposes high-risk destructive operations (shell execution, memory deletion) without permission gates or audit instrumentation. Parameter schemas are minimal, most lack type constraints, ranges, or format specifications. No pagination, chaining-IDs, or downstream reference patterns are evident in the source. Tool composition is poor: 'session_transcript' vs 'session_status' create naming ambiguity; 'check_tasks' and 'check_tasks_json' duplicate functionality. Error handling is absent from visible tool definitions.
Control a browser agent for automated browsing
Cancel a running task
Check the status of running tasks
Check task status and return as JSON
Clear the current task plan
Invoke a coding agent for code generation or analysis
Get detailed information about a coding agent
Create a task plan
CRITICAL: Destructive tools ('exec', 'forget') lack confirmation patterns and error recovery guidance. Shell command execution and memory deletion are irreversible; no dry-run, confirmation step, or permission gates are visible.
CRITICAL: Tool parameters use generic free-form string inputs ('action', 'label', 'agent') without enum constraints, ranges, or format specifications. LLMs cannot infer valid values and will hallucinate inputs.
HIGH: No output schemas are documented. Tool descriptions state what is returned (e.g. 'Get recent sessions') but not the structure (fields, types, counts). Agents cannot plan downstream calls without knowing response fields.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Schedule or manage recurring tasks
Execute a shell command
Remove information from memory
Store information in memory
Send content to a specific session
Get the default session
Lookup a specific session
Get recent sessions
Resolve a session reference
Get the status of a session
Get the transcript of a session
Spawn a new task or agent with context
Update stored memory information
Update a step in a task plan
Fetch content from a web URL
HIGH: Naming ambiguity and redundancy. 'session_transcript' vs 'session_status' are unclear in purpose; 'check_tasks' and 'check_tasks_json' duplicate functionality (only output format differs, not behavior). LLMs conflate similar names and waste reasoning on choosing between them.
HIGH: Parameter descriptions are missing or trivial (under 20 chars). E.g., 'action' in 'cron' and 'agent_browser' is described only as 'Cron action to perform' or 'Browser action to perform', no examples, enums, or valid values listed.
MEDIUM: Tool descriptions are under 50 chars (baseline 194 chars for A+ tools). Descriptions like 'Spawn a new task or agent with context' and 'Get recent sessions' lack context for when to call them, prerequisites, or what 'context' entails.
MEDIUM: No pagination, rate-limiting, or result-capping guidance. Tools like 'session_recent' and 'check_tasks' return potentially large lists; no mention of max results, offset/limit params, or token budget implications.
MEDIUM: No error handling guidance visible. Tools lack actionable error messages, recovery hints, or error categorization (retryable vs user-fixable vs fatal). E.g., 'session not found' offers no remediation path.
MEDIUM: No permission gates or audit instrumentation. Destructive tools ('exec', 'forget', 'spawn') do not declare required scopes (e.g. 'write:memory', 'exec:shell') and no evidence of access control checks or logging.