A Telegram bot powered by Rust that integrates with MCP servers, provides tool execution, memory management, skill/agent loading, scheduling, and command execution within a sandboxed environment.
RustFox provides 26 tools with explicit schema definitions and descriptions. Most tools have clear verb-noun naming and non-empty descriptions. However, multiple tools lack parameter descriptions, some schemas are incompletely specified, and error handling guidance is minimal. The server demonstrates moderate definition quality typical of community-built MCP servers, with room for improvement in parameter documentation and output schema formalization. Average tool description length ~120 chars (within baseline p10-p90 range), but parameter documentation is sparse across many tools.
Tools (26)
cancel_scheduled_taskwritesource verified75/100
Cancel an active scheduled task by its ID.
execute_commandirreversiblesource verified67/100
Execute a shell command within the sandbox directory.
execute_command tool lacks security documentation and error handling guidance. No description of command injection risks, sandbox restrictions, or what errors the tool returns. LLMs will not know when/why to use it safely.
self_upgrade tool is marked IRREVERSIBLE but lacks confirmation step or dry-run mode. Description does not explain failure modes (e.g., rollback behavior, network errors). Pattern:confirmation-request not followed.
Multiple tools (reload_skills, reload_agents, list_scheduled_tasks, plan_view) have descriptions under 70 characters and lack guidance on when/why to call them. Descriptions like 'Reload all skills from the skills directory into memory' do not explain what state change occurs or what the agent should expect.
Document output schemas for all 26 tools. For each tool, specify the JSON structure returned, including field names, types, and whether arrays are paginated. Use the pattern:response-shaper pattern.
Expand descriptions for reload_skills, reload_agents, list_scheduled_tasks, and plan_view to 100+ characters. Explain what state change occurs, when to call, and what structure is returned. Reference pattern:tool-description.
Add error recovery guidance to execute_command: 'Returns the command exit code and stderr. If exit code != 0, the command failed, consider retrying with a simpler command or adjusting parameters.' Include sandbox restrictions in the description.
Implement a dry-run mode for self_upgrade. Add a 'confirm' parameter (boolean, default false). When false, return a summary of what would be upgraded without making changes. When true, perform the upgrade. This follows pattern:confirmation-request.
For schedule_task, document the return value: 'Returns {task_id, next_run_time, status}. Store task_id for use with cancel_scheduled_task and rerun_scheduled_task.' Enables tool chaining.
Add parameter descriptions and constraints to execute_command: clarify the 'command' param format (shell syntax, no pipes unless needed), execution timeout, and sandbox root directory. Include examples of valid commands (e.g., 'cargo build', 'ls -la').
For plan_create and plan_update, add error handling: 'If plan already exists, returns error, call plan_view() first to check status. If step_id is out of range, returns 400 error listing valid indices.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 19 points across a rubric change (v1 → v2)
60/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
C
60
<=2025-11-25
v2
2026-03-09
F
41
-
v1
plan_updatewritesource verified80/100
Update a step's status in the active plan. Call before starting a step (in_progress) and after finishing (done or failed).
plan_viewread onlysource verified72/100
View the current plan as a checklist. Call at the end of execution to review progress before synthesising the final answer.
read_agent_fileread onlysource verified75/100
Read a file from an agent directory.
read_fileread onlysource verified80/100
Read the contents of a file within the sandbox directory
read_skill_fileread onlysource verified75/100
Read a file from a skill directory.
read_soul_fileread onlysource verified77/100
Read the full contents of a soul file (SOUL.md, AGENTS.md, or USER.md) from the home directory.
recallread onlysource verified75/100
Retrieve a specific piece of remembered knowledge.
reload_agentswritesource verified68/100
Reload all agents from the agents directory into memory.
reload_skillswritesource verified68/100
Reload all skills from the skills directory into memory.
rememberwritesource verified77/100
Store a piece of knowledge for long-term memory. Use this to remember user preferences, facts, or anything useful.
rerun_scheduled_taskwritesource verified73/100
Execute a scheduled task immediately.
schedule_taskwritesource verified80/100
Schedule a task to run at a future time.
search_memoryread onlysource verified78/100
Search through past conversations and knowledge using hybrid vector + full-text search. Finds semantically similar content even with different wording.
self_upgradeirreversiblesource verified68/100
Upgrade the bot to the latest version.
send_filewritesource verified75/100
Send a file from the sandbox to the current chat. The file must already exist in the sandbox.
try_new_techwritesource verified75/100
Run a sandboxed experiment with a new technology or approach.
write_agent_filewritesource verified75/100
Write a file into an agent directory.
write_filewritesource verified80/100
Write content to a file within the sandbox directory. Creates parent directories if needed.
write_skill_filewritesource verified77/100
Write a file into a skill directory under the configured skills folder.
No output schemas documented for ANY tool. The rubric requires documenting return types and structured output. LLMs cannot plan downstream tool calls without knowing what fields are returned.
plan_create and plan_update tools create/modify state but offer no error recovery guidance. What happens if a plan name conflicts? What if step_id is out of range? Descriptions do not guide LLM recovery.
try_new_tech tool description lacks specification of supported languages beyond the enum (rust, javascript). No error handling guidance if compilation fails or runtime error occurs. No explanation of sandbox constraints or result format.
search_memory and schedule_task tools mention defaults (limit=5, trigger_type enum) but do not explain output structure, pagination, or what fields are returned. LLMs cannot chain calls without knowing response schema.
patch_skill, write_skill_file, write_agent_file lack guidance on how to handle overwrite conflicts. Do these merge, append, or replace? What happens if the skill/agent does not exist? No error classification provided.
patch_skillwrite_skill_filewrite_agent_file
Update try_new_tech description to include output format: 'Returns {success: bool, output: string, execution_time_ms: number}. If compilation fails, returns success=false with error details in output.' Add note on supported language versions (Rust edition, Node version).
For read_soul_file, add guidance on when to call: 'Call this at the start to understand agent personality and constraints from SOUL.md, or to read integration details from AGENTS.md. Returns raw markdown content.' Helps with tool selection.
Add idempotency notes to write_file, write_skill_file, write_agent_file: 'Idempotent, writing the same content twice produces the same result with no duplicate side effects.' This helps LLMs safely retry on network errors.
For path parameters (read_file, write_file, list_files, send_file, etc.), add format constraints: 'Path must be relative to sandbox or absolute within sandbox (no '..' traversal). Returns error if path escapes sandbox boundary.' Prevents path injection.
Document limit caps for list_files and list_scheduled_tasks: 'Returns at most 100 items per call. If more exist, include a next_cursor or offset in the response for pagination.' Prevents context window exhaustion.
Add parameter type clarification: for 'notes' parameter in plan_update, specify max length (e.g., '0-500 characters') to prevent LLMs from passing multi-paragraph responses.
For schedule_task, clarify trigger_value format: 'ISO 8601 for one_shot (e.g., 2025-02-15T10:30:00Z) or standard 6-field cron for recurring (e.g., "0 9 * * MON").' Include examples.
Add a 'pattern-chaining' note to tools that depend on others: e.g., 'schedule_task likely used after plan_create to schedule task execution based on plan.' This guides agent reasoning.