The server provides 10 tools with mixed quality. Naming is mostly clear and verb-focused (bash, install_with_apt, text_editor, remote_bash, remote_text_editor, remote_download, use_glob, use_grep, llm_proxy_local, llm_proxy_remote). All tools have descriptions of adequate length (40-400+ chars). However, several tools lack proper output schema documentation, the server code shows tool registration but does not expose return type schemas. Parameters have type definitions and descriptions for most tools, though some are generic. Error handling is implicit rather than explicit in the visible code. The text_editor and remote_text_editor tools have particularly rich parameter sets (10+ params with full descriptions), which is above the baseline average of 4 params. The bash and remote_bash tools document risks (IRREVERSIBLE, WRITE) but lack explicit recovery guidance in descriptions. Use_glob and use_grep have well-structured descriptions with usage patterns. The llm_proxy tools accept empty input, which is valid but unusual and could benefit from optional parameter documentation about configuration. Overall, the server demonstrates competent naming and parameter documentation but falls short on output schema clarity and structured error recovery patterns.
Run commands in a bash shell. When invoking this tool, the contents of the "command" parameter does NOT need to be XML-escaped. You don't have access to the internet via this tool. You do have access to a mirror of common linux and python packages via apt and pip. State is persistent across command calls and discussions with the user. To inspect a particular line range of a file, e.g. lines 10-25, try 'sed -n 10,25p /path/to/the/file'. Please avoid commands that may produce a very large amount of output.
Install one or more apt packages inside the Docker container as the root user.
Start a lightweight OpenRouter proxy inside the same Docker container used by the "bash" tool. Returns a JSON string with url and scope. The url is reachable from inside the "bash" container as localhost. It runs a standard OpenAI compatible server, so you can use it with any OpenAI compatible client. You can use all models available on OpenRouter, for instance: - openai/gpt-5-mini - google/gemini-2.5-pro - anthropic/claude-sonnet-4
Start a lightweight OpenRouter proxy on the remote GPU machine. Returns a JSON string with url and scope. The url is reachable from inside the remote machine as localhost. It runs a standard OpenAI compatible server, so you can use it with any OpenAI compatible client. You can use all models available on OpenRouter, for instance: - openai/gpt-5-mini - google/gemini-2.5-pro - anthropic/claude-sonnet-4
Output schemas not documented in source code. No explicit return type definitions visible for any tool. LLMs cannot infer what fields to expect from responses, forcing them to guess at downstream tool parameter requirements.
Error handling lacks recovery guidance. The bash tool risks timeout and command failures but does not specify what error messages it returns or how the LLM should recover. Same for install_with_apt, remote_bash, remote_text_editor (marked WRITE/IRREVERSIBLE).
llm_proxy_local and llm_proxy_remote accept empty input schema ({}). Descriptions claim they 'start' a proxy and return JSON with 'url' and 'scope', but no input parameters are documented. If configuration is required (e.g., model selection, timeout), it should be exposed as optional parameters with descriptions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | C | 64 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 11 | - | v1 |
Run commands via SSH on a remote GPU instance.
Download files from a remote GPU instance via rsync.
Custom editing tool for viewing, creating and editing files on a remote GPU instance.
Custom editing tool for viewing, creating and editing files. State is persistent across command calls and discussions with the user. If `path` is a file and `show_lines` is True, `view` displays the result of applying `cat -n`. If `path` is a file and `show_lines` is False, `view` displays the result of applying `cat`. If `path` is a directory, `view` lists non-hidden files and directories up to 2 levels deep. If a `command` generates a long output, it will be truncated and marked with `<response clipped>`. The `undo_edit` command will revert the last edit made to the file at `path`. Always write arguments with keys, do not rely on positions. Be careful with escaping strings and line breaks in Python code.
- Fast file pattern matching tool that works with any codebase size - Supports glob patterns like "**/*.js" or "src/**/*.ts" - Returns matching file paths sorted by modification time - Use this tool when you need to find files by name patterns
A powerful search tool built on ripgrep Usage: - ALWAYS use Grep for search tasks. NEVER invoke `grep` or `rg` as a Bash command. - The Grep tool has been optimized for correct permissions and access. - Supports full regex syntax (e.g., "log.*Error") - Filter files with glob parameter (e.g., "*.js", "**/*.tsx") or type parameter (e.g., "js", "py", "rust") - Output modes: "content" shows matching lines, "files_with_matches" shows only file paths (default), "count" shows match counts - Pattern syntax: Uses ripgrep (not grep) - literal braces need escaping (use `interface\{\}` to find `interface{}` in Go code) - Multiline matching: By default patterns match within single lines only. For cross-line patterns like `struct \{[\s\S]*?field`, use `multiline: true`
remote_bash and remote_download descriptions lack context on prerequisites. Users/LLMs do not know if a remote instance must be provisioned first or what 'remote GPU instance' refers to. Descriptions should mention EXISTING_INSTANCE_ID setting or link to setup docs.
Idempotence not documented. text_editor and remote_text_editor support str_replace, insert, write, undo_edit commands. Multiple invocations of the same edit command could fail or corrupt state if 'old_str' doesn't match after the first successful edit. Descriptions should clarify idempotence assumptions.
Text editor undo_edit command lacks clear scope. Can the LLM undo across multiple tool invocations? Is undo state per-file or global? Description does not clarify, risking silent failures when LLM assumes undo can roll back to an earlier state.