MCP server for running and managing long-running tasks on remote GPU/compute nodes via SSH
Strong tool definitions with comprehensive descriptions and well-structured schemas. All 11 tools have explicit names, descriptions, and input schemas with proper type definitions and parameter annotations. Tool naming follows verb_noun convention (remote_*). Descriptions are detailed and action-oriented (avg 150-250 chars), explaining WHAT each tool does, WHEN to use it, and dependency hints for agent teams. Parameters consistently include type constraints and descriptions. Output schemas are documented in prose. Error handling is present with recovery guidance. Main gaps: (1) STDIO transport caps at 50 regardless, (2) no explicit per-item error reporting for batch operations, (3) missing optional confirmation pattern for destructive operations (remote_stop_task uses force flag but no dry-run), (4) output field names not always matched to input expectations across tool chains.
Live-check ALL background tasks in one call. Returns status + new output lines (diff-based) for each task. Batches SSH calls per node for efficiency. Instant, non-blocking — ideal for periodic monitoring and agent teams. Pass node name to filter, or omit for all nodes.
Pull files or directories from a remote node back to your local machine using rclone. Syncs from node:remote_path to local_dir, respecting exclude patterns. Useful for retrieving logs, checkpoints, and results. Requires rclone installed locally. Returns summary of files transferred and timing.
Wait for new output from a background task (long-polling). Blocks up to 'timeout' seconds. Returns ONLY lines produced since the last call (incremental/diff-based — saves tokens). Returns immediately when the task finishes. Call this in a loop to continuously monitor a running task. Use timeout=0 for instant non-blocking check — returns new lines immediately without waiting. This is the recommended mode for agent teams where blocking is unacceptable.
List configured remote GPU nodes and check which ones are reachable via SSH. Returns name, host, user, GPU description, and connectivity status for each node.
List all background tasks started in this session with their last known status. Instant response — does not contact remote nodes. Use remote_task_status on a specific task for a live check.
STDIO-only transport: server not remotely accessible, cannot be used by hosted MCP clients. Hardens protocol readiness ceiling at 50.
remote_stop_task (IRREVERSIBLE operation) lacks dry-run or confirmation pattern. Tool uses a force flag but does not support pre-execution confirmation step before sending SIGTERM/SIGKILL.
Output schemas are documented in prose descriptions but not formally exposed via JSON Schema. Downstream tools expecting specific field names (e.g., task_id from remote_start_task) are not validated against a machine-readable contract. LLMs infer structure from natural language descriptions only.
Batch operation error handling not explicit. remote_check_tasks processes multiple tasks but does not specify per-task success/failure reporting. If one task fetch fails, unclear whether partial results are returned or entire operation fails.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 73 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Run a short command on a remote node via SSH and return its output. Output is auto-truncated to max_lines. If the command takes too long, use remote_start_task instead to run it in the background.
Launch a long-running command on a remote node in the background (via nohup). Returns immediately with a task_id. The process survives SSH disconnection. Use remote_follow_task to get incremental output, or remote_task_status for a one-off snapshot. In agent teams, return the task_id to the supervisor and use remote_follow_task with timeout=0 for non-blocking status checks.
Stop a running background task on a remote node. Sends SIGTERM by default (graceful shutdown), or SIGKILL if force=True. Returns success status and message. Task status transitions to 'failed' with the signal as exit code.
View or update .ssh-tasks.yaml in a project directory. Pass project_dir to read/validate config. Pass updates (remote_dir, exclude) to modify. Returns the current config after any updates. Helps manage rclone sync excludes and remote directory path. Changes persist in .ssh-tasks.yaml.
Push a local project directory to a remote node using rclone. Syncs from project_dir to node:remote_dir, respecting .ssh-tasks.yaml excludes. Requires rclone to be installed on your machine. Batches large syncs. Returns summary of files transferred and timing.
Get a one-off snapshot of a background task: status (running/completed/failed), elapsed time, exit code, and the last N lines of output. Use this for a quick check. For continuous monitoring, use remote_follow_task instead.
remote_sync_config parameter 'project_dir' defaults to current working directory but server-side CWD context not documented. Risk that agent calls across multiple sessions with inconsistent working directory assumptions.
Tool descriptions include helpful guidance ('Use timeout=0 for instant non-blocking check') but this is implementation detail guidance, not parameter constraint documentation. No formal enum or pattern constraints; LLMs must parse text to infer valid ranges.