Full SSH terminal for Linux servers with AI/user control and dual-stream visibility
Mixed quality across 13 tools. Strengths: most tools have descriptions and explicit schemas; several tools follow verb_noun naming conventions (execute_command, list_batch_scripts, delete_batch_script). Weaknesses: descriptions vary significantly in depth and LLM-optimization; parameter documentation is inconsistent; error handling guidance is minimal; output schemas are undocumented; parameter constraints (enums, ranges) are underspecified in many tools; execute_script_content description is verbose (2000+ chars, dilutes signal). Tool composition shows some brittleness, multiple tools for similar execution patterns (execute_command, execute_script_content, execute_script_content_by_id) risk LLM confusion about which to select.
Helper tool to create a batch (shell) script from command list. Useful for AI to quickly build properly formatted scripts. Example: commands = [ {"description": "Network interfaces", "command": "ip link show"}, {"description": "Routing table", "command": "ip route show"}, {"description": "DNS config", "command": "cat /etc/resolv.conf"} ]
Send Ctrl+C to a running command. Use when user wants to stop a long-running command.
Check status of a long-running command. OUTPUT_MODE: Same options as execute_command - "auto": Smart decision based on output size - "full": Get complete output (for completed commands) - "preview": Peek at first/last lines - "summary": Just metadata (polling frequently) - "minimal": Status only - "raw": Complete unfiltered output (no truncation, no filtering) Use "summary" or "minimal" when polling frequently to save tokens. Use "full" when command completes and you need results.
Delete a batch script from database (requires confirmation). First call without confirm shows script details and warning. Second call with confirm=true actually deletes. This is a hard delete - execution history is preserved but script content is lost.
Execute command on remote Linux machine with smart completion detection. BEHAVIOR: - Waits for command completion (detects prompt return) OR timeout - Returns smart-formatted output based on output_mode - Full output stored in buffer - Optionally tracks in conversation for rollback support TIMEOUT: - Default: 10 seconds (sufficient for most commands) - Override for long operations: timeout=300 (5 min), timeout=1800 (30 min) - Maximum: 3600 seconds (1 hour) OUTPUT_MODE OPTIONS: - "auto" (default): Smart output based on command type and size * < 100 lines: returns full output * >= 100 lines: returns preview only * Installation commands with errors: returns error contexts * Installation commands without errors: returns last 10 lines - "full": Always return complete output - "preview": First 10 + last 10 lines only - "summary": Metadata only (line count, error flag) - "minimal": Status + buffer_info only - "raw": Complete unfiltered output (no truncation, no filtering) CONVERSATION TRACKING: - conversation_id (optional): Associate command with conversation for tracking - If provided: command saved with conversation for rollback support - If omitted: Behavior depends on user's conversation mode choice: * "in-conversation" mode: conversation_id auto-injected * "no-conversation" mode: command saved standalone * No mode set: Commands run standalone (default behavior) ⚠️ CONVERSATION MODE WORKFLOW: - User's mode choice persists for ALL commands on current server - Mode is set when: select_server (user chooses), start_conversation, or user explicitly sets "no-conversation" - Mode is cleared when: switching servers, new Claude dialog - Claude should NEVER ask before each command - the mode handles it automatically RETURN VALUES: - status="completed": Command finished - status="cancelled": User interrupted with Ctrl+C - status="timeout_still_running": Exceeded timeout, still executing - status="backgrounded": Command backgrounded with &
execute_script_content description is excessively verbose (2000+ chars). Dilutes signal by including implementation details (LOG_FILE_LOCATION, OUTPUT_MODE_GUIDANCE, Aspen tool references) that should be in docs, not the LLM-facing description. LLM cannot reason effectively with such dense text; prevents quick tool selection.
execute_command description is 1500+ chars and conflates implementation details with actual behavior. Includes markdown formatting (TIMEOUT, OUTPUT_MODE OPTIONS, CONVERSATION TRACKING) and warnings that belong in separate documentation, not the tool description visible to LLM during selection.
No output schemas documented for any tool. LLMs cannot know what fields to expect in responses, forcing them to guess about data structure and plan downstream tool calls blindly. Critical for composition.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 66 | <=2025-11-25 | v2 |
Execute multi-command batch script on remote Linux server. OUTPUT_MODE_GUIDANCE: Use output_mode='full' for diagnostic commands with expected concise output. Full output returns directly in the response for immediate analysis. LOG_FILE_LOCATION: Script output is automatically saved to the local user's home directory: - %USERPROFILE%\mcp_batch_logs\batch_output_[timestamp].log Use Aspen tools (apsen-tool_v2:read_file) with project root ~\mcp_batch_logs to access saved log files for post-execution analysis—for example, extracting specific error context, parsing structured data, debugging by reading lines around errors, or processing log entries for analysis. Do not use bash/Linux tools to access local log files; use Aspen tools for local file system access. Workflow: 1. Pre-authenticate sudo if script contains sudo commands 2. Upload script to remote /tmp directory 3. Set executable permissions 4. Execute script with output logging (AI blocks here) 5. Download log file to local machine 6. Parse output and return structured results User sees live progress in terminal. AI is blocked until completion. Parsing happens AFTER execution completes. OUTPUT_MODE: - "summary" (default): Steps, errors, execution time + preview (first/last 10 lines) Token efficient. Log file saved locally for later analysis if needed. - "full": Includes complete output in response (for diagnostics where AI needs to analyze output) Uses more tokens but gives AI all data in one round trip. Example script format: #!/bin/bash echo "=== [STEP 1/3] Check interfaces ===" ip link show echo "[STEP_1_COMPLETE]" echo "=== [STEP 2/3] Check routing ===" ip route show echo "[STEP_2_COMPLETE]" echo "=== [STEP 3/3] Check DNS ===" cat /etc/resolv.conf echo "[STEP_3_COMPLETE]" echo "[ALL_DIAGNOSTICS_COMPLETE]"
Execute a saved batch script by ID. Loads script from database and executes it on the remote server. Increments usage counter and tracks execution in batch_executions table.
Get batch script details and content by ID. Returns complete script information including source code. Use this to view a script before executing or editing it.
Get full unfiltered output of a command. WARNING: Uses more tokens than filtered output.
List batch scripts saved in database. Browse saved scripts with filtering and sorting options. Use this to find scripts to reuse or manage.
List command execution history from database with flexible filters. Query persistent command history across all servers and sessions. Useful for: - Reviewing what commands were executed on a server - Finding commands from specific dates or time periods - Debugging by examining command history - Auditing command execution - Analyzing errors across multiple sessions Returns database records with integer IDs, command text, execution status, timestamps, error information, and conversation context when available. By default, shows commands from the currently connected server.
List tracked commands from current session with status. Shows commands in CommandRegistry (in-memory, max 50). Useful for checking what's currently running or was recently executed in this session. For historical commands from database, use list_command_history instead.
Save a batch script to database (without executing). Saves script for later reuse. Automatically deduplicates based on content hash. Does NOT execute the script - use execute_script_content_by_id to run it.
Three similar execution tools (execute_command, execute_script_content, execute_script_content_by_id) lack clear distinction in names and descriptions. LLM will struggle to choose between them, risking wrong tool selection and wasted calls.
Error handling and recovery guidance is missing. Tools like execute_command and delete_batch_script do not document what errors can occur, what they mean, or how the LLM should recover. No 'retryable vs. fatal' classification.
Parameter constraints are incomplete. Many parameters lack explicit ranges, regex patterns, or enums where applicable. E.g., 'timeout' parameters have defaults but no documented min/max; 'sort_by' uses enum but other parameters do not. LLMs may pass invalid values without early feedback.
Parameter dependency documentation is sparse. E.g., 'conversation_id' is optional in multiple tools but its behavior when omitted is not clearly specified. Undocumented dependencies force LLMs to guess or trial-and-error.
get_command_output description is minimal ('WARNING: Uses more tokens...'). Does not explain when to use it vs. check_command_status, what format the output takes, or how to interpret 'raw' parameter. Ambiguous use cases.
Destructive operations (delete_batch_script, cancel_command) lack confirmation-step documentation in schema. While delete_batch_script mentions confirm in description, the mechanism is not formally modeled in the schema for clarity.
Tool descriptions reference internal concepts ('CommandRegistry', 'batch_executions table', 'step markers') that expose implementation details. Descriptions should focus on LLM-relevant behavior, not DB schema or internal state.