Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server has 9 tools across 3 distinct domains (lynx, Chrome, tmux). Tool naming is generally verb-based and clear (view_webpage, screenshot, send_keys). However, there are critical gaps: (1) Parameter descriptions are present but often terse or generic; (2) Output schemas are not explicitly documented in any tool, LLMs cannot see what fields they'll receive; (3) Error handling is minimal, most tools will raise exceptions rather than guiding recovery; (4) Some tools (screenshot, send_keys) perform state-modifying operations without confirmation or dry-run options. The lynx tools lack input validation (what happens if the URL is malformed?). The tmux tools are better structured with docstrings, but dump_all_panes returns a complex nested structure without clear field documentation. Overall, this is a functional but underspecified toolkit typical of community servers.
Tools (9)
capture_full_paneread onlysource verified75/100
Capture the full scrollback buffer of a tmux pane.
No output schemas documented for any tool. LLMs cannot reason about what fields they'll receive, forcing them to guess at response structure and invoke downstream tools with potentially incorrect field names.
Error handling is absent. Tools like view_webpage and screenshot will raise raw exceptions (e.g., lynx command not found, pyppeteer launch failure) with no guidance for recovery. Per the rubric, errors must tell the LLM what to do next.
State-modifying tools (screenshot, send_keys) lack confirmation/dry-run options. An LLM could accidentally fill a disk with screenshots or send unintended keystrokes without a way to preview or confirm first.
Recommendations
Document output schemas for all 9 tools. For lynx tools, specify that output is a formatted string with returncode, stdout, stderr fields. For tmux tools, document that list_window_ids returns a list of strings (window IDs), capture_visible_pane returns a multi-line string, and dump_all_panes returns a list of dicts with keys: window_id, window_name, window_index, pane_id, pane_index, pane_active, pane_current_command, pane_title, pane_current_path, contents.
Add comprehensive error handling to lynx and Chrome tools. Wrap subprocess calls in try-except blocks and return structured error responses (not stack traces). E.g., 'lynx command not found. Install lynx and ensure it is in PATH.' for missing executables; 'Invalid URL format' for malformed URLs with guidance to try a different URL.
Add a dry-run parameter to screenshot and send_keys (e.g., dry_run=False). When dry_run=True, return what WOULD happen (e.g., 'Would take screenshot of https://example.com and save to /tmp/test.png') without executing. This lets LLMs preview before committing.
Expand parameter descriptions to 50 - 100 characters each, following the pattern: 'What it is (type/format). When/how to use it. Valid values or constraints.' E.g., 'The tmux pane ID (format: %0, %1, etc.). Found by calling list_pane_ids first. Alphanumeric + %.' instead of just 'The tmux pane ID'.
Add input validation and actionable error messages. For view_webpage, validate that url starts with http:// or https://. For send_keys, validate that keys_to_send is non-empty. Return structured error objects: {error: 'Invalid URL', expected: 'http:// or https://', got: '<url>'} instead of raising exceptions.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 16 points across a rubric change (v1 → v2)
50/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
50
<=2025-11-25
v2
2026-03-09
F
34
-
v1
send_keyswritesource verified75/100
Send keys to a tmux pane. Set enter=False to avoid sending <Enter>. Returns the visible contents after sending.
dump_all_panes returns a List[Dict[str, object]] with undocumented field names and types. The return type annotation is present in code, but there is no docstring describing the schema of each dict (window_id, pane_id, contents, etc.). LLMs will struggle to extract the right fields.
Parameter descriptions are present but often minimal. E.g., 'The tmux pane ID' for pane_id doesn't explain the format (e.g., '%3', '%0') or where to get one. Per the rubric baseline of 72 chars average, these descriptions are below target.
dump_all_panes accepts a nullable max_lines parameter but provides no validation guidance in the description. What happens if max_lines is negative or zero? The description should clarify.
Lynx tools (view_webpage, duckduckgo_search) do not validate URLs or query strings. A malformed URL will cause lynx to fail silently or produce a cryptic stderr. No guidance for the LLM to retry or escalate.
view_webpageduckduckgo_search
Document the return type of dump_all_panes inline: 'Returns a list of dicts, one per pane: {window_id, window_name, window_index, pane_id, pane_index, pane_active (bool), pane_current_command (str), pane_title (str), pane_current_path (str), contents (str)}.' This is already in the code docstring but not visible to LLMs that only read parameter/output docs.
For dump_all_panes max_lines parameter, clarify: 'If provided, return only the last N lines (1 or more). If None (default), return all lines. Use max_lines to truncate very large scrollback buffers and reduce token usage.'
Add a required_permissions or security note to send_keys: 'This tool sends keystrokes to a tmux pane, which can execute any command. Use with caution and only to controlled environments (e.g., test shells).' This surfaces the risk to agents.
For lynx tools, add fallback or retry guidance: 'If the page cannot be loaded, verify the URL is accessible and lynx is installed. Try a simpler URL (e.g., example.com) to test connectivity.'
Consider splitting dump_all_panes into two tools: dump_panes_summary (returns just {pane_id, pane_title, pane_current_command} for quick discovery) and dump_pane_full (returns complete {..., contents} for a specific pane). This follows the single-responsibility principle and reduces token overhead for discovery calls.