Give your AI agent eyes (screenshot live pages), ears (Sentry/Datadog/Rollbar errors + runtime checks), hands (verify fixes worked), and root-cause tools (source-map trace resolution, git regression blame), and debug tools (run tests, stream logs, query your DB, call APIs). 24 tools, 121-module engine. Scan, fix, screenshot, and prove the fix worked — all inside your AI assistant.
GateTest MCP server has 24 tools with highly variable quality. Many tools have descriptions (good), but critical gaps exist: (1) Most input schemas lack detailed parameter descriptions required by the rubric, parameters are defined with type and name but descriptions are minimal or generic. (2) No output schemas are documented anywhere in the visible code, users cannot see what fields to expect from tool results. (3) Parameter descriptions are either missing or single short words that do not meet the 10-1024 character guideline or explain constraints. (4) High-risk tools like 'fix_issue', 'http_request', and 'query_db' lack proper error handling, validation rules, and security guardrails. (5) Tool names are generally good (verb-noun pattern: get_*, scan_*, run_*, list_*), but some lack clarity: 'stream_logs' vs 'tail_logs', 'blame_regression' is colloquial and unclear. (6) No evidence of idempotency guarantees, confirmation flows for destructive operations, or per-item failure reporting for batch tools. Average tool score: 42/100.
Query past local scans in the memory store
Which git commit introduced this line — read-only, ranks candidates across a stack trace
👁 Eyes — screenshot any live URL or localhost so the AI sees the rendered page
Verify GateTest engine is operational
Cross-repo prior-art lookup via memory store
Render a PR body for a set of fixes
Forensic-tier AI diagnosis per finding
No output schemas documented for any tool. Users cannot see what fields to expect in responses, forcing them to guess or trial-and-error. This violates pattern:tool and pattern:response-shaper, requiring LLMs to parse unstructured output and wasting tokens.
Parameter descriptions are minimal or missing. Tools like 'get_badge', 'scan_url', 'capture_screenshot' define parameters with type and name but no description of what the parameter controls, expected format, or constraints. LLMs cannot infer meaning from 'url' or 'repo' alone, should include format hints (e.g., 'URL must start with http(s)://', 'Repository ID (owner/repo format or full GitHub URL)').
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
AI-driven auto-fix for a specific finding
Get embeddable README badge for any repo
👂 Ears — top Sentry / Datadog / Rollbar errors with file:line attribution
Retrieve full result of the last scan this session
👁 Eyes — baseline vs current visual diff to catch UI regressions
🤝 Hands — call any API with auth headers, follow redirects, inspect responses
List all 121 modules with descriptions
🤝 Hands — read-only SQL/NoSQL queries (Postgres/MySQL/SQLite/MongoDB/Redis)
Minified/bundled stack trace → original file:line via source maps
👂 Ears — runtime errors, console warnings, API health on any URL
Run one specific module against a path
🤝 Hands — auto-detect + run the project's test suite (Jest/Vitest/pytest/cargo/go)
Run local scan — quick (50-module), standard (100+ module), full (121-module), or diff-aware smart scan
Scan any public git repo via the hosted API
Quick scan any live URL via hosted API
🤝 Hands — tail a running process or log file in real time (up to 60s)
🤝 Hands — re-scan changed files for hard pass/fail proof the fix worked
High-risk tool 'fix_issue' (WRITE risk) lacks error handling guidance, confirmation flow, and rollback strategy. No documentation on what happens if the fix fails midway, how to verify it worked, or what to do if unintended side effects occur. This violates pattern:confirmation-request and pattern:recovery-guide.
Tool 'http_request' exposes credentials as parameters ('headers' can contain Authorization tokens). This violates pattern:secret-injection, credentials must never appear as tool parameters and should use server-side secret injection. Agents log all parameters; secrets in params leak into traces.
Tool 'query_db' accepts a 'connectionString' parameter, exposing database credentials directly. This is a critical security violation, credentials must be injected server-side via environment variables or vault, never as tool parameters. See pattern:secret-injection.
List tools ('list_modules', 'get_production_errors') lack pagination parameters (limit, offset, page_size) and do not document result caps. 'list_modules' claims to return '121 modules' with no paging, if the agent needs to iterate over results, a single unbounded call risks context window exhaustion. See pattern:paginated-result.
Tool naming lacks clarity in a few cases: 'blame_regression' is colloquial and unclear (should be 'identify_regression_commit' or 'find_commit_for_line'). 'stream_logs' is ambiguous, does it mean 'tail_logs_realtime' or 'fetch_logs_stream'? Unclear names force LLMs to guess intent.
No evidence of idempotency guarantees or retry safety documentation. Tools like 'fix_issue', 'run_tests', 'http_request' do not declare whether repeated calls with identical inputs are safe. Agents retry on ambiguous failures, non-idempotent tools risk duplicate side effects (double API calls, duplicate records, repeated charges).
No error classification or recovery guidance. If a tool fails, LLMs do not know whether to retry, ask the user, or give up. Error responses should categorize failures: retryable (e.g., timeout), user-fixable (e.g., invalid input), fatal (e.g., permission denied). See pattern:recovery-guide.
Parameters like 'suite' in 'scan_local' use enums ('quick', 'standard', 'full', 'smart'), which is good, but enum descriptions are missing. LLMs do not know the difference between 'quick' and 'standard', should include: 'quick: 50-module subset for CI (fast); standard: 100+ modules (balanced); full: all 121 modules (thorough); smart: diff-aware, only changed files'.
Tool 'stream_logs' accepts a 'duration' parameter (max 60s) but does not explain consequences of the timeout. Does the tool return partial logs? Truncated output? Error? If max duration is 60s, real-time tailing of long-running processes may be cut off unexpectedly. Should document: 'Max 60s; returns all logs collected up to timeout, or error if stream is silent beyond 60s'.