MCP server exposing SCOUT firmware analysis capabilities as tools for AI agents to drive firmware analysis without shell access
SCOUT presents 15 tools with complete input schemas and reasonable descriptions. However, descriptions are often terse (many under 100 chars), lack explicit usage guidance, and omit dependency information. Parameter descriptions are minimal, most params have type hints but no human-facing guidance on constraints or valid formats. Output schemas are not documented in the tool definitions; agents cannot predict return structures. Error handling is implicit rather than explicit. The server follows a read-heavy, analysis-oriented pattern (11 READ_ONLY, 4 WRITE tools) with reasonable naming conventions, but lacks the rich guidance and validation rules expected of production tools. Overall, this is a solid domain-specific tool suite (firmware analysis) with competent structure but insufficient LLM-optimization for multi-step reasoning.
Start a new firmware analysis (launches in background).
Get attack surface analysis summary.
Get binary analysis info including hardening status (NX, PIE, RELRO, Canary).
Get X.509 certificate analysis results.
Look up CVE matches for firmware components.
Get the full reasoning_trail and triage context for a specific finding. Returns reasoning_trail (PR #11), category (PR #7a), original and current confidence, fp_verdict and triage_outcome. Useful for analysts auditing LLM debate transcripts.
Get communication graph showing service relationships and IPC channels.
Output schemas not documented. Agents cannot predict return structure (fields, types, pagination, chaining IDs). LLMs must guess what get_finding_reasoning or scout_binary_info return, increasing hallucination and downstream tool mismatch.
Tool descriptions lack explicit WHEN guidance. E.g., scout_list_findings is described as 'List security findings from a run, optionally filtered by severity, evidence_tier, and minimum confidence', but does not explain when to call it vs scout_get_finding_reasoning, or what filters enable faster/cheaper filtering. LLMs cannot reason about tool selection.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 63 | 2026-07-28+ | v2 |
Inject an analyst hint for a finding into the feedback registry. Next time adversarial_triage runs with the same AIEDGE_FEEDBACK_DIR, the advocate prompt is prefixed with the hint text. Priorities: low, medium (default), high.
List security findings from a run, optionally filtered by severity, evidence_tier, and minimum confidence.
List available SCOUT analysis run directories.
Override the verdict for a specific finding. Sets fp_verdict (false_positive, true_positive, undecidable) and optional triage_outcome and notes.
Read a stage artifact file (JSON or text, max 30 KB).
Re-run specific stages on an existing analysis run.
Get CycloneDX SBOM (Software Bill of Materials) from a run.
Get status of stages in a run. Returns stage name, status (ok/partial/failed/skipped), duration, and limitations.
Parameter descriptions are minimal or absent. 'run_id' is described as 'Run directory name (e.g. 20250101_120000_abc123)' in scout_stage_status, but agents don't know if they should construct this or retrieve it from scout_list_runs. 'stage' parameter in scout_stage_status is 'Optional: specific stage name to inspect', what are valid stage names? Is there a discovery tool?
No error recovery guidance. If scout_analyze fails (e.g., firmware_path invalid, case_id conflict), the tool returns an error with no actionable next steps. LLMs cannot self-correct without explicit recovery hints like 'If firmware_path not found, call scout_list_runs to verify run directories' or 'Use scout_run_stage to retry failed stages.'
Missing parameter validation rules in descriptions. scout_analyze accepts 'no_llm' boolean with default=true, but the description does not explain what LLM calls do or why an agent would override the default. scout_list_findings has min_confidence (0.0 - 1.0) but does not explain if the default is 0.0 (accept all) or something else.
No documented pagination or result limits. scout_list_findings and scout_list_runs do not document max result count or pagination support. If a run has 1000+ findings, does scout_list_findings return all or cap at a limit? Without documented pagination, agents cannot safely iterate large result sets.
WRITE tools lack explicit destructive intent in descriptions. scout_run_stage says 'Re-run specific stages on an existing analysis run' but does not clarify if this overwrites prior stage results or appends new ones. scout_override_verdict says 'Override the verdict for a specific finding' but does not warn that this modifies audit records irreversibly.
No dry-run or confirmation pattern for irreversible operations. scout_override_verdict and scout_inject_hint permanently modify findings and feedback registries. No tooling for preview, rollback, or confirmation, agents risk overwriting analyst decisions without review.
Tool chaining metadata not returned. If an agent calls scout_list_findings and then wants to call scout_get_finding_reasoning for a specific finding, does scout_list_findings return finding_id? No documented output schema to verify. Similar for scout_list_runs → scout_stage_status → scout_read_artifact chain.