MCP server for advanced binary analysis using angr, FLARE tools (floss, capa), and specialized forensic analysis capabilities
Arkana defines 6 analysis tools for binary reverse engineering via angr. Tool naming follows verb_noun pattern well (list_, get_, disassemble_, diff_). Descriptions are present but lack critical context: they state WHAT the tool does but rarely explain WHEN to use it, what the output structure is, or how it chains with other tools. Most descriptions are 100-150 characters, acceptable length but lack LLM-optimized clarity. Input schemas are visible with type declarations, but output schemas are completely undocumented, LLMs cannot predict what fields to extract from results. Parameters have descriptions but lack format constraints (e.g., 'function_address' says 'Hex address' but doesn't specify '0x'-prefix format or validation rules). All tools are READ_ONLY which is good for safety, but no tool declares idempotence explicitly. Error handling is not visible in the provided code, no recovery guidance, retry classification, or actionable error messages shown.
[Phase: advanced] Compares the loaded binary against another to find matching, differing, and unmatched functions. For patch diffing and variant analysis.
[Phase: deep-dive] Disassembles raw instructions at any address — not limited to known functions. Useful for shellcode, data-as-code, or arbitrary offsets.
[Phase: deep-dive] Recovers calling conventions, parameter counts, and return types for functions.
[Phase: deep-dive] Recovers local variables and parameters for a function — names, sizes, stack offsets or register locations, and access counts.
[Phase: advanced] Computes reaching definitions for a function — which register and memory definitions reach each program point.
[Phase: context] Discovery tool: lists all available angr-based analysis capabilities with descriptions. Use this to understand what angr analyses are available before calling specific tools.
Output schemas completely undocumented. LLMs cannot predict what fields each tool returns, forcing them to guess or hallucinate field names for downstream composition. This violates pattern:tool-description and pattern:response-shaper.
Input parameter descriptions lack format constraints and validation rules. 'function_address' says 'Hex address' but doesn't specify '0x' prefix requirement, bit-width, or validation behavior on invalid input. This violates pattern:constrained-input and review:param-validation-rules.
Tool descriptions lack WHEN-to-use context and dependency hints. E.g., get_reaching_definitions doesn't explain: 'Call list_angr_analyses first to discover available analyses. Requires a loaded binary.' This makes discovery inefficient and forces agents to experiment.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 54 | - | v1 |
No error handling guidance visible in provided code. Tools should return actionable error messages (e.g., 'Function at 0x401000 not found. Available functions: [list]') per pattern:recovery-guide, but none is shown.
Parameter 'category' in list_angr_analyses is a free-form string with documented enum values ('all', 'decompilation', 'cfg', etc.) but no JSON Schema enum constraint. This invites hallucinated category names from LLMs. Violates pattern:constrained-input.
Numeric parameters lack explicit bounds. 'limit' appears in multiple tools but no min/max constraints are declared (should be 1 - 1000 or similar). Unbounded limits allow LLMs to pass absurd values (limit=999999) per review:param-validation-rules.
Tool descriptions mention phases ('[Phase: context]', '[Phase: advanced]', '[Phase: deep-dive]') but never explain what these phases mean. These metadata tags are opaque to LLMs, either remove them or document their significance.