Map your application's AI attack surface from source: MCP servers (with security audit), agent frameworks, LLM call sites, model gateways, AI infrastructure, provider keys, and HTTP/OpenAPI endpoints. Local, static, CI-friendly, with a visual UI.
ai-surface exposes 2 tools with generally clear naming and descriptions. Both tools follow verb-noun conventions (scan_, check_) and have substantive descriptions (180+ chars each). Input schemas are visible and typed. However, output schemas are not documented in the visible source, parameter descriptions lack specificity (e.g., no format/range constraints), and error handling guidance is absent. The tools are read-only with low risk, but descriptions could be more actionable for LLM selection. No critical gaps, but notable omissions in schema documentation and error recovery patterns.
Compare the current AI attack surface at PATH against a rolling baseline and report only what is NEW, MODIFIED, or REMOVED, with the compliance mapping for each. Call it right after edits to catch AI surface a change just introduced (a new MCP server, a new agent tool, a new LLM call). The first call records the baseline and reports nothing changed; each later call reports the delta since the previous call. Pass baseline_file to compare against a specific committed baseline instead (that file is not modified).
Inventory the AI attack surface in code at PATH: LLM SDK calls, agents and their tools, MCP servers, RAG / vector stores, model gateways, provider keys, and the API endpoints that expose them. Returns counts by category and severity, the top risks, the governance frameworks the surface maps to, and the highest-severity surfaces (capped by max_surfaces). Read-only, offline, no data leaves the machine.
Output schemas not documented. Both tools lack explicit documentation of response structure, fields, and types. LLMs cannot plan downstream operations or extract relevant data without knowing the output shape.
Parameter descriptions lack actionable constraints. 'path' param is described as 'Directory path to scan' with a default of '.' but does not specify: required format (absolute vs relative), character restrictions, size limits, symlink handling, or depth constraints. Similarly, 'max_surfaces' has no bounds (min/max), and 'baseline_file' lacks existence/permission requirements.
No error recovery guidance. Descriptions do not explain what to do if scanning fails (permissions, malformed configs, missing files, timeout). Agents have no actionable error context, no retry hints, no fallback suggestions, no classification of errors as user-fixable vs fatal.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2026-07-28+ | v2 |
Parameter 'baseline_file' in check_new_ai_surface lacks clarity on mutual exclusivity with stateful baseline tracking. Description says 'compare against a specific committed baseline instead (that file is not modified)' but does not explain: What happens if both a rolling baseline AND baseline_file are present? Is baseline_file mutually exclusive with the rolling baseline mode? How does an agent choose between them?
scan_ai_surface description promises 'highest-severity surfaces (capped by max_surfaces)' and 'top risks' but does not define what constitutes a 'risk' or how severity is ranked. LLMs cannot reason about prioritization without understanding the severity scale (e.g., is it 1-5, low/medium/high, or quantitative?).