A prompt injection detection server that fetches URLs and scans their content for injection attempts using a cascading multi-tier classification system with Claude models.
The server exposes a single tool 'scan_url' with a well-defined purpose: fetch and scan URLs for prompt injection attacks. The tool has a clear, actionable description (113 chars), proper naming following verb_noun convention (scan_url), and a complete input schema with type and description. However, there are notable gaps: the output schema is not documented in the code, error handling returns raw JSON without actionable recovery guidance, and the tool lacks annotations (readOnlyHint, idempotentHint) that would help agents reason about safety. The implementation is functionally solid but missing patterns expected in production-grade tools.
Fetch a URL and scan its content for prompt injection. Returns clean content or a structured rejection. Use this instead of WebFetch when consuming untrusted web content.
Output schema not documented. The tool returns different structures based on scan result (clean vs. injection detected vs. error), but the response schema is not formally specified. LLMs cannot plan downstream operations without knowing what fields to expect.
Error messages lack recovery guidance. When the tool fails (e.g., HTTP error, SSRF block, content type rejection), it returns raw error JSON without actionable next steps. 'URL rejected by SSRF guard' tells the agent the URL was blocked, but not what to do, retry with a different URL? Contact the admin? Ask the user?
Missing tool annotations. The tool is read-only and idempotent (scanning the same URL twice produces the same result), but lacks readOnlyHint and idempotentHint annotations. Agents cannot infer these properties and may incorrectly treat the tool as stateful or destructive.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 57 | - | v1 |
Response structure ambiguous in code. The handler returns different response shapes (content[0].type='text' with JSON body) for success and error cases. The JSON body is not formally documented, forcing LLMs to parse unstructured text.