OFAC sanctions screen, KYA, and transaction risk for AI agents — MCP + HTTP + CLI. Real OFAC data, no API key.
Mixed quality across 10 tools. Strengths: all tools have descriptions (50-180 chars), most input parameters typed and described, clear business logic (email/SMS verification, sanctions screening). Weaknesses: no output schemas documented in the provided code snippet; several parameters missing explicit type constraints (e.g., 'country' and 'service' in create_number accept free-form strings instead of enums); some parameter descriptions are generic or lack actionable constraints; no error handling patterns visible (what happens if no SMS arrives? no code extraction possible?); no confirmation/dry-run pattern for destructive operations (release_number, dispute_open); weak composability signals (tools return objects but chaining requirements unclear). Tool names are reasonably clear (verb_noun pattern mostly followed) but descriptions vary in LLM-optimization quality.
Create a fresh disposable email inbox (or reuse an existing one by label). Use this when an agent needs an email address to sign up / receive a verification message. Returns the address. Poll fetch_code to get the OTP.
Rent a phone number that can receive SMS/OTP codes. Use this when an agent must verify via phone (WhatsApp, Telegram, banks, apps requiring SMS). Returns the rented `number`. Poll fetch_sms for the OTP. `country` e.g. "usa","russia","any"; `service` e.g. "discord", "google","any" (provider may use it to pick a suitable number).
Open a dispute when an agent-paid transaction went bad (non-delivery, fraud). Records the dispute with a 7-day auto-escalation window. Phase 1 = registry + notification; Phase 2 will integrate agentcourt-api for arbitration. Returns: {dispute_id, status, escalation_at}
Fetch the latest verification code/link from an inbox. `wait` seconds for a message to arrive (polls). Optional `match_from` / `match_subject` substrings filter which message counts. Returns extracted `code` and any verification `link`, or {"empty": true}.
Fetch the latest SMS (OTP) for a labeled number. Polls up to `wait` seconds for a message to arrive. Returns extracted `code` and sender `from`, or {"empty": true}.
No output schemas documented for any tool. The code snippet shows input schemas but does not specify what fields callers should expect in responses (e.g., does fetch_code return {code: string, link: string, from: string, subject: string} or a different shape?). LLMs cannot plan downstream tool calls or extract needed fields without documented output shapes.
create_number accepts 'country' and 'service' as free-form strings with example values in description ('e.g. usa, russia, any' and 'e.g. discord, google, any'). LLMs will treat these as the only valid options or hallucinate values outside the intended set. Should use enum constraints instead.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 55 | 2026-07-28+ | v2 |
Verify an AI agent's identity before transacting with it (Know Your Agent). Call BEFORE paying or trusting another agent. Returns a trust score + flags. evidence keys (any subset): wallet_address, wallet_age_days, domain, pubkey, owner_email, declared_country. Higher trust = more verified attributes and a clean sanctions screen. Use the returned recommendation to decide whether to proceed with a counterparty agent. Returns: {trust_score: 0-100, verified: [...], flags: [...], recommendation}
List all existing labeled inboxes/numbers.
Stop renting a phone number and free it. Call when you're done with OTP.
Score a transaction's fraud risk BEFORE authorizing payment. Call right before an agent pays. Combines counterparty signals + amount anomalies + sanctions screen + rail/category heuristics. Recommendation is one of: allow / review / decline. 'decline' = abort the payment. rail in: x402, ap2, acp, tap. category in: digital_goods, services, physical. Returns: {score: 0-100, recommendation, reasons: [...], screen_id}
Screen a counterparty against OFAC/EU/UN/UK sanctions lists. Cheapest check, call first. At least one of name / wallet / country required. Free provider uses open sanctions data (no key). Useful as a fast pre-filter before the heavier risk_score call. Returns: {matches: [{list, entity, match_type, confidence}], clean: bool}
risk_score 'rail' parameter states 'x402, ap2, acp, tap' as a string default but does not explicitly declare it as an enum. Same pattern: category lists 'digital_goods, services, physical' as free-form string. Both should be enums to prevent hallucination.
sanctions_check 'name', 'wallet', 'country' parameters have empty or minimal defaults ('') and vague descriptions. Does the tool require at least one to be non-empty? The description says 'At least one of name / wallet / country required' but the schema does not enforce this, parameter descriptions do not state the mutual-requirement constraint clearly.
fetch_code and fetch_sms both poll for messages but return structure not documented. If a message does not arrive within the wait timeout, does the tool return {empty: true} or throw an error? The description mentions this but no error-handling pattern is visible in the code provided.
release_number, dispute_open, and create_number are WRITE operations (state-altering) but lack error-handling guidance, confirmation patterns, or dry-run support. An LLM could accidentally release a number before fetching the OTP, or open a dispute erroneously. No confirmation-request or pre-flight validation pattern visible.
release_number description is only 59 characters: 'Stop renting a phone number and free it. Call when you're done with OTP.' Does not explain when or why to call it, what happens if you don't, or what errors might occur. Too terse to guide LLM decision-making.
create_inbox, fetch_code, list_inboxes, create_number, fetch_sms are composable as an email/SMS verification chain, but chaining requirements are not documented. Does create_inbox return an 'id' field that fetch_code expects in 'label'? The 'label' parameter is described as 'Label of the inbox' but is it opaque, human-readable, or a returned UUID?
kya_verify 'evidence' parameter is type 'object' with description 'Evidence dictionary with optional keys: wallet_address, wallet_age_days, domain, pubkey, owner_email, declared_country.' Does not specify types for each key, allowed ranges (e.g., wallet_age_days an integer in days?), or what trust_score/flags values are returned. Output schema missing.
dispute_open 'evidence' parameter accepts 'object|null' but does not document what keys are expected, their types, or the purpose. Description says 'Optional evidence dictionary' but provides no schema or examples. LLMs cannot reason about what evidence to provide without field definitions.