AI-powered OSINT agent and MCP server that maps digital footprints across 3,400+ public data sources. Exposes 30+ OSINT tools via MCP protocol for use with Claude and other MCP-compliant clients.
The server defines 30 OSINT tools with consistent structure. All tools have descriptions (10 - 150 chars, mostly 50 - 120) and input schemas with typed parameters. However, descriptions are largely generic and lack LLM-optimized guidance on WHEN to call each tool, dependency ordering, or expected output structure. Most tools accept a 'json_output' boolean parameter but lack documentation of what the JSON structure contains. No output schemas are documented in the code. Parameter descriptions are minimal (e.g., 'Target email address' with no guidance on format, validation, or fallback behavior). Error handling and recovery guidance are absent from tool definitions. Tool naming is consistent (verb-noun pattern) and clear, but many tool names are similar in intent (e.g., search_email, search_breach, search_emailrep, search_gravatar all target emails), creating potential LLM confusion without explicit differentiation in descriptions.
Generate targeted Google dork URLs for reconnaissance of any target string (name, email, username, domain).
Fetch and extract full text content from a URL via Bright Data (residential proxy) to bypass blocks and paywalls. Requires BRIGHTDATA_API_KEY.
Query AbuseIPDB for IP abuse reports, blacklist status, threat type. Requires ABUSEIPDB_API_KEY.
Check if an email appears in data breaches via HaveIBeenPwned. Requires HIBP_API_KEY env var.
Query Censys for host/service and certificate data (IP, port, services, technologies, certificates). Requires CENSYS_API_ID and CENSYS_API_SECRET.
Enumerate subdomains from certificate transparency logs via crt.sh. Keyless and purely passive (public CA logs), surfaces internal/staging hosts that never resolve publicly.
Output schemas not documented. All 30 tools accept a 'json_output' boolean but no tool defines what the JSON structure contains (fields, types, nesting). LLMs cannot plan downstream tool calls or parse responses reliably.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 57 | <=2025-11-25 | v2 |
Validate a Bitcoin or Ethereum address and return a keyless on-chain summary (balance, transactions, total received).
Perform DNS lookups (A, AAAA, MX, TXT, NS, SOA, CNAME records) for a domain.
Enumerate subdomains of a target domain using sublist3r.
Execute Google dorks live via Bright Data (residential proxy) and return raw HTML snippets. Returns top 10 results. Requires BRIGHTDATA_API_KEY.
Enumerate accounts linked to an email using holehe.
Email reputation and footprint summary via EmailRep.io (profiles seen, breach/abuse flags). Requires EMAILREP_API_KEY.
Extract embedded metadata (EXIF/IPTC/XMP) from a local file using exiftool: camera make/model, software, timestamps, author, and GPS coordinates. Flags embedded GPS location.
Risk-ranked IP exposure report (geolocation, ASN, reverse-DNS, DNS blocklists, VPN/Tor flags). Omit 'ip' (or pass 'me') to check the caller's own public IP.
Comprehensive footprint via Bright Data: search dorks live, find files, web assets, and leaks across index + dark/paste web. Combines dorks + scraping + OSINT. Requires BRIGHTDATA_API_KEY.
Search GitHub for username, email, API keys, secrets, and code snippets (public repos).
Look up an email's public Gravatar profile: avatar, display name, bio, location, and linked/verified accounts.
Check an IP against GreyNoise: internet scanning activity, known malicious behavior, VPN/datacenter status. Requires GREYNOISE_API_KEY.
Passive organisation/domain recon via theHarvester: emails, people, and subdomains from public sources (passive only, no active probing).
Check if an email or IP appears in HudsonRock's stealer log (credential theft, malware, dark web). Requires HUDSONROCK_API_KEY.
Retrieve geolocation and ASN data for an IP address via ipinfo.io. Omit 'ip' (or pass 'me') to auto-detect the caller's own public IP.
Retrieve geolocation, VPN/proxy, and threat data for an IP address. Requires IP2LOCATION_API_KEY.
Broad username/identity discovery across 3,400+ sites via maigret (also extracts profile details).
Search Pastebin dumps for mentions of an email or username.
Gather carrier, country, and line type data for a phone number. Use E.164 format.
Query Shodan for host intelligence or banner searches. If the query looks like an IP address, performs a host lookup. Otherwise performs a keyword/service search. Requires SHODAN_API_KEY.
Enumerate and verify platforms where a username is registered, using sherlock plus a WhatsMyName subset of modern/niche sites.
Query VirusTotal for file hashes, URLs, domains, or IPs (security scanning, threat intelligence, detection ratio). Requires VIRUSTOTAL_API_KEY.
List URLs archived under a domain in the Internet Archive (Wayback Machine) via the keyless CDX API. Passive; recovers deleted, forgotten, or historical pages. Pairs with scrape_url to fetch a recovered page.
Retrieve WHOIS registration data for a domain.
Tool descriptions lack LLM-optimized guidance on WHEN to call each tool vs similar alternatives. E.g., search_email, search_breach, search_gravatar, search_emailrep all target emails but descriptions do not clarify which to use first or when they complement each other. LLMs will pick arbitrarily or call all of them, wasting tokens.
No error handling or recovery guidance in tool definitions. Tools that require API keys (HIBP_API_KEY, EMAILREP_API_KEY, GREYNOISE_API_KEY, HUDSONROCK_API_KEY, BRIGHTDATA_API_KEY, SHODAN_API_KEY, VIRUSTOTAL_API_KEY, CENSYS_API_ID, CENSYS_API_SECRET, IP2LOCATION_API_KEY, ABUSEIPDB_API_KEY) do not document what happens if the key is missing or invalid. LLMs have no recovery path.
No parameter constraints or format guidance. E.g., search_phone expects E.164 format but description says only 'Target phone number in E.164 format (e.g. +14155552671)', LLMs may pass values like '415-555-2671' or '+1-415-555-2671' that violate the constraint and fail silently.
Similar tool names risk LLM confusion. search_email (holehe), search_username (sherlock), and search_maigret (identity discovery across 3,400+ sites) overlap in purpose but lack discriminating descriptions. search_ip (ipinfo.io) vs search_exposure (risk-ranked report) also share target type with no clear decision tree.
No documentation of pagination, result limits, or rate-limiting behavior. Tools like search_maigret and search_dorks_live may return large datasets; no tool description mentions result caps, pagination support, or guidance on when results are truncated.
No distinction between tools that require external API keys and tools that are keyless/passive. All tool descriptions treat API-keyed tools the same as public-data tools, risking LLM calls to unavailable services. Descriptions should state 'Keyless; purely passive (no credentials required)' for tools like search_crt, search_wayback, search_dns.
json_output parameter is present on all tools but undocumented. Descriptions do not explain what json_output=true produces vs json_output=false. LLMs cannot make an informed choice.