Wayback Machine for agents: closest archived copy of a URL with its text, and full capture history
Strong tool definitions with excellent descriptions (150-250 chars each), comprehensive parameter schemas using Zod with proper types and constraints, and clear use-case guidance. All three tools have detailed descriptions explaining what they do and when to use them. Parameters include proper validation (minLength, maxLength, pattern, enum, default). Tool annotations present (readOnlyHint, destructiveHint). Main gaps: no documented output schemas, no error handling guidance in descriptions, and task_context is required but not all agents may provide it.
Closest archived copy of a URL from the Wayback Machine (Internet Archive), nearest to a date if given. Returns the snapshot URL, its exact timestamp, the raw-bytes URL, and optionally the page text (HTML stripped, capped at 30k chars). Falls back to the Common Crawl index when the Wayback Machine has no capture. Use when a page is dead, changed, or walled and you need what it said.
Every capture the Wayback Machine holds for a URL, URL prefix, host or whole domain (CDX index). Filter by status code or MIME type, keep one capture per period (collapse), restrict a date window, and page with a resume key. Use for site archaeology: what a site ever had, when a page changed, every URL under a path.
Archive a live URL now in the Wayback Machine (Internet Archive Save Page Now API). Returns the job ID and status; poll the job endpoint to wait for completion. Use when you need a fresh capture of a live page, or to ensure a page is archived before it changes or disappears.
Output schemas not documented. Tool descriptions explain what is returned (snapshot URL, timestamp, text, etc.) but no formal schema is visible in code or docs. LLMs cannot plan downstream operations without knowing response structure.
task_context parameter is required (minLength:1) but not all agent frameworks provide it. This could cause tool invocation failures if an agent omits it. Consider making it optional with a sensible default.
Error handling guidance missing from descriptions. No mention of what happens on upstream failures (Wayback Machine down, Common Crawl index stale, rate limits). LLMs need recovery hints like 'If no capture found, try a broader date range' or 'If rate limited, wait 60s and retry'.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 83 | 2026-07-28+ | v2 |
capture_history filter parameter uses free-form strings (e.g. 'statuscode:200') rather than structured objects. LLMs may hallucinate invalid filter syntax. Consider an enum or structured filter object.
save_page_now lacks confirmation/dry-run pattern. This is a WRITE operation that queues an archive job. No mention of cost, rate limits, or whether the operation is idempotent. Agents should know if retrying is safe.