Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
Two tools with complete input schemas and output schemas defined in Zod. Tool names follow verb_noun pattern (extract_url, extract_many). Descriptions are present and substantive (180+ chars for both tools). Parameters have descriptions and constraints (url format, min/max integers, bounds on arrays). However, several patterns are underutilized: no error handling guidance, no examples of retryable vs fatal errors, no security notes for malicious URLs, no pagination or result limiting despite batch tool accepting up to 20 URLs. Output schemas are well-typed but descriptions lack detail on what each field represents (e.g., what is 'sourceStrategy'? What are typical values for 'warnings'?). The batch tool (extract_many) returns per-item success/failure which is good composition, but lacks guidance on partial failure handling. Parameter descriptions are in Chinese which reduces LLM clarity in English-dominant environments. Overall, solid foundational definitions with room for error handling, field documentation, and i18n.
Parameter descriptions are in Chinese; reduces clarity for English-speaking LLMs and breaks cross-language agent compatibility. Should be translated to English or provide bilingual descriptions.
Output schema fields (sourceStrategy, publishedAt, warnings, debug) lack descriptions. LLMs cannot determine what these fields mean or when they are populated. For example, 'sourceStrategy' is unclear, is it 'readability', 'playwright', 'fallback'? What are valid values?
No error handling guidance in tool descriptions. Errors like FETCH_ERROR, PLAYWRIGHT_UNAVAILABLE, NON_HTML_RESPONSE are defined in code but not surfaced to LLMs. No guidance on retryability, recovery steps, or when to fall back to alternative approaches.
extract_urlextract_many
Recommendations
Translate all parameter and output field descriptions to English (or provide bilingual descriptions) so non-Chinese LLMs can understand intent and constraints.
Add descriptions to every output schema field. For each field in the response (sourceStrategy, publishedAt, warnings, images, debug, etc.), explain what it contains and what values LLMs should expect.
Document error codes and recovery paths. Add a section to tool descriptions: 'On failure, this tool returns one of: FETCH_ERROR (network issue, retry with longer timeout), NON_HTML_RESPONSE (not a webpage, likely PDF or binary), EXTRACTION_ERROR (content extraction failed, check if playwrightFallback is enabled), PLAYWRIGHT_UNAVAILABLE (fallback disabled or browser not found). See error response for retryability flag.'
Add context window guidance to extract_many description. Example: 'Note: each extracted page can produce 5 - 30 KB of markdown depending on content. Extracting 20 pages may consume 100 - 600 KB of tokens. Consider smaller batches or filtering URLs by importance if context is constrained.'
Add security note to both tools: 'This tool fetches and processes arbitrary URLs. Ensure URLs are trusted and validated. Avoid passing user-supplied URLs without vetting. Consider rate-limiting and timeout parameters to prevent resource exhaustion.'
Document playwrightFallback behavior explicitly. Example: 'If initial extraction yields less than ~120 characters of content or detects JavaScript-required patterns (CAPTCHA, bot checks), Playwright browser automation is triggered as fallback. This is slower but handles JavaScript-heavy sites. Set to false if browser availability is not guaranteed.'
Spec posture evidence
Inferred effective spec: 2025-06-18+.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
extract_many accepts up to 20 URLs but no pagination support or result limiting. If a batch returns large markdown content for all 20 items, response could exceed context window. Tool description should warn about per-item size limits and token consumption.
No security guidance on malicious URLs. Tool accepts any URL format but description does not warn about risks of fetching from untrusted sources (SSRF, DNS rebinding, malicious HTML). Should recommend input validation and sandboxing.
playwrightFallback defaults to true with no explicit opt-out warning. If Playwright is resource-intensive or has licensing implications, LLMs should have clear guidance on when fallback is triggered and what the cost/benefit trade-off is.
extract_urlextract_many
Add example output (sanitized) to tool descriptions so LLMs understand the structure before calling. E.g., 'Returns object with url, title, author, markdown (converted page content), plaintext, images (array of URLs), warnings (array of issues encountered).'
Consider adding a per-tool example in the description (not in parameter defaults). For extract_url: 'Example: extract_url with url="https://example.com/article" returns {title: "Article Title", author: "Jane Doe", markdown: "# Article Title\n\nContent here...", warnings: ["Missing author metadata"]}', sanitized, not literal values.
Add idempotency note: 'extract_url is idempotent, calling it multiple times with the same URL returns the same content (barring live site changes). Safe to retry on timeouts or transient errors.'
Document concurrency behavior for extract_many: 'Concurrent extractions are limited by the concurrency parameter (default 3, max 8). Each concurrent extraction uses separate fetch and optional Playwright instance. Higher concurrency speeds up batches but increases memory and CPU usage; adjust based on available resources.'
Return structured error responses with actionable guidance. Current code already has ERROR_CODES and retryable flags in extract_many, ensure these are fully documented in tool descriptions so LLMs know how to interpret failure responses.