A powerful MCP server to equip LLMs with web access, search, and content extraction capabilities
Server defines 3 web tools with clear action-verb naming (search_web, fetch_url, view_website) and comprehensive parameter schemas using Pydantic Field annotations. All tools have descriptions (10 - 200 chars range, solid) and input schemas with type constraints. However, output schemas are undocumented, callers cannot predict the structure of returned data. Error handling is minimal (no recovery guidance). Tool compositions are sound (each does one thing), but lacking inter-tool chaining hints. Descriptions are adequately detailed but do not reference dependencies or failure modes. The 'limit' parameter in fetch_url (hardcoded to 20,000 in implementation) is not reflected in the schema, creating a discrepancy.
Universal content loader that fetches and processes content from any URL. Automatically detects content type (webpage, PDF, or image) based on URL.
Execute a web search using the given search query. Returns a list of results including title, URL, and a rich content snippet.
Capture and return a rendered screenshot of a website. Helpful when not just the data but also the visuals are relevant.
Output schemas are not documented. Callers cannot predict return structure for search_web (does it return {results: [...], count: int}? fields in each result?), fetch_url (what metadata is included?), or view_website (Image type is defined but no description of width/height/format).
fetch_url parameter 'limit' is hardcoded to 20,000 in implementation (loaders.py line ~123: limit=20_000) but not exposed in schema, and the limit parameter description does not mention this constant. This violates the parameter-constraint documentation rule and prevents agents from controlling truncation.
No error handling guidance. None of the three tools provide actionable error messages or recovery hints. For example, search_web does not document what happens if the search provider is unavailable; fetch_url does not explain what to do if a URL is inaccessible or returns a 404; view_website does not document timeout behavior or what happens if JavaScript fails to render.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 49 | - | v1 |
view_website description is vague ('Helpful when not just the data but also the visuals are relevant'). Does not explain when to use it vs fetch_url, what formats it returns, resolution, whether it supports full-page scrolling, or timeout behavior.
fetch_url has a provider selection parameter (trafilatura, httpx, zendriver) but does not explain the performance/accuracy tradeoffs or when to prefer each. An agent cannot reason about which to choose.
No inter-tool chaining hints. If search_web returns results, the agent must infer it should call fetch_url next. Descriptions should state: 'Use search_web to find URLs, then fetch_url to extract full content.'