An MCP server that extracts and envelops fresh content from GitHub repositories, Google Scholar, and Hacker News using Cloudflare's browser rendering API
FreshContext Worker defines 3 tools with basic schemas and descriptions, but lacks depth in LLM-optimized documentation. All tools are read-only web extractors following a similar pattern. Tool names are reasonably clear (extract_github, extract_scholar, extract_hackernews), but descriptions are brevity-optimized for brevity rather than guidance. Parameters (url, max_length) are typed with Zod schemas and have default values, but lack detail about expected formats or constraints. No output schema documentation is visible. Error handling exists at the HTTP layer but lacks LLM-actionable guidance for recovery.
Extracts README and content from GitHub repositories using Cloudflare browser rendering
Extracts content from Hacker News using Cloudflare browser rendering with timestamp extraction
Extracts academic content from Google Scholar using Cloudflare browser rendering with date extraction
Tool descriptions lack actionable context for LLM selection. All three descriptions are identical in structure ('Extracts content from X using browser rendering and markdown conversion'). No guidance on WHEN to use each tool, what prerequisites exist, or what distinguishes them from general web scrapers. Descriptions are 88-90 characters, below the 194-char production baseline, omitting context about FreshContext envelope format, confidence metadata, or retry behavior.
No output schema documentation visible. Code shows tools return { content: [{ type: 'text', text: stamp(...) }] }, but this structure is not documented for the LLM. Agents cannot predict response fields (source, published_date, confidence, adapter) or plan downstream tool chains. Production tools must document return types explicitly.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 10 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
max_length parameter lacks actionable constraints. Description says 'Maximum length of extracted content' but does not specify units (characters vs tokens), whether it applies to the markdown result or the raw HTML, or what happens if content exceeds the limit (truncation, error, pagination). Defaults differ (6000 for GitHub/Scholar, 4000 for HackerNews) without explanation. LLMs cannot reason about the tradeoff between completeness and token usage.
No error handling guidance for LLM recovery. Code throws 'Browser Rendering API error: {status} {text}' but LLMs receive raw HTTP status codes and Cloudflare error messages with no actionable next steps. No distinction between retryable failures (rate limits, timeouts) and user-fixable errors (invalid URL, blocked content). Agents cannot self-correct.
url parameter description is minimal ('URL of the GitHub/Scholar/Hacker News resource to extract'). Does not specify expected URL format, whether relative URLs are allowed, whether authentication is required, or what content types are supported. LLMs may pass invalid or malformed URLs without clear constraints.
No pagination or result limiting strategy documented. If Cloudflare Browser Rendering returns very large markdown (e.g., full GitHub repo wiki), truncation at max_length may break semantic meaning. No next_cursor or page parameter for partial results. Tool description does not mention this limitation.