Cloudflare Workers implementation of an MCP server for searching and extracting content from social media platforms (Reddit and TikTok) using the PostCrawl API
Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
PostCrawl MCP server defines 4 tools with reasonable structure but exhibits significant gaps in descriptions, parameter documentation, and error handling. All tools use Zod schemas (positive), but descriptions are generic and several parameter constraints lack clarity. The 'search_and_extract' and 'extract' tools violate the single-responsibility principle by combining concerns. Error handling is present but non-actionable (returns JSON string dumps rather than structured recovery guidance).
Tool composition violation: 'search_and_extract' combines search and extract into one tool, blocking agent composition. Should be split into separate 'search' and 'extract' tools called in sequence.
Missing output schema documentation for all tools. LLMs cannot plan downstream calls without knowing response structure. search() returns what fields? extract() returns extracted content in what format? This violates pattern:tool and pattern:response-shaper.
Error handling returns unstructured JSON string dumps (error.message + JSON.stringify(errorDetails)) instead of actionable guidance. Errors lack classification (retryable vs fatal), suggested recovery (e.g. 'Try search_users() first'), or constraint violations. Does not follow pattern:recovery-guide.
searchsearch_and_extractextractcheck_health
Recommendations
Split 'search_and_extract' into two separate tools: 'search_posts' (returns minimal metadata + IDs for composition) and 'extract_content' (takes URLs or post IDs, returns rich content). This enables agents to call search, inspect results, and extract only what they need.
Add explicit output schema documentation to all tools. Example for 'search': { post_id: string, title: string, author: string, platform: string, url: string, published_date: ISO8601, engagement_count: number }. This lets LLMs plan downstream calls.
Document parameter constraints inline in descriptions. For 'results': 'Number of results to return per page (1-100, default 10)'. For 'page': 'Page number (1-indexed, default 1)'. For 'response_mode': 'Format of response: raw (JSON), structured (schema), or summary (brief text)'.
Add dependency hints to tool descriptions. E.g., 'search': 'Use to discover posts. Results include post IDs and URLs, pass these to extract_content() for full content extraction.'
Reduce tool count by renaming: 'search' → 'search_posts', remove 'search_and_extract', add dedicated 'extract_content' for URLs. This makes the intent clearer and prevents the LLM from conflating two distinct operations.
Parameter constraints lack documentation. 'results' parameter accepts number but no min/max stated (baseline: 1-100). 'page' parameter accepts number but no validation explained. 'response_mode' enum values are extracted from responseModeSchema but not visible in tool schema, LLM sees no constraint. Violates pattern:constrained-input.
Duplicate tool functionality: both 'search' and 'search_and_extract' perform search operations, differing only in whether extraction is combined. This forces LLMs to reason about subtle distinctions instead of using one canonical search tool. Violates composition principle.
No pagination limit stated for 'results' parameter. Default is 10, but unbounded numbers let LLMs pass excessive values that could break APIs or exhaust context. Missing explicit max constraint (recommend: 1-100 range).
Tool descriptions are too short (29-54 chars) and lack LLM-optimization context. 'Search for posts' does not explain: Which platforms are searchable? What query syntax is supported? When should I use 'search' vs 'search_and_extract'? Violates pattern:tool-description (baseline: 34-392 chars, ideally 50-200).
No explicit idempotency guarantees documented. If an LLM retries a 'search' or 'extract' call due to a transient error, will results be deduplicated or doubled? This is critical for agent retry safety.
searchsearch_and_extractextract
Document idempotency: state that repeated calls with identical parameters return identical results (no side effects), making them safe to retry on transient failures.
Add batch variants if agents frequently call extract() in loops: 'extract_batch(urls: string[])' returning { results: [{ url, content, status }], failed_urls: [{ url, error }] }. This reduces token waste and latency.
Expand 'check_health' description: 'Checks whether the PostCrawl API is available and responsive. Returns status (operational/degraded/down), latency_ms, and component health (search: ok/slow/error, extract: ok/slow/error). Use before executing search/extract if reliability is critical.'
Add examples in docstrings (not descriptions) showing typical request/response pairs. E.g., in code comments or a separate docs site, show: search({ query: 'AI news', platforms: ['twitter'], results: 5 }) → { posts: [...], total: 1024, next_page: 2 }.