A comprehensive MCP server for Medium article scraping with GraphQL-first architecture, ElevenLabs Creator API for premium audio, parallel processing and HTML output. Supports 35+ domains including TowardsDataScience.
The Medium Scraper MCP server has significant definition quality issues. While tool names follow verb_noun conventions (medium_scrape, medium_search, medium_batch), descriptions are inconsistent in depth and quality. Most critically, input parameter schemas are partially visible but lack complete JSON Schema type definitions in the provided code. The sample code excerpt cuts off mid-response for resources, making full assessment difficult. Of the 7 tools visible: medium_scrape, medium_batch, and medium_search have reasonably detailed descriptions (100-150 chars) with multiple parameters described; medium_fresh and medium_export are more sparsely documented; medium_cast and medium_synthesize mention external APIs but lack error handling guidance for API failures. Output schemas are largely undocumented in the visible code, we see a resource hint ('medium://trending') but no structured output documentation for tool responses. Error handling is minimal, no guidance on what happens when URLs fail, APIs timeout, or paywalls are hit (though handle_paywall() function exists, it's not integrated into tool descriptions).
Scrape multiple Medium articles in parallel.
Generate audio podcast from a Medium article using ElevenLabs TTS.
Export a Medium article to PDF or other formats.
Get fresh/trending articles for a specific tag or publication.
Scrape a Medium article with full v3.0 capabilities.
Search Medium for articles by topic or keyword.
Synthesize a research report using AI (Gemini/OpenAI) on a given topic from Medium articles.
Output schemas are undocumented across all tools. No tool specifies what fields are returned, data types, or structure. LLMs cannot plan downstream calls or extract required data without documented responses.
Enum parameters are mentioned in descriptions but not formalized in JSON Schema. 'medium_search.sort' should be enum["latest"|"top"]; 'medium_synthesize.model' should be enum["gemini"|"openai"]; 'medium_export.format' should be enum["pdf"|"markdown"|"html"]. Free-form strings invite hallucinated invalid values.
Error handling and recovery guidance is missing. No documentation of what happens on paywall detection (handle_paywall() exists in code but not surfaced in tool contract), API failures (ElevenLabs, Gemini, OpenAI timeout), invalid URLs, or rate-limiting. Errors should guide LLM on recovery actions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 71 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter constraints are mentioned in descriptions (e.g., 'max_results (default: 10, max: 50)') but not formalized in JSON Schema with min/max properties. LLMs cannot reliably validate without schema enforcement.
API credentials (ElevenLabs API key, Gemini/OpenAI keys) are referenced in code but tool descriptions do not clarify how authentication is handled server-side. Users must infer that secrets are injected; this should be explicit to avoid LLMs attempting to pass tokens as parameters.
Tool descriptions lack context on when to use each tool vs. alternatives. 'medium_search' vs. 'medium_fresh' distinction is unclear. No guidance like 'Use medium_search for broad queries; use medium_fresh for trending articles in a specific topic.'