Self-hosted URL-to-Markdown service with stable, refreshable share links. Converts web pages, Reddit posts, Hacker News threads, and documents to Markdown via MCP.
PullMD provides 6 tools for web-to-Markdown extraction with clear, action-verb names and reasonable descriptions. Input schemas are present with typed parameters and descriptions for most parameters. However, several critical gaps reduce overall quality: (1) output schemas are not documented, tools return Markdown/metadata but the structure is not formally specified, preventing LLMs from planning downstream use; (2) error handling is missing, no guidance on retryability, user-fixable errors, or recovery paths; (3) parameter constraints are weak, many free-form strings lack enums or format validation (e.g., 'lang' in extract_reddit accepts a string with no constraint on valid language codes); (4) composition issues, tools like extract_web and extract_html suggest overlapping responsibilities but are not clearly distinguished; (5) security audit trail and permission gating are not evident in the tool definitions. The tools follow naming conventions well (verb_object pattern), but the lack of output documentation and error guidance prevents confident agent use.
Generate YAML frontmatter for extracted content
Determine if a URL is a Reddit or Hacker News link
Extract and convert a Hacker News thread to Markdown
Extract and convert a Reddit post to Markdown
Extract and convert a web page to Markdown
Score the quality/relevance of extracted content
Output schemas not documented. Tools return Markdown content, metadata, quality scores, and frontmatter YAML, but the response structure (field names, types, nested objects) is not formally specified. This prevents LLMs from knowing what fields to extract and planning downstream tool chains.
No error handling guidance. Tool descriptions do not indicate which errors are retryable (network timeouts), user-fixable (invalid URL), or fatal (SSRF block). Agents cannot distinguish and will retry unrecoverable errors or give up on retryable ones.
Free-form string parameters without constraints. 'lang' in extract_reddit accepts any string with no enum or validation hint. 'url' in all tools has no format constraint (must be valid HTTP/HTTPS). LLMs may pass invalid values like 'hello' or 'ftp://...' without self-correction hints.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Unclear tool composition. Three extraction tools (extract_web, extract_reddit, extract_hackernews) plus a utility (check_url_type) suggest overlapping responsibilities. The descriptions do not clearly distinguish when to use extract_web vs. extract_reddit, only the parameter names hint at the difference. This forces LLMs to reason about tool selection.
Parameter descriptions lack detail about expected format and constraints. For example, 'URL of the web page to extract' does not specify whether relative URLs are accepted, what URL schemes are valid, or how to handle redirects. Format specifications must be in the description because JSON Schema patterns are not visible to LLMs.
No pagination or result limits documented. extract_web and extract_reddit may return very large Markdown documents (entire pages with all comments), but the tool descriptions do not mention size limits or indicate if truncation occurs. Large results can exhaust context windows without warning.
Security permissions not declared. The tool definitions do not indicate whether rate limits apply, what SSRF checks are enforced, or what credentials/tokens are required. Agents cannot make informed decisions about which requests to retry or whether to batch calls.