A Rust-based workspace containing multiple MCP servers for web crawling, file system operations, Reddit API access, Tavily search integration, and Twitter API access
This workspace spans 5 distinct MCP servers (mcp-crawl, mcp-filesystem, mcp-reddit, mcp-twitter, mcp-tavily) with 31 total tools. All tools have basic descriptions and well-formed JSON Schema inputs with proper type definitions and required fields. However, descriptions are generic and lack LLM-optimization guidance. Parameter descriptions are minimal, most say only what the parameter is named (e.g., 'The URL to scrape') without explaining when to use the tool vs alternatives, prerequisites, or error recovery paths. Tool naming follows verb_noun convention consistently, but compositions have issues: several tools operate on the same resource (e.g., multiple scraping variants) without clear functional boundaries. Output schemas are not documented in the provided source, we see input schemas but no explicit response type definitions or pagination guidance. Error handling is absent from visible code. Security is a concern: no visible secret injection, no permission gates, and no audit logging. Overall, this is a competent but baseline-quality workspace, solid foundations but missing the LLM-centric polish, error recovery guidance, and composition patterns that distinguish A-tier servers.
Advanced web scraping with support for JavaScript rendering and dynamic content
Create a new directory, including any necessary parent directories. If the directory already exists, this operation will succeed without error.
Extract attribute values from elements matching a CSS selector
Extract all forms and their fields from a webpage
Extract all images from a webpage
Extract all links from a webpage
Extract metadata (title, description, keywords, etc.) from a webpage
Output schemas completely absent. No tools document their response structure, field types, pagination, or what IDs/references are included for downstream chaining. LLMs cannot plan multi-step sequences without knowing what fields to expect.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Extract structured data (JSON-LD, microdata) from a webpage
Extract all tables from a webpage with headers and rows
Extract text content from elements matching a CSS selector
Get comments from a specific Reddit post
Get direct message conversations
Get posts from a subreddit sorted by specified criteria
Get Twitter user profile information
Get information about a specific subreddit
Get user's home timeline
Get trending/popular subreddits
Get current Twitter trending topics
Get comments posted by a specific Reddit user
Get information about a Reddit user
Get posts submitted by a specific Reddit user
Get a detailed listing of all files and directories in a specified path. Results clearly distinguish between files and directories with [FILE] and [DIR] prefixes. This tool is essential for understanding directory structure and finding specific files within a directory.
Read the complete contents of a file from the file system. Handles various text encodings and provides detailed error messages if the file cannot be read. Use this tool when you need to examine the contents of a single file.
Scrape a single webpage and extract basic content using readability
Search for text patterns in a webpage using regex
Search for posts across Reddit or within a specific subreddit
Search for tweets
Select elements from HTML using CSS selectors
Post a new tweet
Write content to a file, creating the file if it doesn't exist and creating parent directories as needed. This will overwrite existing files.
Convert XPath expressions to equivalent CSS selectors
Descriptions lack LLM-optimization. Most are 30-60 characters and state only the minimal 'what' (e.g., 'Extract text content from elements matching a CSS selector'). Missing 'when to use' guidance, prerequisites, and error recovery paths. LLMs cannot distinguish between similar tools (e.g., extract_text vs select_elements) without clearer context.
Web scraping tool composition is unclear. Eleven extraction tools (select_elements, extract_text, extract_attributes, extract_links, extract_images, extract_forms, extract_tables, extract_metadata, search_patterns, extract_structured_data, advanced_scrape) all operate on URLs without clear functional boundaries. LLMs will struggle to choose between them. Recommend consolidating into 2-3 well-defined tools (basic_scrape, advanced_scrape_with_js, extract_by_selector).
No error handling or recovery guidance visible in code. Tools like extract_attributes, search_patterns, xpath_to_css have no documented failure modes. LLMs cannot self-correct when invalid input is passed (e.g., bad CSS selector, regex pattern).
Pagination not documented. get_posts, search_posts, get_comments accept a 'limit' parameter but no offset, page, or cursor token. Large result sets will bloat the context window without guidance on how to fetch the next batch.
Parameter descriptions are missing or minimal. 'limit' in get_posts has no guidance on valid range (e.g., max 100?). 'selector' in select_elements has no hint about valid CSS syntax. LLMs will guess or pass invalid values.
No security constraints visible. write_file and create_directory have no path sanitization or permission checks documented. No secret injection pattern used. API credentials (for Reddit, Twitter) must be stored server-side, but no visible mechanism shown.
No destructive operation confirmation or dry-run support. write_file overwrites without warning; send_tweet posts immediately. Agents can cause unintended data loss or public posts without a confirmation step.
Enum constraints incomplete. get_posts sort parameter correctly uses enum ['hot', 'new', 'top', 'rising'], but search_tweets mode uses enum incorrectly without documenting fallback behavior if an invalid mode is sent. Some tools accept free-form strings where enums would prevent hallucination.
Idempotency not declared. Repeated calls to send_tweet with the same text will post duplicate tweets. No idempotent key or deduplication mechanism visible. Agents retrying on ambiguous failures will create duplicates.