MCP server for web search, website content extraction, and YouTube transcript retrieval
This server has three clearly named, verb-led tools with reasonably comprehensive docstrings and good parameter documentation. All tools are READ_ONLY, which is appropriate for a web search and content extraction service. Naming is clear and action-oriented (search_web, read_website, get_youtube_transcript). Parameter descriptions are present and mostly well-specified with format guidance (e.g., 'must start with http:// or https://', '11-char video ID'). However, there are notable gaps in schema documentation, output field specification, and error handling guidance. Input schemas are well-formed with types and defaults where appropriate, but return value schemas are not formally documented in code, they are only described in prose. Error handling returns error strings instead of structured responses, and recovery guidance is minimal. The server lacks tool annotations (readOnlyHint, etc.) and does not document pagination or result limits formally in the schema. Overall, this is solid mid-tier work with good naming and description clarity but incomplete schema and output documentation.
Get the text transcript of a YouTube video. Use this tool to retrieve spoken content from a YouTube video. Accepts a video ID (e.g. "dQw4w9WgXcQ") or any common YouTube URL format including youtube.com/watch, youtu.be, /shorts/, /embed/, and /live/ links. Returns the full transcript as a single string, or an error message if no transcript is available.
Extract the main text content from a webpage URL. Use this tool when you need to read the full content of a specific webpage. Returns the extracted text content, truncated to ~15 000 characters if the page is very long. Returns an error string if the fetch or extraction fails.
Search the web using DuckDuckGo. Use this tool to find current information on any topic. Returns a JSON array of result objects, each containing: - "title": the page title - "href": the URL of the result - "body": a short snippet/summary of the page (truncated to ~200 chars)
Output schemas not formally documented. Return types are described in prose (e.g., 'JSON array of result objects', 'extracted text content', 'full transcript as a single string') but not as JSON Schema in the tool registration. LLMs cannot parse expected fields from docstrings alone, they need formal schema to plan downstream chaining and extract the right data.
Error responses are unstructured plain strings instead of structured recovery guidance. Errors like 'Error: Could not fetch content from {url}' or 'Error performing search: {str(e)}' give the LLM no actionable next steps. Per pattern:recovery-guide, errors should tell the agent what to do next (retry, call a different tool, ask the user, etc.) and categorize as retryable vs. user-fixable vs. fatal.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 50 | - | v1 |
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). All three tools are READ_ONLY, which should be declared via readOnlyHint in the tool definition to enable the LLM to reason about side effects and retry safety. Missing annotations reduce protocol alignment and LLM reasoning quality.
search_web lacks explicit pagination documentation and enforcement. While max_results is capped at 10, there is no offset/page parameter or next_cursor for iterating large result sets. If an agent needs more than 10 results, the tool cannot deliver them. Per pattern:paginated-result, tools returning lists should support pagination.
Missing input validation hints in descriptions. While read_website documents 'must start with http(s)://', there is no guidance on what happens if invalid URLs are passed. get_youtube_transcript accepts multiple URL formats, but no validation rules are stated. Per pattern:constrained-input, parameter descriptions should specify expected format, range, and constraints explicitly.