MCP server that can perform a web search locally without the use of APIs.
This server has 5 tools, all with complete input schemas and reasonable descriptions. However, there are critical issues with naming consistency, parameter descriptions, and output schema documentation. Tool names use a mix of conventions (rag_search_* vs deep_research_*) and some descriptions are generic or incomplete. Most critically, output schemas are not formally documented, tools return Dict with vague 'content' fields. Parameter descriptions exist but lack constraints (min/max for num_results, top_k; valid backend enum values not enforced). The code shows lazy imports and embedding overhead but no error handling guidance for LLMs. No tool annotations (readOnly hints) despite all tools being read-only. Average tool score across the 5 is ~52, placing this in the 'fair' range with noticeable gaps.
Perform deep research across multiple search terms using specified search backends. This tool aggregates results from multiple searches across chosen engines, scores them by relevance, and returns the most relevant content with duplicates removed. Perfect for comprehensive research on a topic. Available backends: bing, brave, duckduckgo, google, grokipedia, mojeek, yandex, yahoo, wikipedia
Perform deep research using DuckDuckGo as the search backend across multiple search terms. Aggregates results, scores them by relevance, and returns the most relevant content.
Perform deep research using Google as the search backend across multiple search terms. Aggregates results, scores them by relevance, and returns the most relevant content.
Search the web for a given query using DuckDuckGo. Returns context to the LLM with RAG-like similarity scoring to prioritize the most relevant results. This tool fetches web search results, scores them by semantic similarity to the query using text embeddings, and returns the top-ranked content as markdown text.
Search on Google for a given query using ddgs. Give back context to the LLM with a RAG-like similarity sort.
Inconsistent naming convention. Tools use both 'rag_search_*' (single-backend) and 'deep_research_*' (multi-backend) patterns. LLMs conflate similar names, distinguishing 'rag_search_ddgs' from 'deep_research_ddgs' requires reading full descriptions. Should use verb_noun consistently (e.g., 'search_web_ddgs', 'search_web_multi_backend').
Output schemas not documented. All tools return Dict with a single 'content' key containing markdown text. The tool descriptions do not formally specify the return type structure (e.g., {"type": "object", "properties": {"content": {"type": "string"}}}). LLMs cannot infer downstream field usage or error handling from untyped returns.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 60 | - | v1 |
Parameter constraints not enforced or described. 'num_results' and 'top_k' accept integers but no min/max bounds are documented (code defaults to 10 and 5, but description doesn't state valid range or why these defaults). 'backends' in 'deep_research' is an array of strings but does not formally constrain to the listed enum [bing, brave, duckduckgo, google, grokipedia, mojeek, yandex, yahoo, wikipedia]. Free-form string arrays invite hallucinated backend names.
rag_search_google description is terse (60 chars: 'Search on Google for a given query using ddgs. Give back context to the LLM with a RAG-like similarity sort.'). Does not explain the difference from rag_search_ddgs, when to use it, or what 'similarity sort' means. Fails pattern:tool-description baseline of 34 - 392 chars; this is at the lower end and lacks actionable detail.
No error handling guidance. Code does not document failure modes (network timeout, invalid search backends, failed content fetch, embedding errors). Docstrings do not tell LLMs: 'If DuckDuckGo is blocked, try Google backend.' or 'If embedding fails, retry with fewer results.' Error recovery is invisible to the agent.
No tool annotations. All tools are read-only (no side effects), but fastmcp tool definitions lack readOnlyHint or similar metadata. MCP protocol now supports tool annotations to signal immutability. Current implementation does not declare this, forcing LLMs to infer read-only status from descriptions.
deep_research 'backends' parameter has null default (defaults to ["duckduckgo", "google"]) but description says 'If None, uses default', inconsistent and confusing. LLM will not know whether passing null vs omitting the param has different behavior.