A Go-based MCP server providing web search, academic search, web page fetching, and PDF parsing capabilities with LLM summarization support.
Server has 4 well-defined tools with generally strong descriptions and input schemas. Naming is clear (smartsearch, academicsearch, cleanfetch, pdf_parser). Descriptions are detailed and LLM-optimized (200+ chars, exceeding typical 194 char baseline). All parameters have type definitions and descriptions. However, output schemas are not explicitly documented in the visible code, parameter constraints are insufficiently formalized (enum values not declared for key parameters like 'time_range'), and error handling guidance is minimal. No tool annotations (readOnlyHint, destructiveHint) observed despite all tools being READ_ONLY operations.
学术论文检索工具,从多个学术数据库并行搜索论文,返回标准化的 Markdown 格式结果(含标题、作者、DOI、期刊、引用数、PDF 链接)。可用引擎(engines 参数可多选,为空则全部使用)包括 arxiv、crossref、openalex、semantic_scholar、pubmed、google_scholar、europepmc、dblp、doaj。引擎选择建议:医学/生物 → pubmed, europepmc | CS/AI → arxiv, semantic_scholar, dblp | 全学科 → crossref, openalex, google_scholar | 开放获取 → doaj。已持有 DOI 或 arXiv id 时,直接将其作为 query(如 '10.1038/s41586-020-2649-2'、'doi:10.1038/s41586-020-2649-2'、'https://doi.org/10.1038/s41586-020-2649-2' 或 '2401.04085'、'arXiv:2401.04085'、'https://arxiv.org/abs/2401.04085'),将走单篇精确查询(忽略 engines/time_range/page);拿到结果中的 pdf_url 后,可将该 URL 作为 pdf_parser 工具的 path 参数解析全文。
网页内容抓取工具,获取指定 URL 的干净 Markdown 内容。可用 urls 批量抓取(与 url 合并去重,最多 5 个),单条失败不影响其它。
PDF 解析工具,path 为本地文件路径、file:// 或远程 http(s) PDF URL(学术结果的 pdf_url 可直接传入)。优先用 PDF 库提取文本转为 Markdown;长文档可用 pages 指定页码(如 '1-10'),省略时默认只解析前 20 页(pdf_parser.max_pages)并提示截断;大文档自动存储到临时文件。
通用联网检索工具,获取实时信息(新闻/技术文档/产品/数据等)。查询词需精准凝练、聚焦核心意图,避免堆砌大量同义/次要词形成关键词列表(会稀释相关性)。可用 intent 参数说明检索目的以获得更精准的结构化摘要。可用 fetch_top_n 对评分最高的前 N 条抓取正文(默认 0 不抓,上限 5)。主引擎不可用时自动回退 Bing。
Output schemas not documented. All four tools lack explicit return-type documentation. LLMs cannot plan downstream calls or extract required fields without knowing what structure to expect. Per pattern:tool, output schemas are mandatory for agent composition.
Enum constraints not formalized in schema. 'time_range' parameter accepts integers but should validate against allowed ranges (1,3,6,12,0). 'timeRange' in academicsearch accepts strings ('year', 'month', 'week', 'day') but these are described in text, not declared as enum in JSON Schema. LLMs cannot auto-complete valid values and may hallucinate invalid ones.
Tool annotations missing. All four tools are READ_ONLY operations (no side effects) but lack readOnlyHint annotation. Per current MCP spec (2026-07-28), tool annotations guide agent execution planning and safety checks. Their absence means agents cannot auto-classify tools as safe to execute speculatively.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 68 | 2026-07-28+ | v2 |
Error handling lacks recovery guidance. Source code shows no error messages that explain 'what to do next'. E.g., if academicsearch fails because an engine is unreachable, the response should suggest fallback engines or simpler queries. Per pattern:recovery-guide, errors must guide the agent toward resolution.
Parameter interdependencies underdocumented. cleanfetch accepts both 'url' and 'urls' (merged and deduplicated). The constraint 'urls' must be combined with 'url' into max 5 items is stated in the description but should also be reflected in parameter relationship docs or schema constraints. This prevents validation ambiguity.