AI-powered academic paper analysis assistant with chat, terminology management, and deep analysis capabilities
PaperMate has 4 tools with basic functionality but significant quality gaps. All tools have descriptions and visible input schemas, but descriptions lack depth and fail to follow agentic tool patterns. Parameters are minimally constrained (only 'query' string across 3 search tools). The add_term tool has the most complete schema with three typed parameters and a character constraint, but no error guidance or recovery patterns. None of the tools have documented output schemas, making downstream composition difficult. No tool annotations (readOnlyHint/destructiveHint) despite clear risk classifications. Error handling is minimal, tools return raw exceptions rather than actionable recovery guidance. Tool descriptions are generic and don't answer WHEN to use each tool over alternatives (e.g., why arxiv_search vs tavily_search vs searxng_search?). Parameter descriptions are trivial ('Search query or term to look up' repeated 3 times).
Add a professional term to the project's terminology memory. Use this tool when you explain a new technical term to the user. The term will be saved for consistent translation across the project. Input format: JSON with fields: term (English), translation (Chinese), explanation (brief definition) Example: {"term": "Transformer", "translation": "变换器", "explanation": "基于自注意力机制的神经网络架构"}
Search arXiv for related academic papers. Use this to find authoritative definitions, translations, and usage of technical terms in research papers.
Search a SearXNG instance for scientific sources. Use this to find definitions and usage of academic terms.
Search the web via Tavily for authoritative academic definitions and translations. Prioritizes arxiv.org results.
Three near-identical search tools with no disambiguation. Descriptions do not explain WHEN to use arxiv_search vs tavily_search vs searxng_search. LLMs will struggle to choose the right tool.
No documented output schemas for any tool. arxiv_search returns formatted text blocks with Title/Authors/Summary/URL, but this structure is not formally declared. LLMs cannot reliably parse results or plan downstream operations.
add_term has a '50 chars max' constraint on explanation parameter but this is only mentioned in tool description, not in parameter schema. JSON Schema 'maxLength' property is missing.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No tool annotations despite clear risk classifications. add_term is WRITE and should have destructiveHint=true. Search tools should have readOnlyHint=true. Without annotations, agents cannot reason about side effects.
Generic parameter descriptions. 'Search query or term to look up' repeats identically across 3 tools. Descriptions should explain what kind of query works best (e.g., 'technical terms', 'author names', 'full paper titles').
Error handling returns raw exceptions (ToolExecutionError) without recovery guidance. 'arXiv search error: ...' or 'Empty query.' tell LLM nothing about what to do next.
No parameter constraints (enums, min/max, patterns). Query parameters accept any string, could trigger billion-character API calls or SQL injection if passed unsanitized to backend services.