An MCP server for searching and retrieving articles from Google Scholar
This server has significant definition quality issues. While all three tools are explicitly registered with the FastMCP decorator and have basic descriptions, the schemas lack proper parameter typing, descriptions are too generic, and critical output structure is undocumented. The server follows a basic verb-noun naming convention but lacks the depth needed for production LLM integration. Parameter descriptions are minimal or absent, parameter types in some cases are improperly specified (e.g., 'tuple' for year_range), and return types are not formally documented. Error handling is present but generic, returned error objects don't guide the agent toward recovery. Tool composition is reasonable (three distinct search operations), but lack of pagination support for search results is a critical gap.
Get detailed information about an author from Google Scholar.
Search for articles on Google Scholar using advanced filters.
Search for articles on Google Scholar using key words.
No documented output schemas for any tool. Tools return dicts with inferred structure (title, authors, abstract, url for searches; name, affiliation, interests, citedby, publications for author info). LLMs cannot plan downstream calls or extract fields without explicit output documentation.
Parameter 'year_range' in search_google_scholar_advanced is typed as 'tuple', not a valid JSON Schema primitive. Should be array ([start_year, end_year]) or object with 'start_year' and 'end_year' properties. Current type confuses LLMs and breaks schema validation.
No pagination support. search_google_scholar_* tools accept num_results but return all matching items without offset/limit or cursor mechanism. Large result sets can exhaust context windows. Missing pattern: paginated-result.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 36 | - | v1 |
Parameter descriptions are too brief (2-5 words). LLM-optimized descriptions should explain WHAT the param does, WHY it matters, and allowed values/formats. E.g., 'author' in advanced search should say 'Filter results by author name (partial match supported, e.g. "Einstein"). Returned results include only papers by matching authors.' Current: 'Author name'.
Tool descriptions lack discrimination context. 'search_google_scholar_advanced' description does not explain when to use it instead of 'search_google_scholar_key_words'. LLMs cannot distinguish tools without comparative guidance. Should add: 'Use this when filtering by author or publication year. For simple keyword queries, use search_google_scholar_key_words.'
Error responses are generic dicts: {'error': 'An error occurred...'} without recovery guidance. Errors do not categorize as retryable, user-fixable, or fatal. Missing pattern: recovery-guide. Example: 'Author not found. Did you mean: [list similar names]?' or 'Search failed (rate-limited). Retry in 10 seconds.'
Tool descriptions lack execution side effects declaration. 'search_google_scholar_advanced' should state 'Read-only. Does not modify Google Scholar data.' Agents need to know which tools are safe to retry (get/search/read) vs. risky (create/update/delete/send).
No input validation or constraints. 'query' parameters accept any string; no length limits documented. 'num_results' has default 5 but no min/max bounds (e.g., 1 - 100). LLMs can pass absurd values (num_results=999999) without feedback.
get_author_info result limit (5 publications) is hardcoded in implementation and not documented in description. Users expect 'detailed information' but get truncated data without explanation.