Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The Semantic Scholar FastMCP server provides 16 well-defined tools with consistent naming patterns and complete JSON Schema input definitions. All tools follow a clear verb_noun naming convention (search, get, list patterns). Input schemas are present and properly typed for all tools. However, descriptions are brief (averaging 80-120 characters), which is below the production baseline of 194 characters and limits LLM ability to disambiguate between similar tools. Output schemas are not documented, critical for agent planning. Tool composition is sound (single responsibility), but parameter descriptions lack depth regarding constraints, valid ranges, and format requirements. Error handling guidance is absent from all tool descriptions. No tool annotations (readOnlyHint) are present despite all being READ_ONLY operations.
Output schemas not documented. Tools return results but agent cannot predict field names, types, or structure. Forces trial-and-error or documentation lookup to plan downstream chains.
Tool descriptions are 50-120 characters, significantly below the 194-character baseline. Too brief to guide LLM selection when multiple similar tools exist (3 search tools, 3 batch/detail tools). Descriptions lack WHEN to use guidance and prerequisites.
Expand all tool descriptions to 150-250 characters. Include: (1) What data does it return? (2) When should I call this vs similar tools? (3) Are there prerequisites or required parameters? Example: 'paper_relevance_search: Search papers by keyword relevance across titles, abstracts, and citations. Use when exploring a topic broadly; use paper_title_search for exact title matches. Supports filters (year, venue, open access). Returns up to 1000 papers with pagination.'
Document output schemas for all tools. Create a structured response example in docstrings or API docs. Minimally: list returned fields, their types, and whether each field is always present or conditionally set. Example: 'Returns: [{ paper_id: string, title: string, year: number, citation_count: number, venue?: string, authors: [{ author_id: string, name: string }] }]'
Add parameter constraint guidance. For each parameter, specify: (1) Allowed values (enum) or format (regex/length/range), (2) Whether it is required and what happens if omitted, (3) Interaction with other parameters. Example for 'limit': 'Integer, range 1 - 100 (default 10). Returned results may be fewer if API has fewer matches.'
Add min/max bounds to all numeric parameters. Specify: min_citation_count (0 - 999999), limit (1 - 100), offset (0 - 10000). Without bounds, LLMs may pass absurd values like 'limit=999999' that timeout or violate API quotas.
Enumerate valid field names for the 'fields' parameter across all tools. Example: 'fields: Array of field names to include in response. Valid values: title, abstract, venue, year, citation_count, authors, open_access_pdf, etc. Omit or pass [] to get default fields.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Parameter descriptions lack format/constraint guidance. 'query' is just 'Search query string' without mentioning required length, special characters allowed, or example formats. 'limit' has no min/max bounds (rubric baseline: should specify 1-100 range). 'fields' accepts arrays but does not enumerate valid field names from the Semantic Scholar API spec.
Pagination token for bulk_search not explained. Parameter 'token' described as 'Pagination token for bulk search (optional)', no guidance on how to obtain it, format, or max lifespan. Agents do not know whether to reuse tokens across sessions.
Tool annotations absent. All 16 tools are READ_ONLY (idempotent, safe, no side effects). Adding readOnlyHint=true to all tools would enable agents to call them speculatively without fear of mutations. This is a straightforward improvement with no downside.
Ambiguous tool naming in recommendation cluster. 'get_paper_recommendations_single' vs 'get_paper_recommendations_multi' requires LLM to understand plural/singular distinction. Better names: 'recommend_papers_by_id' and 'recommend_papers_by_multiple'. Current names are verb_noun but lack action clarity.
Batch detail tools (paper_batch_details, author_batch_details) return 'fields' parameter as string not array, breaking consistency. paper_details and paper_authors use array; batch variants use string. LLM must reason about format difference.
paper_batch_detailsauthor_batch_details
Add tool annotations (readOnlyHint: true) to all 16 tools in the FastMCP server configuration. This is a one-line change per tool that signals to agents: this call is safe to retry, speculatively invoke, or parallelize with others.
Add error recovery hints to tool descriptions. Example: 'If paper_id is invalid, error message will include: "Paper not found. Try paper_relevance_search() to find paper IDs by title." Retryable errors (rate limit, timeout) will be retried automatically; non-retryable errors will surface to the user.'
Rename recommendation tools for clarity: 'get_paper_recommendations_single' → 'recommend_papers_by_paper_id' and 'get_paper_recommendations_multi' → 'recommend_papers_by_papers'. This makes intent clearer without introducing ambiguity.
Standardize batch detail parameter schemas. Both paper_batch_details and author_batch_details should use 'fields: array<string>' not 'fields: string', matching single-detail tool conventions.
Document the pagination token format and lifecycle for bulk_search. Example: 'token: String pagination cursor returned by a prior paper_bulk_search call. Tokens expire after 30 days. Do not reuse across different queries.'
Add rate limit guidance to tool descriptions or server metadata. Example: 'Semantic Scholar API allows 1 request/second with standard key, 100/second with partner key. Requests exceeding quota return HTTP 429. FastMCP will retry with exponential backoff.'
Document result limits and pagination. Example: 'paper_relevance_search returns max 10 results per call (limit param). For exhaustive searches, use paper_bulk_search with pagination tokens, or iterate with offset. LLM should plan multi-step queries to stay under token limits.'