Enhanced MCP server for advanced biomedical research with AI-powered workflows, providing comprehensive PubMed search, metadata retrieval, PDF downloading, and paper analysis capabilities
This server demonstrates basic MCP structure but has significant gaps in schema quality, parameter documentation, and error handling. Tool names follow verb-noun convention adequately, but parameter descriptions are sparse or missing, input schemas lack proper type definitions in several cases, and output schemas are not documented. The server implements 7 tools with READ_ONLY and WRITE risk classifications, but lacks actionable error messages and recovery guidance. Average tool score: 52/100.
Perform a comprehensive analysis of a PubMed article.
Attempt to download the full text PDF for a PubMed article.
Enhanced search with advanced filters and AI-assisted result analysis.
Find articles related to a given PMID using NCBI's sophisticated relationship algorithms.
Fetch metadata for a PubMed article using its PMID.
Perform an advanced search for articles on PubMed.
Search for articles on PubMed using key words.
Ambiguous parameter types: `pmid` accepts both string and integer (Union[str, int]) without type discrimination in schema. LLMs may pass wrong type, causing silent failures or type coercion bugs.
Input parameter schemas lack descriptions or are incomplete. Parameters like `query_terms`, `search_fields`, `publication_types` in enhanced_search_with_filters have minimal context about valid values, format constraints, or dependencies. Rubric baseline: 100% of A+ tools have descriptions for all parameters; this server has ~40%.
Output schemas are not documented. Tools return List[Dict[str, Any]] or Dict[str, Any] with no specification of fields, types, or structure. LLMs cannot plan downstream calls or extract the right data. Rubric baseline: 100% of A+ tools have documented return types.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Error handling returns generic or unhelpful messages. Examples: 'An error occurred while searching: {str(e)}' and 'No metadata found for PMID' do not guide the LLM on what to do next (retry, try a different tool, ask the user). Rubric baseline: error responses must tell the LLM what to do next.
deep_paper_analysis registered as @mcp.prompt() but listed as a tool in the manifest. Function signature accepts pmid and returns Dict, suggesting it is a tool, not a prompt template. This creates ambiguity about its role and how it should be invoked.
No pagination support documented. Tools like search_pubmed_key_words and search_pubmed_advanced default to num_results=10 but do not show how to fetch subsequent results. Rubric baseline: tools returning lists must accept page/offset and limit, and return total count or next_cursor.
Parameter enum values not declared. E.g., sort_by accepts 'relevance', 'pub_date', 'most_recent' but these are documented only in description text, not as formal enum constraints. LLMs will hallucinate invalid values like 'date_asc' or 'newest'.
Date format documentation is inconsistent or missing. search_pubmed_advanced documents 'YYYY/MM/DD' but other tools may accept different formats. LLMs cannot reliably format dates without explicit constraints.
Destructive operation (download_pubmed_pdf) marked as WRITE risk but no dry-run, confirmation, or rollback mechanism. If PDF download fails mid-operation, no recovery guidance. Rubric baseline: irreversible operations should support confirmation step.
No tool composition documented. It is unclear whether agents should call search_pubmed_key_words then get_pubmed_article_metadata, or whether search_pubmed_advanced is the primary entry point. Tool ordering and dependency chains are not explicit.