AI-powered research workflow assistant with MCP servers for academic database access, ICMJE-compliant
The server defines 24 tools across 5 MCP servers with mostly complete schemas and descriptions. Tool naming is action-verb consistent (bib_*, search_*, export_*, get_*), and parameter descriptions are present for most tools. However, several critical gaps emerge: (1) No output schemas are documented anywhere, the rubric requires them so LLMs know what to expect; (2) Some parameters lack type constraints (enums, ranges), e.g., 'source' in bib_search and search_europepmc should be enums, 'detail_level' in export_session is free-text when it should be an enum; (3) Error handling is absent, no recovery guidance, no classification of errors as retryable/user-fixable; (4) Parameter relationships and dependencies are undocumented (e.g., bib_import_bibtex requires one of file_path OR bibtex_text, but this is not stated); (5) A few descriptions are generic or too brief. The bibliography manager and crossref servers have reasonable descriptions and parameters, but missing output schema documentation and error handling drop the overall score into the C+ range.
Add a single reference to the project bibliography.
Add a note to a reference.
List all files attached to a reference.
Get notes for a reference, optionally filtered by type.
Import references from a BibTeX file or text.
Import references from an RIS file or text.
Link a PDF or other file to a reference.
NO OUTPUT SCHEMAS DOCUMENTED. The rubric requires documenting return types and fields so LLMs know what data to expect and can plan chained calls. For example, search_works should document that it returns a list of work objects with fields like title, authors, doi, publication_year, abstract. Without this, LLMs cannot reliably extract data or chain tools.
MISSING ENUM CONSTRAINTS on free-text parameters. Parameters like 'detail_level' (export_session, export_latest), 'source' (bib_search), 'result_type' (search_europepmc), 'sort' (search_works, search_europepmc), 'note_type' (bib_add_note, bib_update_note) should declare valid values as enums, not just describe them in text. This prevents LLM hallucination of invalid values.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | <=2025-11-25 | v2 |
List all searches and imports recorded for a project.
Search stored references in the project bibliography.
Update an existing note.
Validate whether a DOI exists and retrieve basic metadata.
Draft a reproducible CrossRef search script for user review.
Export the most recent chat session to a QMD file.
Export a VS Code Copilot Chat session to a QMD file.
Get articles that cite a given article (forward citations).
Retrieve the full text of an open access article from PMC.
Get articles referenced by a given article (backward references).
Get the reference list for a work by its DOI.
Get full metadata for a work by its DOI.
List available VS Code Copilot Chat sessions for this workspace.
Search Europe PMC for biomedical and life sciences literature.
Search Europe PMC and save a reproducible script to the project.
Search CrossRef for bibliographic works.
Search CrossRef and save a reproducible script to the project.
MISSING ERROR HANDLING GUIDANCE. No tool descriptions include recovery hints or error classification. E.g., search_works should say: 'If no results found, try broadening the query or removing filters. API rate limit: 1 request per second.' get_full_text should indicate which PMCID prefixes are valid (PMC only, not PMID). Agents need actionable guidance on how to recover from failures.
UNDOCUMENTED PARAMETER DEPENDENCIES. bib_import_bibtex and bib_import_ris accept EITHER file_path OR raw text content, but the description does not state this mutual exclusivity or which is preferred. Similarly, export_session accepts both project_path and output_path with unclear precedence. Agents may pass both and get unexpected behavior.
MISSING PAGINATION INFO. search_works, search_europepmc, and similar tools accept 'rows' or 'page_size' parameters but do not document: (1) whether results are paginated or capped; (2) what the total count is (for agent planning); (3) whether a next_cursor or offset is returned for continuation. Without this, agents cannot reliably fetch large result sets.
GENERIC/BRIEF DESCRIPTIONS. Tools like 'list_sessions' ('List available VS Code Copilot Chat sessions for this workspace') and 'bib_list_searches' lack depth on when to call them, what context they reveal, or downstream usage. Per the rubric, descriptions should answer: What does it do? When to call it vs. a similar tool? What does it return?
MISSING CONSTRAINTS ON NUMERIC PARAMETERS. 'rows' in search_works defaults to 20 with max 1000, but no min is documented. 'page_size' in search_europepmc defaults to 25 with max 1000, but min is missing. Agents may pass 0, negative, or enormous values. Document ranges explicitly (e.g., 'rows: integer, 1 - 1000, default 20').
INCONSISTENT NAMING ACROSS SIMILAR TOOLS. Bibliography tools use 'project_path' while chat/search tools vary (some 'project_path', some implicit). This is minor but wastes LLM reasoning on parameter mapping. Consider standardizing across the suite.