A collection of MCP servers providing research, document, data, and web tools via FastMCP and FastAPI REST APIs
This server presents 32 tools with significant quality gaps. While tool names generally follow verb_noun conventions (search_papers, extract_info, validate_email), descriptions are inconsistent, some are adequate (40-80 chars) but many are generic or missing critical context. Parameter descriptions exist but lack detail on constraints, formats, or dependencies. Input schemas are present and properly typed in visible code (SearchPapersRequest, ExtractTablesRequest), but output schemas are undocumented or implicit. Error handling is minimal, most endpoints return generic HTTPException without recovery guidance. Security is a concern: no evidence of secret injection patterns, and tools like fetch_webpage and scrape_webpage have no timeout validation or sanitization guardrails. Tool composition suffers from redundancy (search_papers appears twice with nearly identical signatures in tools 1 and 3) and lack of chaining guidance. Pagination is absent despite web and research tools that could return large result sets. Average per-tool score: ~38, placing this in the D range (poor to fair).
Calculate basic statistics for a list of numbers.
Check the status of multiple URLs.
Check if a URL is accessible and get its HTTP status.
Clean and normalize string data with various options.
Clean and normalize text by removing extra whitespace and optionally special characters.
Count the number of pages in a PDF file.
Count words, characters, sentences, and paragraphs in text.
Tool duplication: search_papers appears in both tools 1 and 3 with identical functionality. LLM cannot disambiguate; eliminates one or consolidate.
Output schemas undocumented: extract_paper_info, search_papers_by_author, and all utility tools lack documented return types. LLMs cannot plan downstream tool calls or extract relevant fields.
No pagination or result limits: search_papers (max_results capped at 20), fetch_webpage, and scrape_webpage could return unbounded results, risking context window exhaustion. No offset/limit/next_cursor in responses.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Convert CSV data to JSON format.
Detect the data type of a string value.
Get information about a downloadable file without downloading it.
Extract all email addresses from text.
Search for information about a specific paper across all topic directories.
Extract all links from a webpage.
Extract metadata from a webpage (title, description, Open Graph tags, etc.).
Get detailed information about a specific paper.
Extract metadata from a PDF file including title, author, creation date, etc.
Extract tables from a PDF file as structured data.
Extract all text content from a PDF file.
Extract all URLs from text.
Fetch the content of a webpage.
Find duplicate values in a list.
Generate a citation for a paper in various formats.
Convert JSON array to CSV format.
Normalize all whitespace in text to single spaces.
Parse a URL into its components (scheme, domain, path, query, etc.).
Scrape specific content from a webpage using CSS selectors.
Search for papers on arXiv based on a topic and store their information.
Search for academic papers on arXiv by topic.
Search for papers by a specific author on arXiv.
Validate if an email address is properly formatted.
Validate if a phone number is properly formatted.
Validate if a URL is properly formatted.
Parameter descriptions lack constraint details: 'max_results' (integer) has no min/max bounds stated in description. 'format' in get_paper_citation has enum values (bibtex, apa, simple) but not formalized as enum in schema. Unsafe for LLM input.
No input validation or recovery guidance: HTTPException thrown for missing papers (extract_paper_info, get_paper_citation) without suggesting next steps (try searching first). Generic 404 responses prevent LLM self-correction.
No timeout validation or safeguards on external calls: fetch_webpage, scrape_webpage, extract_links accept timeout param but do not document retry behavior or validate against runaway requests. No rate limiting.
Descriptions are generic or missing critical context: 'Clean and normalize text' (clean_text) and 'Validate if an email address is properly formatted' (validate_email) do not explain WHEN to use vs similar tools or what dependencies exist. Descriptions average ~60 chars; baselines expect 50 - 200.
No tool composition guidance: fetch_webpage and scrape_webpage are separate tools but responses do not clearly document which fields (url, css_selector) are required or how they chain. extract_text_from_pdf vs extract_tables_from_pdf vs extract_pdf_metadata lack dependency hints.
No evidence of secret injection or credential handling: Assuming any API keys or auth tokens are hardcoded or passed as env vars (not visible in source). Server should validate credentials are NOT passed as tool parameters.
Tool names could be more specific: 'extract_info' (tool 2) is vague, does it extract from papers, PDFs, or text? Rename to 'extract_paper_info_by_id' or similar. 'clean_text' vs 'clean_string' creates ambiguity.