MCP server for LaTeX projects combining Python and TypeScript functions
The server registers 9 tools with explicit schemas and descriptions visible in server.py. However, multiple critical gaps lower the overall quality: (1) Tool names lack clarity, 'read_file', 'read_pdf', 'read_pdf_from_citation' are confusingly similar and don't disambiguate their purpose from names alone; (2) Descriptions are brief but many lack actionable context for LLM selection (e.g., 'read_file' doesn't explain when to use it vs 'read_pdf_from_citation'); (3) Parameter descriptions exist but are sparse, no constraints on formats, ranges, or error recovery paths; (4) Output schemas are underdocumented, tools return structured data but no formal response schema is declared; (5) Error handling is minimal, no recovery guidance, no categorization of errors as retryable vs. fatal; (6) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite the MCP 2026-07-28 spec supporting them. 'compile_latex' and 'download_bibliography' are write operations that should be explicitly marked. Individual tool scores range from 40-65; the average is 52.
Run LaTeX compilation (pdflatex/xelatex via latexmk) on the main document and return success plus log snippet.
Download PDFs for bibliography entries listed in resources/cited_papers/index.json.
Parse a BibTeX file and return structured entry metadata including download URLs (DOI, arXiv).
List all .tex files under the LaTeX workspace (relative paths).
Read a text/binary-safe slice of a file given a relative path inside the workspace.
Extract text & metadata from a PDF using pypdf with page/char limits (stores JSON artifact).
Resolve a citation key via resources/cited_papers/index.json to its PDF and extract text (auto-download if needed).
Three read_* tools (read_file, read_pdf, read_pdf_from_citation) are confusingly named. LLMs cannot distinguish from names alone. 'read_file' could mean any file; 'read_pdf' is ambiguous about source (disk vs citation). Convention should be read_<format>_<source>, e.g., read_pdf_by_path, read_pdf_by_citation.
Descriptions lack actionable LLM-selection context. 'read_file' doesn't explain when to use vs read_pdf_from_citation (when should I use which?). 'download_bibliography' doesn't explain prerequisites (must extract_bibliography first?). 'summarize_text' doesn't clarify algorithm or determinism. Descriptions should answer: WHAT, WHEN (vs similar tools), and WHAT NEXT.
Write operations (compile_latex, download_bibliography) lack destructiveHint annotations. MCP 2026-07-28 spec supports tool annotations, use them to signal to agents that these calls modify state and require caution.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | 0.1+ | v1 |
Suggest a stable BibTeX key based on authors/year/title metadata.
Produce a concise natural language summary of provided LaTeX / text content.
Output schemas are undocumented. No tool explicitly declares its response structure. E.g., read_pdf returns JSON with 'text' and 'metadata' fields (inferred), but the schema is not documented. LLMs cannot predict downstream availability of fields without explicit schema.
Error handling provides no recovery guidance. Tools return bare errors; LLMs don't know what to do next. E.g., 'file not found' should suggest 'Try list_tex_files() to find available files.' 'Citation key not found' should list valid keys.
Parameter constraints are missing or underspecified. E.g., 'year' in suggest_bib_key has no range (1800 - 2100?). 'max_bytes' in read_file defaults to 50k but has no documented maximum. 'max_pages' and 'max_chars' in read_pdf have minimums but no maximums, inviting unreasonable requests.
Tool composition assumes prior knowledge of workspace structure. E.g., download_bibliography reads from 'resources/cited_papers/index.json', this path is never explained. Where is this file created? How does the user populate it? This blocks effective agent planning.
summarize_text lacks implementation clarity. Is it a deterministic static summarizer, or does it call an LLM on each invocation? If it calls an LLM, which model? If deterministic, what algorithm (extractive, abstractive, heuristic)? This affects agent expectations and cost/latency.