Model Context Protocol (MCP) for mmseqs2. MMseqs2 is an ultra-fast and sensitive sequence search and clustering suite for protein and nucleotide sequences. This MCP Server provides tools for generating Multiple Sequence Alignments (MSA) using MMseqs2.
The server defines a single tool 'generate_msa' with a detailed implementation and comprehensive parameter schema. The tool name is action-oriented (verb_noun pattern), descriptions are substantive, and the input schema is well-structured with type information and defaults. However, there are significant gaps in error handling guidance, output schema documentation is implicit rather than explicit, and the tool lacks composition patterns. The implementation shows good parameter design with mutually exclusive inputs (sequence vs fasta_file) and sensible defaults (gpu=true, threads=64, sensitivity=7.5), but lacks actionable error recovery guidance and clear output structure documentation that an LLM would need to chain subsequent operations.
Generate Multiple Sequence Alignment (MSA) for a protein sequence using MMseqs2. This tool runs the complete MMseqs2 pipeline to search against a protein database and generate a multiple sequence alignment in a3m format.
Output schema not explicitly documented. Tool description states 'Returns: Either the MSA content as a string (a3m format) or the path to the output file' but does not specify the structure, field names, or format when return_format='path'. LLMs cannot plan downstream operations without knowing the exact output structure.
Error handling lacks recovery guidance. Code raises FileNotFoundError and ValueError but does not provide actionable next steps. Example: 'Database not found at: {database_path}' should suggest 'Set MMSEQS2_DB_PATH environment variable or provide database_path parameter. Available databases: [list]'.
Tool lacks destructive/idempotent annotations. The 'return_format' parameter allows returning file paths that persist in work_dir, and GPU acceleration introduces non-determinism. No tool annotations (destructiveHint, idempotentHint) present to signal side effects to the LLM.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Temporary directory cleanup not guaranteed. When use_temp_dir=True, temp_dir is created but cleanup only occurs if an exception is caught in the try/except block. If the tool succeeds, the temporary directory is never deleted, leaking disk space.
Parameter 'database_path' accepts arbitrary file paths without validation beyond existence check. No guidance on expected file format (MMseqs2 database vs arbitrary files). Agents could pass invalid paths, causing cryptic subprocess errors.
No pagination or result size limits documented. The 'max_seqs' parameter defaults to 100,000, which could produce enormous a3m output strings that exhaust token budgets. Tool description does not warn about potential context window impact.