Pseudolikelihood Maximization for Coevolution Analysis (PLMC) tools for EV+Onehot fitness prediction modeling. Generates PLMC model parameters and evolutionary couplings from protein sequence alignments.
This server has critical definition quality gaps that severely limit LLM usability. Both tools have reasonable names starting with verbs (plmc_generate_model, plmc_convert_a3m_to_a2m), and descriptions are present and substantive (218 chars and 268 chars respectively, within the 10-1024 baseline). However, the input schemas are severely incomplete: parameter descriptions are present but lack critical constraint information, no output schemas are documented anywhere, and error handling is absent. The tool definitions appear statically registered in src/tools/readme.py but cannot be fully verified from the provided source excerpt. Parameter validation rules and API contract documentation are missing entirely, making it difficult for LLMs to know how to invoke these tools correctly or handle failure modes.
Convert A3M alignment to A2M format and clean query gaps. This function performs two steps: 1. Converts A3M to A2M format using reformat.pl 2. Removes positions where the query sequence has gaps (both '.' and '-'). Cleaning query gaps is essential for downstream modeling as gaps in the query sequence can cause issues in fitness prediction workflows. Required preprocessing step as PLMC requires A2M format.
Generate PLMC model parameters and evolutionary couplings for EV+Onehot fitness prediction. This tool runs plmc with preset parameters optimized for EV+Onehot modeling and generates the required output files: model_params (binary parameter file containing model parameters) and EC file (text file containing evolutionary coupling scores). Input is protein alignment file in A2M format (use plmc_convert_a3m_to_a2m if you have A3M). Preset parameters are based on ev_onehot/scripts/plmc.sh.
No output schema documented for either tool. plmc_generate_model returns binary model_params and EC text files, plmc_convert_a3m_to_a2m returns A2M file, but LLMs have no formal schema describing return values, file paths, or success/error structure.
Parameters lack constraint specifications (min/max, regex patterns, enum values where applicable). lambda_e and lambda_h are documented as 'L2 regularization coefficient' but no valid range is specified. max_iterations is integer but no bounds given. LLMs will guess at valid ranges and may pass invalid values.
No error handling guidance. Tools modify state (WRITE risk: create binary/text files in output_dir, convert and rewrite A2M files) but provide no error recovery instructions. LLMs cannot distinguish retryable errors from permanent failures or know what corrective action to take.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Parameter alignment_path marked required=false but the tool requires A2M alignment input to function. Misleading cardinality hides a de facto required dependency. For plmc_convert_a3m_to_a2m, a3m_file and a2m_file are both required but depend on prior preprocessing; no indication that a2m_file must be writable or that a3m_file must exist.
Parameter descriptions for paths (alignment_path, a3m_file, a2m_file) are minimal. No indication of expected directory structure, absolute vs relative paths, file permissions, or disk space requirements. 'Path to protein sequence alignment file in A2M format' does not guide LLMs on how to locate or construct valid file paths.
No idempotency contract. plmc_generate_model writes to output_dir with out_prefix, if called twice with same inputs, does it overwrite? Append? Error? Agents must know whether they can safely retry on transient failures.