Provides both synchronous and asynchronous APIs for protein stability prediction using SPIRED-Stab. Supports CSV and FASTA inputs, single variants, and batch processing with job management for long-running operations.
The server defines 8 tools with mixed quality. Tool names follow verb_noun conventions well (get_*, list_*, submit_*, cancel_*, predict_*, analyze_*), and all tools have descriptions present in the source code. However, critical gaps exist: (1) Parameter schemas lack enum constraints for status filters (list_jobs accepts 'status' as free-form string when it should enumerate: pending|running|completed|failed|cancelled). (2) Descriptions are present but generic, most lack WHEN/WHY guidance that LLMs rely on for tool selection. (3) Output schemas are not documented for any tool, LLMs cannot plan downstream calls without knowing what fields are returned. (4) Error handling is minimal, no recovery guidance. (5) The 'device' parameter in predict_stability and related tools accepts free-form strings ('cuda:0', 'cpu', 'cuda:1') without enum constraints, inviting hallucinated GPU device IDs. (6) Several tool descriptions exceed the recommended length (predict_stability's description is ~550 chars, well above the 200-char sweet spot for LLM processing). (7) No examples of output structure for critical tools like predict_stability or submit_stability_prediction. (8) Parameter descriptions for variant_file and wt_fasta_file are adequate but could specify more about path validation and existence requirements. Average per-tool quality is fair but below production baseline.
Analyze a single protein variant for stability changes (synchronous API). This is a convenience function for analyzing just one variant sequence without needing to create input files.
Cancel a running job.
Get log output from a running or completed job.
Get the results of a completed job.
Get the status of a submitted job.
List all submitted jobs.
Predict protein stability using SPIRED-Stab (synchronous API for small batches). Use this for quick predictions with fewer than 50 variants. For larger batches, use submit_stability_prediction() for background processing. This tool predicts the stability changes (ddG and dTm) for protein variants compared to a wild-type sequence. It supports both CSV and FASTA input formats.
Output schemas are not documented for any tool. LLMs cannot plan downstream calls without knowing return field names and types.
Free-form string parameters (status, device) lack enum constraints. 'status' in list_jobs explicitly enumerates valid values in the description but NOT in schema, so LLM may pass 'pending_review' or other invalid values. 'device' in predict_stability, analyze_single_variant, and submit_stability_prediction accepts arbitrary strings, LLM may hallucinate non-existent GPU device IDs.
Overlapping tool intent: predict_stability and analyze_single_variant both perform stability prediction. Tool selection guidance is missing, when should LLM choose one vs the other? Description mentions '<50 variants for predict_stability' and '>50 variants use submit_stability_prediction', but no guidance distinguishes analyze_single_variant.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
Submit protein stability prediction for background processing (large batches). This operation is suitable for large batches (>50 variants) that may take more than 10 minutes. Returns a job_id for tracking progress.
Error handling is minimal. No recovery guidance in any tool description. e.g., get_job_result does not explain what error to expect if job not found or still running. cancel_job does not clarify if operation is idempotent (can you cancel an already-cancelled job safely?). predict_stability catches FileNotFoundError and ValueError but LLM sees opaque 'File not found' responses without guidance on what to try next.
predict_stability description is 550+ chars (far above the 200-char baseline for optimal LLM processing). Verbose preamble and example usage waste tokens and bury key decision logic. Should be condensed: 'Predict protein stability synchronously for <50 variants. Returns ddG and dTm predictions. For larger batches, use submit_stability_prediction().'
Parameter inconsistencies across similar tools. predict_stability makes wt_fasta_file optional with fallback to 'wt.fasta'; analyze_single_variant makes it required. No explanation for the difference. LLM may assume both behave the same way and omit the parameter in analyze_single_variant calls, causing failures.
list_jobs output schema is undocumented. Description says 'List of jobs with their status' but does not specify fields per job (job_id, status, timestamps, etc.). LLM cannot extract what it needs without guessing the response structure.
No pagination support documented for list_jobs. If many jobs accumulate, LLM has no way to limit or paginate results, risking context window exhaustion.