A unified server providing REST API and MCP interfaces for autopvs1.bgi.com PVS1 variant data
AutoPVS1-Link demonstrates strong naming conventions, detailed parameter descriptions, and comprehensive schemas for variant/CNV scoring. Tool names follow verb_noun patterns (get_*, search_*, clear_*) with clear distinctions. All 6 tools have descriptions exceeding the 10-char minimum, with most ranging 150-500+ chars. Input schemas are fully defined with JSON Schema types, enums, and constraints. However, output schemas are not explicitly documented in the source, responses are described narratively but lack formal structure definitions. Tool descriptions are exceptionally detailed (some exceed 1024 chars), providing context on response modes, cache behavior, and error handling. Error handling includes specific recovery guidance (e.g., 'requires_disambiguation', 'external_resolver_unavailable', 'invalid_bulk_input'). Composition is excellent: 6 focused tools with clear responsibilities; bulk variants (get_variants_pvs1_data_bulk, get_cnvs_pvs1_data_bulk) reduce wasteful sequential calls. Security: clear_cache is gated behind AUTOPVS1_LINK_ENABLE_DESTRUCTIVE_TOOLS env var. Main weaknesses: (1) output schemas not formally documented, (2) some descriptions are verbose enough to dilute LLM token budget, (3) no explicit pagination cursoring guidance in search_variants despite mentioning cursor tokens.
Clear all service caches. Disabled by default. Enable with AUTOPVS1_LINK_ENABLE_DESTRUCTIVE_TOOLS=true.
Score one copy-number variant with the AutoPVS1 PVS1 rules. First-turn LLM callers get the verdict under ~1.5KB by default (``response_mode='summary'``). Widen to ``response_mode='standard'`` for the full decision tree. AutoPVS1 outputs are research-use only, not clinical decision support.
Score 1-10 copy-number variants in one call. Prefer this over ``get_cnv_pvs1_data`` when you have 2+ CNV IDs. For LLM batch screens, default to ``response_mode='summary'`` so 10 verdicts share one turn budget; widen per-item only when reasoning needs the full decision tree. Items run sequentially server-side and respect the upstream rate limit (default ~1 req/s) plus the existing cache, so a fully uncached 10-item batch can take ~10s wall time and a fully cached one returns in milliseconds. Per-item envelope: each row in the top-level ``results`` array has ``{ok, input, data, error, meta}`` where ``meta.cache_status`` and ``meta.elapsed_ms`` echo that one upstream call's outcome (absent when the item short-circuited before upstream). This per-item shape predates and is scoped separately from the Response-Envelope Standard v1 outer frame. Output items preserve input order. ``response_mode`` and ``include_unmet`` apply per item; the outer ``meta_mode`` controls the envelope. Per-item failures do not stop the batch unless ``continue_on_error=false``. Bulk dispatch errors (malformed ``items``) use ``error_code='invalid_input'`` (subcode ``invalid_bulk_input``). Aggregate cache observability: top-level ``_meta.cache_status`` echoes the unanimous status when every item agrees; on a mixed batch it is ``"mixed"`` and ``_meta.cached_count`` / ``_meta.uncached_count`` split items by warm (``hit``+``coalesced``) vs cold (``miss``+``bypass``). ``_meta.elapsed_ms`` is the SUM of per-item upstream wall-clocks (the honest total for a sequential bulk). Warning aggregation: per-item warnings are NOT echoed; they are collapsed into ``_meta.warnings`` at the top level. A warning code is aggregated only when more than one distinct item emitted it; single-item codes appear without ``count`` or ``affected_indices``. Aggregated codes carry ``count`` (distinct items) and the sorted ``affected_indices`` list. Order is first-seen-code-first.
Output schemas not formally documented. Tool descriptions explain response structure (per-item envelope, _meta fields, cache_status) but no JSON Schema definitions for return types are visible. LLMs cannot plan downstream operations without knowing returned field names and types.
Tool descriptions exceed optimal length. get_variants_pvs1_data_bulk (730+ chars) and get_cnvs_pvs1_data_bulk (730+ chars) descriptions are verbose, diluting LLM token efficiency. Baseline for A+ is 50-200 chars. Excessive detail on cache semantics, warning aggregation, and response envelope structure belongs in developer docs, not the tool description.
search_variants cursor token opaque behavior not clearly stated. Description says 'base64url JSON today (decodable) but treat it as an echo-back token; it MAY become opaque later', this hedging confuses LLMs about whether they should decode/validate the token or blindly echo it. Provide clear guidance on cursor handling.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2025-06-18+ | v2 |
Return local MCP server health. Default behaviour: no upstream call, sub-millisecond. Pass ``check_upstream=true`` for an opt-in HEAD probe — useful when an agent wants to confirm AutoPVS1 is reachable before scheduling a cold scoring call.
Score 1-10 SNV/indel variants in one call. Prefer this over ``get_variant_pvs1_data`` when you have 2+ variant IDs of the same kind. For LLM batch screens, default to ``response_mode='summary'`` so 10 verdicts share one turn budget; widen per-item only when reasoning needs the full decision tree. Items run sequentially server-side and respect the upstream rate limit (default ~1 req/s) plus the existing cache, so a fully uncached 10-item batch can take ~10s wall time and a fully cached one returns in milliseconds. Auto-resolution applies per item: non-canonical inputs (rsID, HGVS c./p./g.) round-trip through Ensembl Variant Recoder before scoring, mirroring ``get_variant_pvs1_data``. Multi-candidate resolutions return per-item ``requires_disambiguation`` with allele-keyed candidates so the caller picks one and re-calls that single item; a resolver outage returns the retryable ``external_resolver_unavailable`` code. Per-item envelope: each row in the top-level ``results`` array has ``{ok, input, data, error, meta}`` where ``meta.cache_status`` and ``meta.elapsed_ms`` echo that one upstream call's outcome (absent when the item short-circuited before upstream). This per-item shape predates and is scoped separately from the Response-Envelope Standard v1 outer frame. Output items preserve input order. ``response_mode`` and ``include_unmet`` apply per item; the outer ``meta_mode`` controls the envelope. Per-item failures do not stop the batch unless ``continue_on_error=false``. Bulk dispatch errors (malformed ``items``) use ``error_code='invalid_input'`` (subcode ``invalid_bulk_input``). Aggregate cache observability: top-level ``_meta.cache_status`` echoes the unanimous status when every item agrees; on a mixed batch it is ``"mixed"`` and ``_meta.cached_count`` / ``_meta.uncached_count`` split items by warm (``hit``+``coalesced``) vs cold (``miss``+``bypass``). ``_meta.elapsed_ms`` is the SUM of per-item upstream wall-clocks (the honest total for a sequential bulk). Warning aggregation: per-item warnings are NOT echoed; they are collapsed into ``_meta.warnings`` at the top level. A warning code is aggregated only when more than one distinct item emitted it; single-item codes appear without ``count`` or ``affected_indices``. Aggregated codes carry ``count`` (distinct items) and the sorted ``affected_indices`` list. Order is first-seen-code-first.
Search AutoPVS1 by gene symbol or variant text. Use ``response_mode='ids_only'`` (lowest-bandwidth lookup) to resolve a query to an AutoPVS1 ``variant_id`` you can hand to ``get_variant_pvs1_data``. ``next_cursor`` is base64url JSON today (decodable) but treat it as an echo-back token; it MAY become opaque later. AutoPVS1 outputs are research-use only, not clinical decision support.
get_server_health check_upstream parameter semantics unclear. Description says 'sub-millisecond' for no probe, but does not state explicit timeout or fallback behavior if upstream is slow/hung. LLMs need predictability to avoid blocking during planning.