Deterministic HLA verification as agent tools. Nomenclature and reference-release validation for HLA allele names, typing strings, GL strings and report text against a pinned IPD-IMGT/HLA release.
HLA-Verify demonstrates strong domain-specific tool design with clear, detailed descriptions and proper input schemas. All 7 tools have descriptions (avg 280 chars, well above 194-char baseline) and documented parameters. Tool names are action-oriented (verify_*, normalize_*, match_*, check_*, validate_*, about). However, output schemas are not explicitly documented in the visible code, only input schemas are shown. The `about` tool has no input parameters but is properly registered. Tool annotations (READ_ONLY) are correctly applied to all tools, signaling idempotency and read-only semantics. Parameter descriptions are domain-specific and actionable (e.g., 'allele strings only, no patient identifiers'). Error handling is implicit (tools return structured dicts with flags and verdicts) but not explicitly documented in descriptions. No evidence of secrets in parameters or destructive operations.
What this server is and is not, what to send it, benchmark evidence for why to use it, and terms.
QC-check one HLA typing (all loci) against the pinned release: resolves every reported allele, flags unresolvable/outdated/locus-mismatched/null alleles, flags too-many/single/homozygous per locus, computes the B-leader (-21 M/T) and KIR-ligand (C1/C2/Bw4) profile, and DRB3/4/5 expected-vs-reported. Nomenclature and internal-consistency checking of the report, not clinical interpretation. typing: {"A": ["A*01:01", "A*02:01"], "B": [...], "DRB1": [...], ...} (any nomenclature era; allele strings only, no patient identifiers).
Donor/recipient immunogenetic compatibility under two published rule sets: HLA-B leader match (-21 M/T, Petersdorf 2020) for a single HLA-B mismatch, and KIR ligand (C1/C2/Bw4) class comparison, computed over each side's full typing QC. Rule checking against published frameworks; it does not rank or recommend a donor. recipient/donor: {"A": [...], "B": [...], "C": [...], "DRB1": [...], ...} (allele strings only, no patient identifiers). Decision support only; not a medical device.
Count a donor-recipient HLA match by the published counting rules (R1-R6): allele arithmetic over chromosomes, not a donor recommendation. recipient/donor: {"A": ["A*01:01","A*02:01"], "B": [...], ...} (two reported alleles per locus, any nomenclature era; allele strings only, no patient identifiers). framework: 6/6, 8/8, 10/10, 12/12, or antigen. Returns count, per-locus verdicts, GvH/HvG mismatch counts, and flags; unresolvable typing yields 'potential', never a confident count.
Output schemas not documented. Tool descriptions explain inputs but not return structures. LLMs cannot plan downstream calls or extract fields without knowing what each tool returns.
Error handling not explicitly documented. Tools return structured dicts with flags (e.g., 'unresolvable', 'potential') but descriptions do not explain what these mean or how LLMs should interpret them.
No pagination or result-limiting guidance. Tools like verify_text could return large lists of allele classifications; no mention of limits or how to handle high-volume inputs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 70 | 2026-07-28+ | v2 |
Normalize one reported HLA allele name (any era) to current 2-field form, with G group, P group, serologic equivalent, and flags. An allele string only, never a patient name, medical record number or other identifier.
Validate and normalize a GL String (Genotype List, ^ | + ~ / grammar): resolves every allele token, flags outdated/unresolvable names and structural problems (mixed loci within a slash-list, a repeated locus within a haplotype or across ^ blocks, more than two haplotypes, differing loci across a genotype or genotype list, empty elements), and returns the normalized string. Grammar and nomenclature checking only; send allele names, not patient identifiers.
Scan HLA typing report text, or model output about HLA, for allele-shaped tokens and classify each one: valid / legacy (with modern form) / deleted (with successor) / fabricated. Nomenclature checking against a pinned IPD-IMGT/HLA release, not interpretation of a case. Use on any AI-generated or transcribed content mentioning HLA. Send the HLA content only, with patient identifiers removed first: the caller is responsible for de-identifying the text.
Tool composition not documented. No guidance on which tools to call in sequence (e.g., 'call check_typing first, then match_score'). Agents must infer the workflow.
Generic tool name 'about' lacks action verb. Should be 'describe_server' or 'get_server_info' to match verb_noun convention and clarify intent.