Commercial-air viability for time-critical organ transport. A deterministic margin engine and WebMCP tool surface that withdraws a recommendation the moment the evidence it rested on stops being true.
Strong domain-specific tool design with excellent naming, clear descriptions, and well-structured schemas. All 12 tools follow verb_noun convention (create_, list_, evaluate_, inspect_, explain_, select_, get_, revalidate_, advance_, replan_). Descriptions are detailed (100-250 chars) and explain WHAT, WHEN, and WHY. Input schemas use proper JSON Schema with type constraints, patterns, enums, and min/max bounds. However, output schemas are not documented, LLMs cannot predict response structure for downstream chaining. Error handling guidance is absent. Tool composition is excellent: each tool has one responsibility, and outputs contain IDs for chaining (e.g., option_id returned by evaluate_flights is consumed by explain_constraints and select_transport_plan).
Release the next recorded event in the captured flight-data timeline — an aircraft reassignment, an inbound delay increasing — and wait for this page's evidence monitor to observe it. The monitor polls on its own timer and delivers the same event unasked; this only brings it forward, which is how a judge sees it on demand. On a live feed there is nothing to release, so this forces an immediate observation instead.
Open an organ-transport mission: origin and destination airport, when the shipment is ready, and the latest acceptable delivery. Called with no arguments it opens the captured mission, ORD to BOS. Returns the mission id and the normalised operational constraints, including the cargo acceptance lead and destination release times that every later margin is computed against.
Compute, for each candidate flight, whether the mission still works: cargo acceptance, aircraft rotation, propagated inbound delay and destination delivery. Returns a signed margin in minutes, the binding constraint and a verdict of VIABLE, TIGHT, BROKEN, UNKNOWN or STALE per flight. Omit flight_numbers to evaluate the whole corridor. Nothing here is a probability; every figure is arithmetic over observations.
Break one evaluated option into its three constraints, each with the two instants its margin is the difference between, plus the propagated delay floor and the empirical P50, P80 and P95 margins with their sample count. Use it to say why a flight fails, and which party to call about it.
Output schemas not documented. LLMs cannot predict response structure (fields, types, nested objects) for downstream tool chaining. Responses like 'margin in minutes', 'binding constraint', 'verdict' are described in prose but not in machine-readable schema.
No error handling guidance. Tools do not document what errors can occur, whether they are retryable, or what the LLM should do next. E.g., evaluate_flights could fail if flight data is stale, but no recovery path is documented.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 81 | 2026-07-28+ | v2 |
Explain why the previously selected plan stopped being valid: the reason code, the field-by-field difference between the evidence the recommendation was issued against and the evidence now, and how many minutes of mission margin moved. Present only when an invalidation has actually been recorded.
Report which feed the flight data comes from, whether this page's evidence monitor is running, what it has observed so far, and what the current mission state is. Read it before trusting any figure: this build replays captured observations on a timer rather than polling a live provider, and says so here in a typed field.
Return the transport plan currently in force, with its certificate and the observations behind it. The presence of this tool is itself the claim that the plan is still valid: it is unregistered the moment its evidence stops holding, so an agent cannot read a stale plan from it.
Show the aircraft assigned to a flight and where it is now: already at the origin since a given time, or still operating an inbound sector with a scheduled and a projected arrival. This is the evidence a timetable does not contain, and the usual reason a flight that looks fine is not.
List the commercial flights in the capture between two airports, with scheduled times and the aircraft assigned to each. This is the timetable view and nothing more: it does not say whether a flight works for the mission. Use evaluate_flights for that.
Re-evaluate the whole corridor against current evidence after a plan has been invalidated, and recommend the best option that still passes every constraint. Restores select_transport_plan so a new plan can be committed. This is the only route back from an invalidated plan.
Recheck the selected plan against current evidence and return whether it still holds. The page also checks this on every evidence change without being asked; this tool exists so an agent can confirm rather than assume.
Commit the mission to one evaluated option and receive a recommendation certificate: the margin, the binding constraint, the exact evidence the recommendation rests on, a validity horizon and a fingerprint over that evidence. Only offered while an option passes every constraint. Selecting withdraws this tool and publishes get_selected_plan and revalidate_plan.
No pagination or result limits documented. list_candidate_flights could return hundreds of flights; no limit, offset, or cursor parameters are present. Large result sets will exhaust context windows.
Stateful tool availability (select_transport_plan withdraws itself, get_selected_plan only available after selection). This is clever domain design but not documented in tool descriptions. LLMs may attempt to call unavailable tools without understanding why.
No confirmation or dry-run pattern for irreversible operations. select_transport_plan commits a mission plan with no preview or confirmation step. Agents could select a suboptimal plan without recourse.