Model Context Protocol server for AgenticRail. Exposes the live gate to any MCP client (Claude Code / Desktop, Cursor, …) as two tools: evaluate_step and verify_receipt. This server is a thin PROTOCOL ADAPTER that calls the same PUBLIC API a real user calls with no service bindings to internal gate or internal secrets.
Two tools with complete input schemas and detailed descriptions. evaluate_step has excellent parameter documentation with enums, constraints, and dependency hints (e.g., 'step must equal function', 'step_order is LOCKED on first call'). verify_receipt is minimal but functional. Both tools lack documented output schemas, critical for LLM planning. No error handling guidance (e.g., what to do on DENIED, HALT, or verification failure). Tool names are verb-noun (evaluate_, verify_) and clear. Descriptions are 400+ chars, exceeding the 10-1024 baseline but comprehensive. Parameters are well-typed with descriptions; action_type enum is load-bearing and correctly includes all eight values with per-step constraints documented in prose.
Ask the AgenticRail gate to ALLOW or DENY a single step of an agent sequence BEFORE it runs. The gate is deterministic (same state+request → same verdict) and enforces step order, replay protection (nonce), timestamp freshness, and sealing. A denied step must not be executed. Every decision is sealed into an Ed25519-signed, hash-chained receipt. Returns the decision (ALLOW/DENY/HALT), any reason codes, and receipt metadata. Use the demo key by sending no Authorization header, or send Authorization: Bearer <your-key>. NOTE: an anonymous call has its sequence_id rewritten to 'demo-mcp-<your id>'. This is intended, not a leak: it scopes the run to the public demo lane and is how anonymous MCP traffic is identified. Always reuse the sequence_id RETURNED in the response for later steps and for verify_receipt -- the id you sent will not resolve.
Verify a receipt returned by evaluate_step. Receipts are Ed25519-signed, hash-chained, and sealed by the gate. This tool POSTs the receipt to the report endpoint and returns verification status.
Output schemas not documented. LLMs cannot plan downstream calls or extract decision/receipt fields without knowing the response structure. evaluate_step returns (decision, reason_codes, receipt_metadata) and verify_receipt returns (verification_status), neither is formally specified.
No error handling guidance. Descriptions mention ALLOW/DENY/HALT decisions and ACTION_NOT_ALLOWED denials but do not tell the LLM what to do on failure. E.g., 'If denied with ACTION_NOT_ALLOWED, read allowed_action_types and retry' is buried in prose; should be in error response structure.
verify_receipt description is minimal (60 chars). Does not explain when to call it, what verification failure means, or what the receipt object structure is. LLMs cannot determine if this is a post-hoc audit or a required step before execution.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 73 | 2026-07-28+ | v2 |
No idempotency or confirmation pattern for evaluate_step. The tool enforces replay protection (nonce) and sealing, but the description does not clarify whether repeated calls with the same sequence_id+step are safe or will be rejected. Agents need to know if they can safely retry.
Authorization header handling is implicit. The description mentions 'send Authorization: Bearer <your-key>' but the tool schema does not expose an auth parameter. This is correct (secrets should not be tool params), but the mechanism (HTTP header forwarding) is not documented in the tool definition itself, only in prose.