HealthClaw Guardrails presents 29 tools with significant definition quality gaps. While tool names follow verb_noun conventions (context_get, fhir_read, fhir_search, etc.) and descriptions exist for all tools, the descriptions lack the depth and specificity required for production-grade agent tooling. Most descriptions are generic (40-80 chars) and fail to answer critical questions: WHEN should the LLM call this tool vs. similar ones? What are PREREQUISITES? What CONSTRAINTS apply? Parameter descriptions are entirely missing from the visible schema, no descriptions for resource_type, resource_id, context_id, or any other parameters. Output schemas are not documented anywhere in the provided source. Only tool names and descriptions are visible; input schemas exist but lack parameter-level documentation. No error handling guidance is present. The codebase shows HTTP transport and tool annotations (destructiveHint/readOnlyHint visible in Risk labels), but the actual schema definitions are opaque to this evaluation. Tools like fhir_propose_write, fhir_commit_write, curatr_apply_fix, action_propose, and action_commit all marked WRITE risk lack confirmation/dry-run patterns. Clinical domain tools (fhir_interpret_labs, care_gaps, guardrail_conformance) lack guidance on when to call them or what downstream decisions they enable. The tools span healthcare (FHIR, questionnaires, wearables, prescriptions, care gaps) with sophisticated guardrail logic (step-up authorization, data quality, clinical reasoning), but the definitions do not expose this richness to the agent.
Tools (29)
action_commitwriteauthsource verified53/100
Execute an approved clinical action.
action_proposewriteauthsource verified55/100
Propose a clinical action (phone call, SMS, form, etc.) for approval.
action_statusread onlyauthsource verified52/100
Check the status of a pending or completed action.
care_gapsread onlyauthsource verified50/100
Identify care gaps and quality measure gaps for a patient.
context_getread onlyauthsource verified57/100
Retrieve a pre-built context envelope with patient-centric FHIR resources. Returns bounded, policy-stamped, time-limited context.
curatr_apply_fixwriteauthsource verified50/100
Apply a data quality fix via Curatr. Requires step-up authorization.
Parameter descriptions completely missing from all 29 tools. Input schema visible but parameter-level documentation absent. LLMs cannot infer what resource_type, resource_id, context_id, or other parameters mean or accept.
Add parameter-level descriptions to ALL input schemas. For each parameter, document: expected format (string, enum values, regex pattern if applicable), range/constraints (min/max for numbers, length limits for strings), and what it controls. Example: 'resource_type: The FHIR resource type to read (e.g., Patient, Observation, Condition). Must be a valid FHIR R4 US Core v9 or R6 ballot3 resource name.'
Document output schemas for all 29 tools. Specify returned fields, their types, and structure. For paginated responses, include total_count, next_cursor, or offset fields. Example for fhir_read: 'Returns a FHIR resource object with id, resourceType, and domain-specific fields (e.g., name, birthDate for Patient; code, value for Observation). PHI-redacted per policy.'
Expand tool descriptions to 150-250 characters. Answer: (1) WHAT does it do? (2) WHEN should the LLM call it? (3) WHAT are prerequisites? Example: 'Read a single FHIR resource by ID. Use this after search results to fetch full details, or after a specific resource ID is known. Supports FHIR R4 US Core v9 and R6 ballot3. Returns PHI-redacted records per guardrail policy. If resource not found, check the resource_id and resource_type match available records.'
Add error handling guidance to all tools. Specify: (a) common error codes/conditions (resource not found, validation error, permission denied), (b) what each means, (c) actionable next step. Example: 'If resource_id does not exist, returns 404 'Not found'. Try fhir_search to locate the resource by partial name or condition. If permission denied (403), verify you have read access via fhir_permission_evaluate.'
Score history
Overall score trend
First recorded score · v2 rubric
54/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
54
2026-07-28+
v2
source verified
48/100
Evaluate clinical data quality and reconciliation using Curatr.
fetchread onlyauthsource verified42/100
Fetch a FHIR resource by URL.
fhir_commit_writewriteauthsource verified57/100
Commit a proposed FHIR resource write after step-up authorization approval.
Evaluate access control permissions for FHIR resources.
fhir_propose_writewriteauthsource verified57/100
Propose a FHIR resource write with clinical reasoning and guardrail explanations. Requires step-up authorization.
fhir_readread onlyauthsource verified62/100
Read a specific FHIR resource by type and ID. Supports FHIR R4 US Core v9 stable resources and FHIR R6 ballot3 experimental resources. Returns redacted resource with PHI protection.
fhir_searchread onlyauthsource verified65/100
Search FHIR resources with filters. Supports FHIR R4 US Core v9 and R6 ballot3. Capped at 50 results for token safety.
fhir_seedwriteauthsource verified48/100
Seed synthetic FHIR data for testing. Privileged tool.
fhir_statsread onlyauthsource verified45/100
Get aggregate statistics on FHIR resources with clinical context.
Output schemas not documented. No specification of what fields are returned, their types, or structure. Prevents agents from planning downstream tool calls or extracting required data.
Tool descriptions generic and brief (40-80 chars typical). Lack critical guidance: WHEN to call vs similar tools, PREREQUISITES (e.g., 'Requires step-up authorization' is mentioned but not explained), expected OUTCOMES. Example: 'Retrieve a pre-built context envelope...' tells WHAT but not WHEN or WHY to call before/instead of other tools.
No error handling or recovery guidance in any tool description. Tools like fhir_read, fhir_search, and fetch provide no hint on what errors they might raise or how to recover. Agents lack actionable next steps on failure.
Write/action tools (fhir_propose_write, fhir_commit_write, curatr_apply_fix, action_propose, action_commit, rx_transfer_request) lack confirmation or dry-run patterns. Descriptions mention 'step-up authorization' but do not explain the approval workflow or what rejection looks like.
Generic tool names 'search' and 'fetch' lack domain specificity. Should be 'fhir_search_generic' or clarified to indicate scope. Ambiguity invites tool selection errors when agents must choose between fhir_search (filtered) and search (generic).
Clinical decision tools (fhir_interpret_labs, care_gaps, guardrail_conformance) lack guidance on what downstream actions they enable or trigger. No indication of when to call them in clinical workflows or what decisions they support.
Pagination and result limits not mentioned in any search or list tool (fhir_search, search, fhir_lastn, fhir_stats). Description of fhir_search mentions 'capped at 50 results' but does not specify how to paginate or retrieve additional results.
Tool fhir_get_token marked as READ_ONLY but semantically generates/retrieves a token. Risk classification may be misleading. Privileged tools (fhir_get_token, fhir_seed) lack explicit scope/permission declarations.
fhir_get_tokenfhir_seed
Implement confirmation or dry-run pattern for all WRITE tools (fhir_propose_write, fhir_commit_write, curatr_apply_fix, action_propose, action_commit, rx_transfer_request). Document the approval workflow: (1) propose returns a proposal_id, (2) agent/user reviews, (3) commit applies the change. Clarify what rejection/timeout looks like and how to cancel a pending proposal.
Clarify step-up authorization workflow in tool descriptions. Explain: What is the authorization gate? What user/role can approve? How long is approval valid? What happens if the user rejects or times out? Example: 'fhir_propose_write: Propose a write but do NOT execute. Returns proposal_id and pending authorization status. Requires step-up authorization (agent requests user/clinician confirmation via action_propose before committing).'
Rename generic tools 'search' and 'fetch' to domain-specific names. Suggested: 'fhir_search_generic' (for broad FHIR queries) and 'fhir_fetch_by_url' (for direct URL fetch). Clarify when to use fhir_search (filtered, bounded) vs. fhir_search_generic (all fields, unfiltered).
Add pagination support documentation. For fhir_search, fhir_lastn, fhir_stats, specify: (a) max results per call (document the 50 limit for fhir_search), (b) cursor/offset parameter name, (c) how to detect end of results. Example: 'fhir_search: Returns up to 50 results. To fetch more, include cursor from previous response in next call. Example: fhir_search(resource_type='Observation', cursor='next_page_token_xyz').'
Document clinical decision context for domain-specific tools. For fhir_interpret_labs, care_gaps, guardrail_conformance, explain: When should the LLM call this in a patient interaction? What decisions does it support? What other tools should follow? Example: 'care_gaps: Identify missing care (e.g., due screenings, needed treatments). Use AFTER retrieving patient history to propose next clinical actions (phone call, form, prescription). Results guide action_propose and action_commit.'
Add scope/permission declarations to privileged tools. For fhir_get_token and fhir_seed, document: What permissions are required to call this? What scopes do they grant? Example: 'fhir_seed [PRIVILEGED]: Requires admin:seed scope. Populates test FHIR data in isolated tenant. For testing only. Do not call in production without explicit authorization.'
Standardize error message format across all tools. Provide actionable recovery guidance, not just error codes. Example: 'Invalid resource_type: "BadType". Must be one of: Patient, Observation, Condition, Medication, .... (See FHIR R4 US Core v9 for full list.) Or try fhir_subscription_topics to list available resource types.'
Add examples of tool chaining in descriptions. Show common multi-step workflows. Example: '1. fhir_search to find patient observations. 2. fhir_read to fetch full Observation details. 3. fhir_interpret_labs to analyze results. 4. action_propose to recommend follow-up to clinician.' This guides agent planning.
Document mutation semantics for WRITE tools. For fhir_propose_write, fhir_commit_write, action_commit, curatr_apply_fix, explicitly state: Are these idempotent? What happens if called twice with the same input? Can they be safely retried? Example: 'fhir_commit_write: Idempotent. Calling with the same proposal_id commits once and returns the same committed resource. Safe to retry on timeout.'
Add data source freshness/timing info to read tools. For sources_check, wearables_sync_status, fhir_compiled_truth, specify: How fresh is the data? When was it last synced? What delay should agents expect? Example: 'wearables_sync_status: Returns device sync status. Data typically 1-5 minutes old depending on device. If real-time data needed, recommend manual patient check-in via action_propose.'
Clarify FHIR version support uniformly. All tools mention 'FHIR R4 US Core v9 and R6 ballot3' but don't clarify behavior differences, deprecations, or fallbacks. Add a single reference doc or clarify per tool. Example: 'Supports FHIR R4 US Core v9 (stable) and R6 ballot3 (experimental). If resource exists in only one version, returns available version. If incompatible versions, returns error with available alternatives.'