Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
The server provides 17 tools across HCM analysis, RAG/documentation retrieval, and advanced reasoning (repair, reconciliation, inverse design). Tool naming is generally clear and verb-driven (analyze_, describe_, search_, query_, etc.). However, descriptions vary widely in depth and LLM-friendliness, some are detailed narratives (e.g., repair_design), while others are terse or vague (e.g., call_tool, list_tools). Input schemas are present for all tools but lack rigor: many parameters have minimal or no descriptions (especially in batch_query_hcm, research_analytics), enum constraints are missing where they should exist (facility_type should be an enum, not free-form), and optional/required distinctions are inconsistent. Output schemas are not documented, LLMs cannot see what fields to expect from responses. Error handling is absent from all tool definitions. Security is partially addressed (READ_ONLY risk marking), but no guidance on permission gates, secret injection, or rate limiting is visible. The reasoning tools (repair_design, repair_freeway, reconcile_codes, inverse_design, validate_design_full) are sophisticated but their descriptions assume domain knowledge and don't explain prerequisites or failure modes clearly.
Tools (17)
analyze_facilityread onlysource verified70/100
Run the complete HCM analysis for one facility: pass facility_type plus its inputs; dispatches to the verified transportations-library executor for that facility and returns the full result.
Describe facility input schemas: with facility_type, the field schema for constructing an analyze_facility request; without it, every library-backed HCM facility type with chapter and adapter status.
diagnose_failureread onlysource verified68/100
Backward-chain: upstream causes of a failed target parameter.
Add explicit JSON response schema to every tool definition. Example: 'Returns: {type: object, properties: {facility_type: {type: string}, analysis_result: {type: object}, los: {type: string, enum: [A, B, C, D, E, F]}, capacity: {type: number}}, required: [los, capacity]}'. This allows LLMs to plan downstream tool calls.
Replace free-form facility_type parameters with proper enums: {type: 'string', enum: ['TwoLaneHighway', 'BasicFreeway', 'ManagedLaneFacility', 'PlanningFacility', 'AlternativeIntersection', 'DisplacedLeftTurn', 'BicycleLOS']}. Verify against transportations-library for current list.
Expand research_analytics description: 'Returns aggregated statistics about the HCM research database: total documents indexed, embedding model version, vectordb config, last sync time. No parameters required. Use to verify RAG deployment readiness before calling other query tools.'
Clarify call_tool vs analyze_facility in tool descriptions. call_tool appears to be a meta-tool for generic function invocation; document its actual purpose and when an LLM should prefer it over domain-specific tools like analyze_facility. If call_tool is internal scaffolding, consider hiding it from LLM exposure.
For repair_design, repair_freeway, reconcile_codes, inverse_design, document nested object schemas inline. E.g., design parameter description should list: 'Object with rust_field keys (lane_width, shoulder_width, apd, volume, …). All values numeric. Required: [lane_width, volume]. Optional: [passing_type].' Use a nested JSON Schema definition if possible.
Score history
Overall score trend
↑ 13 points across a rubric change (v1 → v2)
60/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
C
60
2026-07-28+
v2
2026-03-09
F
47
-
v1
Goal-directed synthesis: feasible geometries that reach a target LOS, each validated by forward execution (HCM Ch.15 PoC).
list_toolsread onlysource verified53/100
List all available tools with optional filtering.
propagate_changeread onlysource verified68/100
Forward-chain: downstream parameters affected by a changed root parameter.
query_hcmread onlysource verified68/100
Query HCM documentation.
reconcile_codesread onlysource verified68/100
Defeasible adjudication of overlapping/conflicting code provisions about one parameter, with an argument trace.
repair_designread onlysource verified75/100
Abductive repair: minimal compliant fix for a failing Two-Lane Highway (HCM Ch.15). Every candidate is re-executed and proved compliant.
repair_freewayread onlysource verified75/100
Abductive repair: minimal compliant fix for a failing Basic Freeway (HCM Ch.12). Keep grade/length on the library's heavy-vehicle PCE grid.
research_analytics tool has empty input schema and minimal description ('Get analytics about HCM research database.'). Cannot infer parameters or behavior.
call_tool and list_tools descriptions are generic or vague. 'Execute a function call' (call_tool) does not explain what functions are available, when to use it vs analyze_facility, or what the function signature should be.
No error handling guidance in any tool definition. LLMs cannot distinguish retryable failures (network timeout) from user-fixable errors (invalid facility_type) from unrecoverable errors. No recovery paths documented.
Parameter descriptions often lack format/constraint details. E.g., goal_los in repair_design says 'Target LOS letter, repair goal is no worse than this' but does not specify allowed values (A - F presumably) or consequences if omitted.
No pagination or result limits documented for retrieval tools (query_hcm, search_hcm_by_chapter, batch_query_hcm). How many results are returned? Can they overflow context? Are they paginated?
Reasoning tools (repair_design, repair_freeway, reconcile_codes, inverse_design) have complex input objects (design, context, bounds) but parameter descriptions do not list required sub-fields or valid keys, leaving LLMs to guess schema.
Add error handling notes to each tool: 'On invalid facility_type, returns {error: 'unsupported_facility', suggestion: 'Did you mean TwoLaneHighway or BasicFreeway?'}. On missing required inputs, returns {error: 'missing_fields', fields: ['volume', 'grade']}. Retry with complete input.'
For retrieval tools (query_hcm, search_hcm_by_chapter, batch_query_hcm), document result limits: 'Returns at most top_k results (default 5, max 50). Each result is {document_id, chapter, section, text, relevance_score}. If result count == limit, call again with an offset to get next page.'
Add example inputs and outputs (separate from descriptions) in tool metadata. E.g., example input: {facility_type: 'TwoLaneHighway', ... }, example output: {los: 'C', capacity: 1800, measures_of_effectiveness: {...}}. This guides LLM without polluting the description field.
Declare tool immutability/idempotence: analyze_facility, query_hcm, etc. are read-only and safe to retry. repair_design, reconcile_codes are reasoning/analysis tools that transform input but do not persist state, idempotent but expensive.
Add prerequisite hints: E.g., 'Before calling repair_design, call analyze_facility with the same design to understand the baseline LOS and why it fails. repair_design will then find minimal fixes.'
For goal_los enum values, explicitly list: 'Target level of service (A, B, C, D, E, or F). A=best, F=worst. Default: C. repair_design will find minimal changes to achieve no worse than this goal.'
Document immutable parameter override in repair_design/repair_freeway: 'immutable array lists parameter names that are site constraints and cannot change. Default: [] (all parameters can be changed). Common site constraints: [volume, grade, length, jurisdiction]. This guides the repair algorithm.'
For batch_query_hcm, clarify behavior: 'Submits queries array. Runs each query independently (not a combined search). Returns array of result-lists, one per query, in input order. Each result-list contains at most top_k items. Example: queries=['bicycles', 'freeway'], returns [[{chapter: 10, ...}, ...], [{chapter: 12, ...}, ...]].'
Add schema validation error responses to tool descriptions: 'If input violates schema (e.g., goal_los not A-F), the tool returns {error: 'invalid_input', field: 'goal_los', expected: 'A|B|C|D|E|F', got: 'G'}. Fix and retry.'