YAML-first AI agent platform — define AI agent roles in YAML and run them anywhere (CLI, API server, or autonomous daemon)
InitRunner provides 11 tools with clear verb-noun naming and consistent parameter structures. However, quality is uneven: tool descriptions are brief (mostly 10-50 chars) and lack LLM-optimized context about WHEN to use each tool. Parameter descriptions exist but are minimal. Output schemas are not documented, we can infer some structure (e.g., search_inbox likely returns email metadata, read_email returns content) but cannot verify from source. Error handling and recovery guidance are absent. The tools themselves are well-scoped (one action each), but descriptions fall short of production standards (baseline 194 chars average; these average ~65 chars). Tool composition is reasonable, related tools (search_inbox, read_email, list_folders) work together, but descriptions don't guide the LLM on sequencing or dependencies.
Convert a numeric value between common measurement units. Supported conversions: km/mi, kg/lb, c/f, l/gal, m/ft, cm/in.
Get the current date and time. Leave timezone empty to use the default.
Pretty-print a JSON string with 2-space indentation.
Generate a random UUID v4 identifier.
Hash text using the specified algorithm (md5, sha1, sha256, sha512).
List all available mailbox folders.
Look up a query using the configured prefix and source. The tool_config parameter is injected by InitRunner from the role YAML and is hidden from the LLM.
Tool descriptions are below production baseline (average ~65 chars vs. baseline 194 chars). Most lack context about WHEN to use the tool, prerequisites, or how it differs from similar tools. For example, 'Get the current date and time' does not explain that timezone is optional or what happens if invalid timezone is provided.
Output schemas are not documented in source. Cannot verify what fields search_inbox, read_email, list_folders, or other tools return. This forces LLMs to guess what data is available for chaining downstream tool calls. Responses should declare field names, types, and whether pagination is supported.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Parse a date string. Leave format empty for ISO 8601 auto-detection.
Read the full content of an email by its Message-ID.
Search for emails using IMAP SEARCH syntax.
Count words, characters, and lines in a text string.
No error handling or recovery guidance documented. What happens when search_inbox query is malformed, folder does not exist, or IMAP connection fails? What should the LLM do next? Tools lack actionable error messages or categorization (retryable vs. fatal).
Parameter descriptions are minimal. For example, search_inbox.limit says 'Maximum number of results to return. 0 or negative uses configured max_results.', but what IS max_results? What is the typical range? Is there a hard cap? Similar gaps in convert_units (no mention of valid unit pairs) and hash_text (no mention of what 'default' means).
lookup_with_config has opaque semantics. Description says tool_config is 'injected by InitRunner from role YAML.' But what does lookup do? What does 'prefix and source' mean? LLM cannot infer intent from this description. Clarify what a lookup returns and how it differs from search_inbox.
Pagination is mentioned for search_inbox but not formally specified. Does it support limit + offset, cursor, or neither? What does it return if limit=0? Can you pass limit=1000 or is there a cap? Baseline tools with paginated results declare page/offset/limit/total_count structure.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present in visible schemas. While all 11 tools appear to be read-only (no state mutations), explicit annotation would clarify this for clients and improve protocol compliance.