A Swift framework for building multi-turn agentic systems with tool use, streaming, approval workflows, and sub-agent composition
AgentRunKit provides a well-structured coding toolkit with 10 tools covering workspace inspection, file operations, search, and command execution. Tool definitions are consistently present with proper schemas and descriptions. However, parameter descriptions are largely absent, and error handling guidance is not evident in the source. The implementation follows a consistent pattern (Tool<InputType, OutputType, ContextType>) with explicit schema declaration via SchemaProviding protocol. Most tools are read-only with clear risk classification. The 'run_command' tool is marked IRREVERSIBLE but lacks confirmation/dry-run capabilities. Average tool description is 65-90 characters, which is adequate but could be more prescriptive about when/why to use each tool.
Replace an exact string in an existing file. The old string must match exactly once.
Return the current git diff for the workspace.
Find workspace files matching a glob pattern such as Sources/**/*.swift.
Search workspace text files by literal text or regular expression.
List text-oriented source files in the workspace.
Apply a sequence of exact string replacements to one existing file.
Read one UTF-8 text file inside the workspace.
Parameter descriptions completely absent across all 10 tools. Schema provides types but no guidance text explaining purpose, constraints, or valid values. LLMs cannot infer parameter semantics from names alone (e.g., does 'limit' mean max results, max file size, or max recursion depth?).
'run_command' is marked IRREVERSIBLE but provides no dry-run, confirmation step, or allowlist visibility. An LLM could invoke 'rm -rf /' without any guard. Agents need explicit confirmation patterns for destructive operations.
No error handling guidance visible in tool definitions. What happens if a file is not found? If edit_file's oldString doesn't match? If run_command times out? Error responses must tell the LLM what to do next (retryable, user-fixable, or fatal).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
Run an allowlisted verification command in the workspace.
Inspect the current workspace root, detected package type, and git status.
Create or overwrite a UTF-8 text file in the workspace.
Parameters with numeric bounds ('limit' accepted in list_files, grep, glob, run_command timeoutSeconds) lack min/max constraints in schema or description. ToolLimits.bounded() is applied at runtime, but LLMs see unbounded integers and may pass absurd values.
'grep' tool has optional boolean 'regex' parameter with no description of what it controls. Is it 'use regex or literal match'? Default is false (literal), but undocumented.
'multi_edit' tool description lacks detail on execution semantics. Does it apply edits sequentially? Atomically? What happens if edit N fails, do earlier edits persist? Critical for agent planning.
Tool output schemas (FileList, FileContent, SearchResults, etc.) are properly defined in CodingToolModels.swift, but descriptions of output fields are absent. LLMs need to know what 'truncated' means, what 'combinedOutput' contains in WorkspaceStatus, etc.