Authorization, delegation, provenance, and verifiable-audit engine for AI agents
Four tools with explicit schemas and descriptions, but significant gaps in parameter documentation and output schema clarity. Tool names follow verb_noun convention (read_file, write_file, run_shell_command, get_runtime_context). Descriptions exist but are inconsistent in depth, write_file and run_shell_command include implementation caveats ('DEMO/REFERENCE') that clutter the LLM-facing description. No parameter descriptions for get_runtime_context (empty properties). Output schemas are not documented anywhere in the source. Error handling and recovery guidance are absent. Security-critical tools (write_file, run_shell_command) lack clear permission/scope declarations.
Get current working directory, git repository root, active branch, and workspace status details.
Read a local file through agent-sudo policy enforcement.
DEMO/REFERENCE executor: classifies and gates the command through agent-sudo, then executes only a narrow allowlist. It is not a general shell. To gate real commands, embed the agent-sudo authorization engine in your agent (see README).
DEMO/REFERENCE executor: classifies and gates the write through agent-sudo, then writes inside the configured workspace (defaults to /tmp/agent-sudo-demo if no workspace is configured). It does not write to arbitrary paths outside the workspace.
get_runtime_context has empty inputSchema properties with no parameter descriptions. LLM cannot infer what fields are returned or how to use the output.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract required fields (e.g., does read_file return {content, size, encoding} or just {content}?).
write_file and run_shell_command descriptions include implementation details ('DEMO/REFERENCE executor', 'narrow allowlist') that distract from the LLM's decision to call the tool. Descriptions should focus on WHAT and WHEN, not HOW.
No error handling or recovery guidance. If read_file fails (permission denied, file not found), the LLM has no guidance on what to do next.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Security-critical tools (write_file, run_shell_command) lack explicit permission/scope declarations. No indication of what authorization checks are performed or what the LLM should know about risk.