Local security operations and ContextCypher architecture context, driven by people and AI assistants
Guardian Agent exposes 36 tools with bare-minimum definitions. All tools have names and brief descriptions, but critically lack structured schemas, parameter type definitions, output documentation, and error handling guidance. The tools span filesystem operations, shell execution, web/browser automation, email, forum posting, automation management, threat intelligence, package installation, and code execution, an extremely broad surface area with minimal safety constraints visible. Most descriptions are under 50 characters and provide no context on when to use each tool, what they return, or how to handle failures. No tool shows enumerated constraints, validation guidance, or idempotent/destructive hints. The source code confirms tools are registered (visible in scripts/test-tool-contracts.mjs and src/tools/*.ts) but their schemas are not materialized in the provided code excerpt. This suggests definitions exist at runtime but are not visible for inspection, per the scoring rubric, inferred definitions cap individual tool scores at 50. The overall score reflects that while tools exist and are named, they lack the rigor needed for safe, reliable LLM invocation.
Delete an automation
Run an automation
Save an automation
Enable or disable an automation
Perform action in browser
Extract data from browser page
Interact with browser page elements
Extract links from browser page
Navigate browser to a URL
No visible input schemas for any of the 36 tools. The tool contract definitions appear to exist only at runtime (inferred from tool names and file references in test scripts), but JSON Schema definitions are not materialized in the provided source code. This violates the core assumption that schemas should be statically inspectable and documented.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 47 | <=2025-11-25 | v2 |
Read browser page content
Get browser page state
Build code
Lint code
Execute code in remote sandbox
Run code tests
Create a document
Post to a forum
Copy a file or directory
Delete a file or directory
List files and directories in a path
Create a directory
Move or rename a file or directory
Read file contents
Search for files matching a pattern
Write file contents
Draft a Gmail message
Send a Gmail message
Call Google Workspace Service
Draft an action for a security finding
Add a watch target for threat intelligence
Remove a watch target
Install a software package
Execute a safe shell command
Get system information
Fetch content from a URL
Search the web
Descriptions are universally terse (20 - 50 characters) and lack context on WHEN to use each tool, WHAT it returns, and HOW to handle errors. For example, 'Read file contents' tells an LLM only the immediate action, not whether it streams large files, returns a byte count, truncates at a limit, or how to handle permission errors.
Parameter descriptions are minimal or missing detail. 'File path to read' says WHAT but not FORMAT (relative vs absolute? size limits? character set?). 'Search pattern' is ambiguous, is it regex, glob, SQL LIKE? Does shell_safe accept environment variables? What is the 'action' parameter in browser_act, click, scroll, type, submit? LLMs cannot infer; they need explicit constraints in parameter descriptions.
No enum constraints visible on any parameters. Multiple tools accept 'type' or 'kind' parameters (e.g., automation_save, intel_draft_action) with no enumerated options. Free-form strings invite hallucinated values; LLMs should pick from a predefined set. This is a critical control for correctness.
No documented output schemas. LLMs cannot plan chained calls if they do not know what fields a tool returns. For example, does fs_list return file sizes, timestamps, and paths? Does web_search return ranking scores? Does gmail_send return a message ID for later reference? Without this information, agents cannot compose multi-step workflows.
No error recovery guidance. Tools like fs_delete (DESTRUCTIVE), gmail_send (WRITE), and package_install (WRITE) offer no error classification or recovery hints. If fs_delete fails, should the agent retry? Ask the user? Does gmail_send support dry-run? Without guidance, agents either retry blindly (risking duplicate side effects) or give up.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in the schema definitions. These are essential for LLMs to understand safety implications. fs_delete and automation_delete are marked DESTRUCTIVE in metadata but likely lack the MCP-level annotation that allows clients to warn or gate the call.
Tool names lack specificity in some cases. 'gws' (Google Workspace Service) is cryptic and provides no verb hint; 'Call Google Workspace Service' is too generic. Does it send email? Manage calendars? Update contacts? LLMs struggle to decide if this is the right tool for a task. Should be split into gws_send_email, gws_create_event, etc., or at least renamed to reflect the most common action.
No visible pagination or limits on list-returning tools. fs_list, browser_extract, web_search likely return unbounded results, risking context window exhaustion. There is no documented limit, page parameter, or offset/cursor mechanism. A single fs_list on a large directory could return thousands of entries, drowning the LLM context.
Parameter types are inferred (string, number, boolean) from the brief schema hints, but no minLength, maxLength, pattern, or min/max for numeric fields are visible. For example, fs_read has a 'maxBytes' parameter, but is 0 valid? Is there a hard system limit? 1GB? The description says nothing. Similarly, 'shell_safe' accepts a command string with no validation constraints visible.
Security-critical tools (fs_delete, shell_safe, package_install, code_remote_exec, gmail_send) lack visible permission gates or audit logging hints. No description mentions which scopes are required (e.g., 'write:email' for gmail_send). If agents are meant to run with least-privilege tokens, there is no guidance on how to configure that.