Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
Automagik Hive exposes 12 tools with critically insufficient definition quality. While tool names follow action-verb conventions (python_executor, web_search, file_reader), the implementations lack essential schema documentation, parameter descriptions, and output specifications. Source code inspection reveals that tools are registered in hive/config/builtin_tools.py, but input schemas are minimal, most contain only basic string/integer/boolean types without constraints, enums, ranges, or detailed descriptions. Tool descriptions themselves are brief (5 - 20 words) and lack context about when to use each tool, what it modifies, dependencies, or error recovery. Most critically, there is NO EVIDENCE of documented output schemas, pagination, or error handling guidance. The rubric requires that tools returning lists offer pagination and limits; none do. For a server exposing destructive tools (python_executor, shell_tools, sql_query, delete operations), absence of confirmation patterns, dry-run modes, or permission checks is a critical gap. Security-critical parameters (API keys, credentials) appear to be handled externally, but the tool definitions do not declare what permissions each requires, blocking least-privilege verification.
Add documented output schemas for every tool. Include field names, types, and constraints. For tools returning objects, show the full structure. For lists, specify max items, pagination fields, and total count.
Expand tool descriptions to 50 - 150 characters following the pattern: 'What it does, when to call it, what it returns.' Example: 'Execute Python code safely in an isolated sandbox. Use for data processing, math, or testing logic. Returns output, any errors, and execution time.'
Add detailed parameter descriptions with format/constraint rules. For 'file_path': 'Path to file to read (must be under /data/, no ../ traversal, max 255 chars, supports .txt, .csv, .json, .yaml).' For enums, replace prose with a declared enum list.
Declare input constraints (min/max for numbers, regex/length for strings, enum lists). Replace 'optional' with defaults; state 'defaults to 20' not 'optional max_results'.
Add permission declarations to every tool. Specify what scopes are required (e.g., 'Requires: read:email' or 'Requires: write:github.repo'). Gate destructive tools behind explicit permission checks.
For tools modifying state (python_executor, sql_query, shell_tools, github_api, slack_api, email_tools), add a 'dry_run' boolean parameter (default true) so agents can preview side effects before committing.
Restructure github_api and slack_api away from a generic 'action' string parameter. Instead, expose separate tools: create_github_pr, list_github_issues, send_slack_message, etc. Each tool does one thing clearly.
Score history
Overall score trend
↑ 14 points across a rubric change (v1 → v2)
42/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
42
2026-07-28+
v2
2026-03-09
F
28
-
v1
source verified
48/100
Send Slack messages and notifications
sql_querywritesource verified39/100
Execute SQL queries safely
web_scraperread onlysource verified45/100
Scrape web pages and extract content
web_searchread onlysource verified53/100
Search the web using DuckDuckGo
youtube_toolsread onlysource verified44/100
Search and analyze YouTube videos
Parameter descriptions are absent or trivial (under 20 chars); LLMs cannot infer when/how to use parameters correctly
Parameters lack enums, ranges, and format constraints; free-form strings invite hallucinated/invalid values (e.g., github_api 'action' param has no enum list)
No input validation rules documented; parameters like 'file_path' have no restrictions against path traversal, no format requirements, no length limits
file_readercsv_toolsshell_tools
Add error handling guidance: 'If file not found, try list_files() to discover available paths.' 'If SQL fails, check syntax and verify table names.' 'If rate limited, retry in 60 seconds.'
For tools accepting lists or returning variable-length results (web_search, csv_tools), add pagination: 'max_results defaults to 20 (max 100). Returns total count and next_cursor for pagination.' Enforce result caps in implementation.
Add examples of valid input values as constraints, not prose. E.g., github_api 'action' should be an enum: ['list_issues', 'create_pr', 'close_issue', 'add_comment']. Do not hide valid options in text.
Document idempotency for each tool. State 'Safe to retry, no duplicate side effects' or 'Not idempotent, verify before retrying' so agents know whether to retry on ambiguous failures.
For tools with dependencies (e.g., 'call file_reader first to get file contents before csv_tools'), add hints in descriptions: 'If you only have a filename, call list_files() first to resolve the full path.'
Add input validation and sanitization inside tool implementations. Reject path traversal in file_path parameters, validate SQL syntax, sanitize shell commands. Return clear errors: 'Invalid: path traversal detected. Use file names only, not ../../sensitive.txt'.