ToolAnything demonstrates moderate structure with consistent naming patterns and moderate parameter schemas, but suffers from critical gaps in parameter documentation, output schema definitions, and error handling. All tools have descriptions, but most are brief and lack actionable context for LLM selection. Parameter descriptions are largely absent or minimal. No visible output schema documentation for any tool. The server supports diverse tool types (function, HTTP, model, SQL) but lacks the depth of specification expected for production agent tooling. Average tool definition is incomplete relative to the 54 Agentic Tool Patterns baseline.
連線診斷用 ping 工具,回傳固定結果
示範 SQL source tool
VAD 前置過濾示範(ONNX)
VAD 前置過濾示範(PyTorch)
加總兩個整數
依名稱回傳問候語
計算兩個整數的總和
打招呼,回傳客製化問候語
Missing parameter descriptions across all tools. Parameters like 'name', 'a', 'b', 'text', 'user_id', 'include', 'conf', 'iou', 'max_det', 'model', 'device', 'root_id', 'relative_path', 'content', 'expected_sha256' are exposed with types but no descriptions explaining their purpose, valid ranges, or expected format.
No output schema documentation visible for any tool. Tools return results but the structure, field types, and purpose of each output field are not documented. LLMs cannot plan downstream calls or extract required data (e.g., returned IDs for chaining).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
寫入一行備忘(示意 side_effect=True)
Write file content with optional atomic checksumming
反轉字串
示範 HTTP source tool
使用 YOLOv8 偵測影像中的人,回傳 bounding boxes 與信心分數
Generic and incomplete tool descriptions lack actionable context. Descriptions like '打招呼,回傳客製化問候語' (Chinese), '計算兩個整數的總和', '反轉字串' do not explain WHEN to use the tool vs alternatives, prerequisites, or what fields the agent should expect back. Many descriptions are under 50 characters.
Destructive tool (standard.fs.write) lacks confirmation/dry-run mechanism. WRITE-risk tools should support preview or confirmation step to prevent irreversible accidents. No visible rollback or compensation tooling.
No error classification or recovery guidance visible. Tools do not document what errors are retryable, user-fixable, or fatal. Error responses lack guidance on next steps (e.g., 'User not found. Try search_users() first').
Tool naming inconsistency: '__ping__' violates verb-first convention and uses dunder naming (typically reserved for internal methods). Should be 'ping' or 'check_connection'. 'users.fetch' is valid but 'hello' lacks a verb; should be 'get_greeting' or 'send_greeting'.
No tool annotations visible (readOnlyHint, destructiveHint, idempotentHint). The WRITE-risk tools (standard.fs.write, quickstart.store_note) and READ_ONLY tools lack explicit capability hints that would guide agents on sequencing and retryability.
Parameter naming inconsistency: 'vad_features' is a float array but no enum or pattern constraint is visible. The array is constrained to shape [6] in input_spec but this is not exposed in JSON Schema format that LLMs reliably parse from descriptions.