NagaAgent has 19 tools with visible schemas and descriptions, but quality is inconsistent. Naming is generally verb-noun style, which is good. However, most descriptions are in Chinese, brief, and lack explicit guidance on when/why to use each tool. Many parameters have types but inconsistent descriptions. Error handling is not documented in schemas. Security concerns (exec, process, write tools) lack explicit mitigation guidance in descriptions. Schemas are present but parameter descriptions vary widely in quality. The average tool scores around 50-55, pulling the overall to a fair/poor range.
游戏攻略查询引擎,支持多游戏、截图识别、历史上下文
游戏攻略查询(强制启用自动截图)。
浏览器自动化操作。启动、导航、截图、交互等。
伤害计算(强制使用计算模式)。
定时任务管理。
编辑文件(查找替换)。
执行shell命令。运行系统命令并返回输出。
查找文件。
Critical security concern: exec and process tools expose arbitrary shell command execution with parameters workdir, timeout, background, pty. No descriptions warn about injection risks, no validation guidance, no permission-gating hints. Descriptions do not mention that these are IRREVERSIBLE and require extreme caution.
Destructive write tools (write, edit, cron) lack explicit error handling and recovery guidance. No descriptions state that these modify state or warn about retry semantics. Missing idempotency hints, agents may retry failed writes, creating duplicates or partial state corruption.
Parameter descriptions are sparse or missing across 16+ tools. Examples: exec has 'workdir' and 'timeout' parameters with minimal context on valid ranges; browser tool action parameter lacks enum documentation of valid actions (status/start/stop/open/snapshot/screenshot/navigate/act); process action parameter similarly undocumented.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
阵容推荐(强制使用完整模式和自动截图)。
搜索文件内容。
分析图片内容。
列出目录内容。
读取记忆片段。
语义搜索记忆。
管理exec进程。查看、轮询、终止进程。
读取文件内容。
获取网页内容。抓取指定URL的页面内容。
联网搜索。搜索实时信息、新闻、价格、天气等。
写入文件内容。
Lack of output schema documentation. Schemas define inputs but do not document return types. Agents cannot plan downstream tool chaining or extract relevant fields. Example: web_search description says '结果数量(1-10)' but does not specify what fields are returned (e.g., does it return title, url, snippet, rank?).
Tool naming suffers from ambiguity: ask_guide, ask_guide_with_screenshot, calculate_damage, get_team_recommendation are specialized gaming tools that should be clearly distinguished in names but currently blend into a generic ask_* family. No prefix distinguishes game query tools from core tools.
Descriptions are in Chinese, limiting accessibility for non-Chinese-speaking teams and complicating localization. English descriptions are needed for international tooling compatibility.
Multiple similar tools without clear distinction: ls vs find (both discover files); memory_search vs memory_get (both retrieve from memory). Descriptions do not explain when to prefer one over the other, forcing LLMs to guess and try both.