Portable AI agent SDK and CLI you install and point at problems. Opinionated agent SDK extending AI SDK's ToolLoopAgent with built-in tools for file operations, shell execution, web search, code analysis, and more.
agntk provides 21 tools with basic schemas and descriptions, but significant quality gaps prevent a higher score. Most tools have descriptions (90% coverage) and input schemas, but many descriptions are generic or lack actionable context. Parameter descriptions are inconsistent, some tools like ast_grep_search have detailed param docs, while others like progress_read and recall have minimal guidance. No output schemas are documented, making downstream chaining difficult. Error handling is not evident in the provided schema definitions. Naming is generally good (verb-noun pattern), but composition issues exist: progress_read and progress_update operate on shared state without clear conflict prevention; remember/recall/queryKnowledge form a knowledge system but lack integration guidance. The browser tool accepts 13 different actions via a free-form 'action' parameter rather than using enum constraints, inviting hallucination. Overall, the toolkit resembles a competent developer tool library (good for human use) but lacks LLM-optimization (clear chaining, error recovery, output schemas).
Replace code patterns using AST-aware rewriting. Dry-run by default. Use meta-variables in rewrite to preserve matched content. Example: pattern="console.log($MSG)" rewrite="logger.info($MSG)"
Search code patterns using AST-aware matching. Supports 25 languages. Use meta-variables: $VAR (single node), $$$ (multiple nodes). Patterns must be complete AST nodes (valid code). Examples: "console.log($MSG)", "def $FUNC($$$):", "async function $NAME($$$)"
Run shell commands in the background without waiting for completion
Browser automation tool for web interactions. Requires agent-browser CLI. Supports actions: open, snapshot, click, dblclick, fill, type, select, press, hover, scroll, screenshot, getText, getUrl, getTitle, wait, eval, check, uncheck, close
Structured chain-of-thought reasoning tool for complex problem-solving
Extract entities and relationships from a piece of text. Use this to analyze conversations, documents, or any text for structured knowledge. Returns entities (people, projects, goals, problems, etc.) and their relationships.
No output schemas documented for any tool. LLMs cannot plan downstream chaining or validate response structure. For example, file_read, glob, grep, and web_search all return unspecified result types, forcing LLMs to infer what fields exist and blocking composition.
browser tool uses a free-form 'action' string parameter (e.g. 'open', 'click', 'fill', 'select', 'screenshot', 'eval') instead of an enum. This invites hallucinated action names. Should be: action: {type: 'string', enum: ['open', 'snapshot', 'click', 'dblclick', 'fill', 'type', 'select', 'press', 'hover', 'scroll', 'screenshot', 'getText', 'getUrl', 'getTitle', 'wait', 'eval', 'check', 'uncheck', 'close']}
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 56 | <=2025-11-25 | v2 |
Create a new file. Fails if the file already exists. Creates parent directories automatically. Use file_write if you want to overwrite existing files.
Edit a file using context-based search and replace. Provide the exact text to find (oldText) and the replacement text (newText). Uses surrounding code as anchors — no line numbers needed. The oldText must match exactly (including whitespace) for the edit to succeed.
Read a file from the workspace. Supports optional line range (startLine/endLine) for reading specific sections. Returns the file content with line numbers.
Write content to a file. Creates parent directories automatically. Overwrites the file if it already exists. Use file_create if you want to fail on existing files.
Search for files using glob patterns
Search file contents using regex patterns
Create and manage multi-step action plans
Read the current progress state
Update and track progress of operations
Search the knowledge graph for code entities (functions, classes, files) or facts. Use this to find information about the codebase or recalled facts. Returns matching entities with their file paths and types.
Retrieve past memories, facts, or observations related to a query. Use this to remember past decisions, context, or learnings. Returns episodes with their content and extracted entities.
Store a fact or observation in memory for later recall. Use this to save important information, decisions, or learnings. The fact will be automatically analyzed to extract entities and relationships.
Multi-step search skills combining multiple search strategies for complex queries
Execute shell commands in the workspace
Search the web for information using multiple providers with fallback chain (DuckDuckGo, SearXNG, Tavily)
progress_read has no input parameters and minimal description ('Read the current progress state'). This is too vague, what format is the progress state? What fields are returned? No output schema. Critically, progress_read and progress_update form a shared-state system with no conflict prevention or coordination model documented.
search_skills has a generic description ('Multi-step search skills combining multiple search strategies for complex queries') that does not explain what it returns, when to use it instead of web_search, or how it differs from other search tools. Parameter schema also lacks detail on what 'query' should contain.
Knowledge graph tools (queryKnowledge, remember, recall, extractEntities) form a four-tool subsystem but lack integration guidance. No documentation of how remember and recall interact, whether recall searches the entire history or a sliding window, or how extractEntities output feeds into remember. This creates composition ambiguity.
shell and background tools accept arbitrary shell commands with no validation, error handling, or guidance on recovery. No mention of environment injection, command injection risk, or what output format to expect. Error handling is missing, what happens if a command times out or fails?
plan tool accepts 'steps' as an array but does not specify the structure of each step (is it a string description, an object with tool/params, or something else?). No output schema explains how the plan is executed or tracked.
deep_reasoning tool description does not explain whether the LLM should call this tool, when to invoke it, or what it does with the thoughts. Parameter names ('thoughtNumber', 'totalThoughts') suggest it expects structured reasoning, but the contract is unclear.
Error handling strategy is not evident. Tools like file_read, file_write, grep, glob have no documented error responses. What happens if a file is not found, permission is denied, or a glob pattern matches nothing? LLMs need actionable guidance.
ast_grep_search and ast_grep_replace accept optional 'globs' parameter (include/exclude patterns) but the description is terse ('Include/exclude globs (prefix ! to exclude)'). Should provide examples: ['src/**/*.ts', '!src/**/*.test.ts']