A Tauri desktop application with integrated LLM capabilities, file operations, web browsing, and code analysis tools for agent-based workflows
Chaty provides 38 tools with basic schemas and descriptions, but exhibits significant quality gaps. Most tool descriptions are present but terse (avg ~80 chars, well below the 194-char baseline). Parameter descriptions exist but lack detail on constraints, formats, and valid ranges. No output schemas are documented. Error handling guidance is absent. Tools like `bash` and `bash_bg` expose dangerous operations without confirmation patterns. File operation tools (`edit_file`, `write_file`) lack atomic transaction semantics or rollback guidance. Browser automation tools lack idempotency markers. The codebase shows tool registration in `src/lib/toolRegistry.ts`, but actual schema definitions are not visible in the provided source, only tool names and brief descriptions are evident. This limits confidence in schema completeness.
when a decision is the user's (competing approaches, unclear requirements, destructive ops), ask a multiple-choice question — don't guess. args: { "question": string, "options": string[] }
run a shell command in the workspace (sandboxed, writes limited to the workspace); waits for exit — don't start dev servers with it. args: { "command": string, "timeout_secs"?: number }
start a long-running command in the background (dev server, slow build); returns an id, you're notified when it ends; sudo unsupported (use foreground bash). args: { "command": string }
type into a running background job and get its screen back: text types a line (Enter added), keys sends keys (down, enter, ctrl-c…). A bash command that stops to ask, or a REPL, moves to the background — answer it with this. args: { "id": number, "text"?: string, "keys"?: string[] }
kill a background job (whole process tree). args: { "id": number }
status + recent output of a background job. args: { "id": number }
Destructive operations (bash, write_file, browser_click, browser_eval) lack confirmation patterns or dry-run support. Agents can execute irreversible commands without safeguards.
Output schemas are not documented. Tools like search_code, web_search, and understand_repo return complex results, but LLMs cannot plan downstream calls without knowing the response structure.
Parameter descriptions lack constraint details. 'timeout_secs' in bash has no min/max; 'limit' in search_code has no bounds; 'pattern' in glob has no format guidance. LLMs cannot validate inputs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
click an element on the page
close the browser. args: {}
read the page's JS console output and exceptions. args: {}
evaluate JavaScript in the page context
send keyboard keys to the page
open a URL (or local file / dev server). Returns title + page text + interactive elements. args: { "url": string }
all visible text of the current page + element list (incl. current input values); a selector reads only its matches. args: { "selector"? }
reload the current page (cache ignored); required after local code edits — without it you see the stale page. args: {}
FULL-page shot (tall pages auto-segment; heavy). First look only — re-checks: browser_snapshot. args: {}
scroll to load more. args: { "to"?: "bottom"|"top", "by"?: number }
viewport shot (instant, light). PREFER for re-checks and after scrolling. args: {}
type text into a focused element
exact replacement (old_string must match the file verbatim and be unique unless replace_all=true). For several changes in one file, pass an atomic edits array in ONE call (any failure = nothing applied) instead of many small calls. args: { "path", "old_string", "new_string", "replace_all"? } or { "path", "edits": [{ "old_string", "new_string", "replace_all"? }] }
edit by line anchors (hidden from doc, anchor mode only)
find files by pattern (e.g. "src/**/*.ts"). args: { "pattern": string }
regex search over file contents. args: { "pattern": string, "path"?: string, "glob"?: string }
list one directory level (no path = workspace root). args: { "path"?: string }
atomic multi-edit (undocumented, executable)
read several files in one call — reach for it instead of firing read_file again and again. A file that cannot be read is reported on its own; the rest still come back. args: { "paths": string[] }
definition outline (functions/classes + line numbers) without reading the whole file. args: { "path": string }
read a file (pdf/docx/xlsx/pptx too — text auto-extracted, scans OCR'd). Pass symbol to get one function/class definition plus its call sites. args: { "path": string, "offset"?: number(1-based), "limit"?: number, "symbol"?: string }
ask the codebase by meaning; returns relevance-ranked files with their key definitions. First choice for unfamiliar code. args: { "query": "where login auth is handled", "k"?: number }
search the user's knowledge-base documents. args: { "query": string }
literal keyword search over file names + contents; names_only=true for names only. args: { "query": string, "path"?: string, "names_only"?: boolean }
one-call repo overview (README, manifest, tree, languages, entry points). First move in an unfamiliar workspace. args: {}
create/update the todo plan; mark finished steps done, the next one in_progress. args: { "todos": [{ "content": string, "status": "pending"|"in_progress"|"done" }] }
find and run just the tests related to the change; no args = everything changed this turn. args: { "files"?: string[] }
view an image in the workspace (vision models see it; others get OCR'd text). args: { "path": string }
background-download a file into the workspace; you're notified on completion — don't read it or re-request before that. args: { "url": string, "path": string }
fetch a URL, handled by content type (article→Markdown, GitHub file→raw source, video→transcript, PDF→text); raw=true for HTML. args: { "url": string, "raw"?: boolean }
web search; site scopes to one site (github.com returns structured repos/issues/code; reddit/youtube/bilibili/any domain). args: { "query": string, "site"?: string }
create a file, or rewrite one wholesale (replaces ALL content); to modify an existing file prefer edit_file. args: { "path": string, "content": string }
Browser automation tools (browser_click, browser_type, browser_key, browser_eval) have minimal descriptions (<40 chars). LLMs cannot determine when to use them or what they return.
Error handling is absent. Tools provide no recovery guidance. If bash fails, grep fails, or web_fetch times out, LLMs have no actionable next steps.