A Tauri desktop application for managing AI IDE adapters, agents, presets, and debate/optimization workflows with integrated analytics and registry management
AgentHarbor exposes 3 read-only filesystem tools with clear, action-verb names (read_file, list_directory, grep) and well-structured JSON schemas. All tools have descriptions (194 - 250 chars, within baseline 34 - 392 range) and documented input parameters with types. Schemas include proper constraints (max_results 1 - 200, additionalProperties: false). However, output schemas are not documented, LLMs cannot predict the structure of grep results or list_directory responses. Error handling is minimal; no recovery guidance or actionable error messages visible. Tool composition is sound (each does one thing), but the lack of output documentation and error patterns prevents a higher score.
Case-insensitive literal substring search across the project. Walks up to 6 levels deep, skipping common build/cache folders and files over 1 MiB. Returns matching line objects {file, line, text}.
List the entries (files and subdirectories) inside a directory of the project. Pass an empty path to list the project root. Returns up to 500 entries, with directories first. Common build/cache folders (.git, node_modules, target, dist, build, .next, .venv, __pycache__) are skipped.
Read a UTF-8 file from inside the project. Use to verify what a file actually contains before claiming it does or doesn't. The path is relative to the project root. Output is capped at 64 KiB; long files are truncated with a marker.
Output schemas not documented. LLMs cannot predict the structure of grep results {file, line, text} or list_directory responses. Without documented return types, agents must infer structure from trial and error.
No error handling guidance. Tools return Result<String, String> but no documentation on error messages, recovery steps, or how LLMs should interpret failures. E.g., what happens if a path is outside the project root?
read_file truncation behavior undocumented in schema. Description mentions '64 KiB cap with marker' but output schema does not indicate whether a 'truncated' flag or marker string is returned. LLMs cannot reliably detect truncation.
grep max_results parameter lacks clear guidance on default behavior. Description says 'Default 50, hard cap 200' but schema does not show default value in JSON Schema. LLMs may not know what happens if omitted.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
No pagination or result limiting for list_directory beyond 500 entries. If a directory has 1000+ files, the tool silently truncates. No next_cursor or indication of truncation in documented output.