A LangGraph-based agentic assistant with sandboxed filesystem and web search tools connected via MCP, exposed through a Streamlit chat UI.
Four tools with basic schemas and descriptions, but significant gaps in parameter documentation and error guidance. All tools have descriptions (10-50 chars baseline met), and input schemas are present with type definitions. However, parameter descriptions are minimal or absent, output schemas are undocumented, and error handling lacks recovery guidance. The filesystem tools show good security (path traversal prevention) but weak LLM-facing documentation. Web search tool lacks result pagination and field documentation. Average per-tool score: 58.
List files and folders inside the sandbox (optionally inside a subdirectory of it). Returns a newline-separated list of relative paths.
Read and return the text content of a file inside the sandbox.
Search the live web for a query and return a short list of results (title, url, and a snippet) formatted as text.
Write text content to a file inside the sandbox, creating parent directories as needed. Overwrites the file if it already exists.
Parameter descriptions are minimal. 'subdirectory' in list_files and 'relative_path' in read_file/write_file lack format constraints, examples, or validation rules. LLMs cannot infer whether paths accept '..' escapes, symlinks, or absolute paths.
Output schemas are undocumented. Tools return plain strings with no structured field definitions. LLMs cannot plan downstream operations or extract specific data (e.g., file count, error type). Violates pattern:response-shaper.
Error responses lack recovery guidance. write_file returns 'Wrote N characters' on success but provides no structured error classification (retryable vs fatal). read_file returns 'No such file' or 'binary file' with no suggestion to call list_files first.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | 2026-07-28+ | v2 |
web_search lacks pagination and result limits. max_results defaults to 5 but tool description does not state this limit or explain pagination strategy. No documented output schema for result fields (title, url, content).
write_file description does not warn that it overwrites existing files. Agents need to know which operations are irreversible. Missing idempotency guidance, calling write_file twice with same content is safe, but LLM cannot infer this.