CLI tool for interacting with Ollama as a cybersecurity agent with tool calling
This server has 17 tools with basic naming and descriptions, but exhibits significant quality gaps. Tool names follow verb_noun convention (execute_command, get_file_contents, write_file, etc.) which is good. However, descriptions are minimal (most under 100 chars), input schemas are present but lack depth, and error handling is minimal. The server is missing output schema documentation, pagination support, and security safeguards for irreversible operations. Most tools score in the 45-55 range, placing this firmly in the 'poor to fair' category per calibration baselines.
Analyze and rate the security risk level of findings.
Copy a file from source to destination.
Create a directory at the specified path.
Create a security checklist for a specific domain or system.
Delete a directory and optionally its contents.
Delete a file.
Execute a shell command and return the output.
No output schema documentation. Tools return unstructured strings; LLMs cannot predict field structure for downstream tool chaining. E.g., 'get_file_contents' returns plain text, not a structured object with {filepath, lines, content, start_line, end_line}.
No error classification or recovery guidance. Errors return generic messages like 'Error: File not found' with no suggestion for next steps. Per pattern, errors should be categorized as retryable, user-fixable, or fatal, and include recovery hints.
Destructive tools (delete_file, delete_directory, execute_command) lack confirmation or dry-run support. No confirmation_request pattern for irreversible operations, risking accidental data loss when LLM makes a planning error.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Find files matching a glob pattern.
Get remediation steps for a CVE vulnerability.
Read and return the contents of a file.
Get information about a file (size, modification time, type, etc.).
Get git diff for uncommitted changes or between commits.
Get git commit log.
Get git status of a repository.
Search for text in files (grep-like functionality).
List files and directories in a path.
Write content to a file.
No pagination or result limits explicitly documented. Tools like 'find_files' and 'list_directory' accept max_results but descriptions do not warn LLMs that results are truncated or offer pagination guidance.
Parameter descriptions are incomplete. Many parameters lack format constraints, ranges, or examples. E.g., 'execute_command' timeout parameter has no guidance on valid range (0-999? 1-3600?).
No input validation or sanitization visible. 'execute_command' shells out with user input; no evidence of command injection protection. 'grep_search' uses user regex without validation. Per security pattern, all agent input must be treated as untrusted.
Security domain tools (analyze_risk_level, get_cve_remediation, create_security_checklist) have trivial descriptions (45-50 chars). Unclear when to use them vs. each other, or what format they expect. 'analyze_risk_level' takes a free-form 'description' param with no constraints.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in schema. Clients cannot infer risk level from metadata; they must parse descriptions. Per current spec (2026-07-28), tool annotations are encouraged for UX and safety.
Parameter naming inconsistencies. 'grep_search' uses 'is_regex' and 'case_sensitive' booleans; 'create_security_checklist' uses 'domain' (too vague).