MCP stdio server exposing safe-ai-util operations as tools
This server exposes 27 tools covering git operations, filesystem access, and build/test runners. Tool definitions are present with schemas and risk classifications, but descriptions are terse and many lack parameter-level documentation. Most tools use concise verb_noun naming (git_status, fs_read, run_make) which is good. However, parameter descriptions are largely absent, the schema often shows only type and default, not what the parameter controls. Error handling is minimal; tools return RunResult with code/stdout/stderr but no structured guidance for LLM recovery. The descriptions themselves are brief (10 - 50 chars for most tools), below the 194-char baseline for production tools. Notably, sensitive operations like git_push, fs_write, and run_npm_ci lack detailed safety/confirmation guidance. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite the Risk field being populated in the metadata.
buf generate [--module <name>]
buf lint [--module <name>]
Test whether path exists (code=0 yes, code=1 no).
Match files against a glob pattern, one per line.
List directory entries non-recursively, one per line.
Read a file (full contents). Supports files up to 10 MiB by default; override with the optional max_bytes argument. Path validated against repo-root sandbox if SAFE_AI_UTIL_REPO_ROOT is set. There is NO 4 KiB cap — do not assume one. For files you only need a slice of, prefer fs_read_lines.
Read a 1-indexed inclusive line range from a file. Use this for large files when you only need a slice (e.g. just the function around line 240). Cheaper on context budget than fs_read of the whole file.
Most tool descriptions are below 20 - 40 characters; too terse for LLM selection logic. e.g., git_push='git push (optionally --set-upstream origin <branch>)', git_rebase='git rebase <onto>'. These descriptions do not explain WHAT the tool does, WHEN to use it, or consequences.
No parameter-level descriptions for any tool. Schema shows type and default but not what parameter controls. e.g., git_add.pattern just says type=string, default='.', with no explanation. LLMs cannot infer semantic meaning from names alone, 'target' in git_diff could mean branch, tag, commit, or working directory.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 8 | - | v1 |
Write content to path (via stdin so payload is never on the cmdline).
git add <pattern> (default '.')
List branches (no name) or create a new branch (name). main/master rejected.
Switch to <branch>
git commit -m <message>
Show working-tree diff (optionally against <target>)
Show recent commits, oneline, capped at <max_count>
git push (optionally --set-upstream origin <branch>)
git rebase <onto>
Show git status
Install dependencies via pip inside venv
Run pytest via safe-ai-util
Create or reuse a Python venv at path (default .venv)
Signal whether the task is complete, partial, or blocked. The harness uses this to decide whether to open the PR as ready or draft, what labels to apply, and whether to skip PR creation entirely. Call this exactly once at the end of your loop, BEFORE you stop calling tools. status must be one of: 'complete' (work is done, ready for review), 'partial' (some progress, more needed), 'blocked' (could not proceed; explain in reason).
Run go build <package> (default ./...)
Run go test <package> (default ./...)
Run go vet <package> (default ./...)
Run make <target>
Run npm ci
Run npm test
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite metadata labeling tools as READ_ONLY or WRITE. The MCP spec supports structured hints to help LLMs reason about side effects. Absence means LLMs cannot distinguish safe operations (git_status, fs_read) from destructive ones (git_push, fs_write) without parsing descriptions.
No documented error handling or recovery guidance. Tools return RunResult(code, stdout, stderr) but schema/description does not tell LLMs how to interpret codes or what to do on failure. e.g., git_branch explicitly rejects 'main'/'master' with code=2, but LLM does not know why or that it should retry with a different name.
No output schema documentation. Tools return text via stdout/stderr but descriptions do not specify what structure or format to expect. e.g., git_log claims 'recent commits, oneline, capped at <max_count>' but does not document whether output is newline-delimited, JSON, or raw text. LLMs must guess.
Destructive tools (git_push, git_rebase, fs_write, run_npm_ci, run_go_build) lack confirmation or dry-run guidance.
fs_write and py_pip_install accept unbounded/large inputs (content string, requirements string) with no documented max size or validation. LLMs could be tricked into passing massive payloads causing memory exhaustion or timeouts.