A deterministic, in-path enforcement engine for AI agents with deny-by-default tool policy, Airlock RASP, reversibility ladder (R1-R3), budgets, and EU AI Act Art.50 transparency rule. Includes MCP tool servers for Git operations and test running.
FerrumDeck provides 13 Git and test-runner tools with basic schemas and descriptions, but exhibits significant quality gaps. All tools have names following verb_noun convention and descriptions present. However, schemas are minimal with sparse parameter descriptions, no output schema documentation, and missing error handling guidance. The server caps at STDIO transport (hard limit 50 for protocol readiness). While tool naming is clear (git_init, git_clone, run_tests), parameter descriptions are often single-line and generic. No evidence of idempotency markers, permission gates, or recovery guidance. The git_add tool accepts 'files' as an array but lacks detail on format expectations. Most parameters (path, remote, branch, framework, pattern, options) lack domain constraints (enums, ranges, format specifications). Output structure is never documented, callers cannot plan downstream chains. Error handling is absent from visible schemas. The test-runner's 'options' parameter is particularly concerning: described only as 'Framework-specific options' with no guidance on what keys/values are valid, forcing LLMs to guess.
Retrieve test execution results
Stage files for commit in a Git repository
List, create, or delete Git branches
Checkout a branch or commit in a Git repository
Clone a Git repository
Commit staged changes to a Git repository
Show differences between commits or working directory
Initialize a new Git repository
Get commit history of a Git repository
Output schemas completely undocumented. No tool provides structured return types, field descriptions, or chaining IDs. Callers cannot plan multi-step workflows (e.g., after git_log, what fields are available to pass to git_checkout?).
run_tests 'options' parameter documented only as 'Framework-specific options' with no enum, format, or example. LLM cannot determine valid keys/values without trial-and-error.
git_branch 'action' parameter lacks enum constraint. Description says 'Action: list, create, delete' but no formal enum {"enum": ["list", "create", "delete"]} in schema. LLM may guess invalid values like 'rename' or 'merge'.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 55 | - | v1 |
Pull changes from a remote Git repository
Push commits to a remote Git repository
Get the status of a Git repository
Run tests in a project
git_log 'limit' parameter has no min/max bounds. Unbounded integer invites LLM to pass 999999, potentially hanging the API or exhausting context.
No error handling guidance in any tool. No recovery hints (e.g., 'If branch not found, call git_branch to list available branches'). Raw error codes provide no actionable next step for the agent.
Destructive operations (git_commit, git_push, git_branch delete) lack idempotency markers or confirmation hints. Agents may retry and create duplicate commits or force-push unintended changes.
Parameter descriptions are minimal (15 - 30 chars). E.g., git_diff 'ref1' is 'First reference (commit/branch)', no hint on whether commit SHAs, branch names, or tags are accepted. Should specify format: 'A commit SHA (40 hex chars), branch name, or tag.'
git_add 'files' parameter is an array but no guidance on format: comma-separated? globs? relative to 'path'? Should document: 'Array of file paths relative to repo root, or glob patterns (*.py, src/**).'
run_tests 'format' parameter hints at enum ('json', 'xml', 'summary') but not formally constrained in schema. get_test_results uses same format hint with no constraint.
No idempotency or composition guarantees. E.g., calling git_commit twice with same message will create two commits (non-idempotent). Agents may retry on timeout and produce duplicates.