The fleet MCP server defines 33 tools with generally clear naming and descriptions. Most tools follow verb_noun patterns (fleet_audit_run, fleet_git_status, fleet_deps_scan). Descriptions are present and substantive (average ~150-200 chars), explaining what the tool does and when to use it. However, there are significant gaps: (1) Input schemas are visible in the code but many lack detailed type constraints, parameters are typed as 'string' or 'number' without min/max bounds or format constraints. (2) No output schemas are documented anywhere in the provided source. (3) Tool parameters lack descriptions in several tools (e.g., fleet_logs has 'container' param but no description visible in definition). (4) Error handling is mentioned in descriptions but no structured error recovery guidance is provided. (5) No pagination guidance for list operations (fleet_git_pr_list, fleet_runner_list). Overall: above-average naming and descriptions, but incomplete schemas and missing output documentation hold the score to low-60s.
Look up Apple App Store Review Guidelines via greenlight. action "list" returns all sections, "show" returns one section (query is a section number like "2.1"), "search" matches a keyword (query is the term). Use to interpret a finding's guideline reference or to guide a fix.
Suppress a confirmed greenlight false positive from future audits. The finding is matched by its exact title, optionally narrowed to a target and to findings whose file or code contains a substring. Every rule must carry a reason. Suppressed findings are dropped and the pass/fail summary is recomputed on subsequent fleet_audit_run calls.
Run an App Store compliance audit on a mobile app project via greenlight preflight. Scans source code, the privacy manifest, and metadata for Apple App Store rejection risks. Target is a registered fleet app name or a path to a mobile project directory. Returns the full report (findings grouped by CRITICAL/WARN/INFO plus a pass/fail summary).
Show the most recent App Store audit results from cache without re-running a scan. Returns cached audit records (summary plus findings). Use as a cheap first pass before fleet_audit_run.
Dependency findings for a specific app
Output schemas are not documented. No tool lists or describes the structure of its response. LLMs cannot plan downstream tool calls or extract the right data without knowing what fields they'll receive.
Numeric parameters lack bounds and constraints. E.g., fleet_logs 'lines' defaults to 100 with no documented min/max. Fleet_logs_recent 'lines' defaults to 50, 'sinceMinutes' defaults to 15, 'maxResults' defaults to 20, no stated upper bounds. Unbounded numbers let LLMs pass absurd values.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Get or set dependency monitoring configuration
Create a PR with dependency updates for an app (dry-run by default)
Add an ignore rule for a dependency finding
Run a fresh dependency scan across all registered apps
Dependency health summary from cache — outdated packages, CVEs, EOL warnings, Docker image updates
Create a feature branch from develop (or other base) and push it
Stage tracked file changes and commit
Onboard an app to GitHub: create repo, push code, protect branches
Create a pull request on GitHub
List pull requests for an app
Push current branch to origin
Create a release PR from develop to main
Git state for one or all apps: branch, clean/dirty, onboard status
DEPRECATED — prefer fleet_logs_recent (token-conservative defaults) or fleet_logs_summary. Get recent container logs for an app.
Get recent log lines for an app, filtered to a level and bounded in size. Defaults are SMALL (50 lines, last 15 minutes, warn+) — broaden only if needed. Returns {text, truncated, suggestion}.
Bounded grep across recent container logs. Returns matching lines with 0 lines of context, capped at max_results. Cheaper than fleet_logs_recent + manual filtering.
Cheap aggregate: counts of log lines by level + the top 10 distinct error/warning messages over a window. Use as a first pass before fleet_logs_recent.
List the registered remote build hosts.
Register or update a remote build host that fleet remote runner tasks target over ssh. Stored in the runner registry (FLEET_RUNNERS_FILE, else ~/.local/share/fleet/runners.json).
Remove a registered remote build host.
Doctor a registered remote build host over ssh: reachability plus a toolchain/disk preflight (os, node, full Xcode, free disk). Surfaces a host that cannot build before a run starts.
Detect drift between vault (encrypted, survives reboot) and runtime (/run/fleet-secrets/, lost on reboot). Shows which keys were added, removed, or changed at runtime but NOT sealed back to the vault. If drift is detected, use fleet_secrets_seal to persist changes, or fleet_secrets_unseal to revert runtime to vault state.
Get a single decrypted secret value from the vault. Returns the value stored in the encrypted vault, NOT the runtime value. Use fleet_secrets_drift to check if runtime differs from vault. NOTE: under the privilege-separated daemon this is the deny-by-default "secret" tier — the operator must opt it in via mcp-policy.json.
Restore vault from backup (.bak file). Backups are created automatically before any seal operation. Use this if a seal operation produced incorrect results and you want to revert to the previous vault state.
Seal runtime secrets back to the encrypted vault. CRITICAL: If you modified environment variables at runtime (e.g. edited .env files in /run/fleet-secrets/), those changes will be LOST on reboot unless you seal them back to the vault with this tool. This re-encrypts the current runtime state into the vault so it persists across reboots.
Set a single secret key/value for an app. IMPORTANT: This updates the encrypted vault directly. The change persists across reboots. If the app is running, you may also need to update the runtime env and restart the app.
List an app's TestFlight builds via the App Store Connect API — build number, version, processing state and expiry. Requires the app's ASC credentials and ASC_APP_ID in its fleet secrets.
Check TestFlight publishing readiness for an app: GitHub CLI availability, the GitHub repo backing the build workflow, App Store Connect credentials, and — when ASC_APP_ID is set — that the ASC API is reachable.
List operations (fleet_git_pr_list, fleet_runner_list) provide no pagination support. No limit/offset, no cursor, no total count. Returning unbounded lists risks context window exhaustion. Missing: accept limit/offset, return total/next_cursor.
No error recovery guidance. Descriptions mention 'returns' but don't explain what happens on failure (credential missing, network error, permission denied) or what the LLM should do next. No recovery guide pattern.
Destructive operations lack confirmation or dry-run safeguards. fleet_runner_remove (DESTRUCTIVE risk), fleet_git_branch/commit/push/pr_create/release (WRITE), fleet_secrets_set/seal/restore (WRITE) are missing confirmation patterns. LLMs can delete runners or lose secrets without reversibility.
Some tools lack explicit parameter descriptions in the schema. E.g., fleet_logs 'container' param has no description visible. While the tool description mentions containers, individual parameters should be self-documenting.
fleet_logs is marked DEPRECATED in favor of fleet_logs_recent and fleet_logs_summary. Presence of deprecated tool alongside newer variants signals API churn. Remove or clearly mark as obsolete.
No tool annotations present. No readOnlyHint, destructiveHint, or idempotentHint. These are optional but recommended for LLM reasoning, especially for tools with WRITE, DESTRUCTIVE, or REVERSIBLE risk levels.
dryRun parameter pattern is used throughout git tools (fleet_git_branch, fleet_git_commit, fleet_git_push, fleet_git_pr_create, fleet_git_onboard, fleet_git_release) but is not explicitly documented as a confirmation pattern in the server description. The parameter itself lacks a description in some cases explaining that it 'previews without making changes'.