Production-grade autonomous MCP server: persistent memory, cognitive control, terminal & browser verification
Gravitas-Core-MCP has 21 tools with explicit schemas and descriptions registered in server.py. However, critical quality issues severely limit this score: (1) Parameter descriptions are largely absent, most parameters lack any explanation of what they control or how to use them. (2) Output schemas are completely undocumented, no tool declares what it returns or what fields agents should expect. (3) Many tool names use verb_controller or verb_action_type patterns that are less crisp than standard verb_noun (e.g., 'controller_transition' vs 'transition_task'; 'browser_navigate' is good). (4) Several tool descriptions are generic or vague, e.g., 'Load task and its context for resumption' doesn't explain WHEN to call it or WHAT it returns. (5) Error handling and recovery guidance is absent from all tool descriptions. (6) No tool indicates whether it's idempotent, retryable, or has side effects. (7) Tools like terminal_execute and terminal_start_background are HIGH-RISK (irreversible, destructive) but lack confirmation/dry-run patterns or explicit warnings in descriptions. (8) Browser and terminal tools expose low-level APIs (screenshot paths, cwd) that should be abstracted into safer, higher-level operations. The schema definitions themselves are present and well-formed (JSON Schema with types), but the semantic quality and LLM-readiness are below production grade.
Return collected JS console errors since last navigate.
Hover over an element by CSS selector.
Navigate to URL (Playwright).
Take screenshot; optional path to save file.
Capture DOM accessibility tree and console errors.
Create a new task and set state to PLANNING.
Return current task state and policy info.
Output schemas completely undocumented. No tool declares its return type, fields, or data structure. Agents cannot plan downstream tool calls or extract relevant data without trial-and-error exploration.
Parameter descriptions are missing or minimal across all tools. Most parameters (e.g., 'new_state', 'cwd', 'url', 'path', 'max_depth', 'snapshot_id') have no explanation of what they control, valid values, or constraints. LLMs cannot infer semantics from names alone.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 48 | 2026-07-28+ | v2 |
Record a step failure; may trigger FAILED_RETRY or ROLLBACK.
Transition task to a new state (PLANNING, CODING, EXECUTING, VERIFYING, FAILED_RETRY, ROLLBACK, COMPLETED).
Return the last verified, immutable working state for rollback/recovery.
Return the last known state (most recent snapshot + active task). Authoritative memory.
Generate Model Resume Package for model swap/editor restart/crash recovery: goal, task, constraints, failures, safe/do-not-touch files.
Save a context snapshot for current task (internal use).
Set the canonical (immutable) state to a snapshot (after verification).
Recursive project structure with noise filtering.
Record a failed strategy/command to prevent repetition.
Load task and its context for resumption (model handover/restart).
Execute a shell command with timeout and cwd. Returns stdout, stderr, exit_code.
List active background process ids.
Start a background process; use process_id to stop later.
Terminate a background process by process_id.
Destructive operations lack warnings and safety mechanisms. terminal_execute, terminal_start_background, browser_screenshot, and memory tools can modify state (execute commands, overwrite files, take snapshots) but lack dry-run, confirmation-request, or explicit 'this is irreversible' warnings in descriptions.
No error handling or recovery guidance. Tool descriptions do not indicate what can fail, how to detect failure, or what the LLM should do next (retry, ask user, try alternative tool). Agents will be blind when errors occur.
Idempotency and retryability not declared. Tools like controller_transition, memory_save_snapshot, and memory_set_canonical modify state but do not indicate whether they are safe to call multiple times with the same arguments. Agents will not know whether to retry on transient failures.
Naming inconsistency and clarity. Some tools use controller_ or terminal_ or browser_ prefixes alongside method names, creating longer, less crisp names. E.g., 'controller_transition' could be 'transition_task'; 'controller_record_step_failure' could be 'record_step_failure'. Prefixes add little value and reduce clarity.
Tool descriptions are generic and lack context about WHEN to use each tool vs. similar alternatives. E.g., 'Resume task and its context for resumption' is vague, does resume_task also validate? Does it restore files? When do you call it instead of get_last_state or get_canonical_state?
Parameter types not enumerated where they should be. 'new_state' in controller_transition accepts a string but no enum of valid states is provided, even though the description lists them: PLANNING, CODING, EXECUTING, VERIFYING, FAILED_RETRY, ROLLBACK, COMPLETED. Enums are self-documenting and prevent hallucinated values.
Low-level APIs expose unsafe parameters. terminal_execute, browser_navigate, and browser_screenshot accept cwd, url, and path parameters that require careful handling (path traversal, command injection). Descriptions do not mention sanitization, validation, or safe patterns.