Gryphon demonstrates solid definition quality with consistent naming, comprehensive descriptions, and well-structured schemas across most tools. All 11 tools follow verb_noun naming patterns (list_, search_, get_, execute_, run_, submit_, cancel_). Descriptions are substantive (100-200 chars on average) and explain WHAT, WHEN, and implications. Input schemas are complete with type definitions and parameter descriptions. However, output schemas are not documented in the visible code, and error handling guidance is minimal, critical gaps for production LLM integration. The server operates as a code execution sandbox and API discovery layer, which constrains some composition patterns but is architecturally sound.
Cancel an executing or queued run by its receipt handle.
Run restricted Python with inputs, result, and await call_tool("server.function", args).
Inspect real parameter, request-body, and response metadata for 1–5 functions.
Poll a persistent run receipt by handle, observing completion or intermediate state.
Discover cached, reusable code recipes grouped by input schema.
Discover servers in the operator-configured compact or function-summary mode.
Read a full or streamed JSON artifact by owned handle without projection.
Output schemas not documented in visible source code. LLMs cannot infer what fields to expect from tool responses (e.g., what does list_servers return? What are the structure of page cursors, server metadata, pagination fields?). This forces agents to guess and plan conservatively.
Error handling and recovery guidance minimal. Descriptions of execute_code, run_cached_code, transform_artifact state WHAT they do but not HOW to recover from failures (e.g., what errors can occur? Should the agent retry, ask the user, or abort?). No actionable error classification.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 67 | <=2025-11-25 | v2 |
Reuse unchanged cached code with a new structured inputs object.
Search the compiled registry before requesting full function schemas.
Submit execution and return a persistent receipt, not an MCP Tasks promise.
Reduce stored JSON offline, without refetching or reading chunks.
transform_artifact description is vague ('Reduce stored JSON offline, without refetching or reading chunks'). Does it filter? Aggregate? Project? LLMs need concrete use-case guidance. Also, 'no imports or call_tool capability' is a constraint, not a purpose, should be in error handling.
execute_code, run_cached_code, and submit_code accept 'inputs' (object) and optional 'input_schema' (object) but do not specify what constitutes valid input. The descriptions mention 'dynamic values available as inputs' and 'Structured execution inputs' but lack concrete examples of structure, constraints, or validation rules. An LLM does not know if inputs should be flat key-value or nested.
Confirmation/dry-run pattern absent for irreversible operations. execute_code, run_cached_code, transform_artifact all modify stored state (execute compiled code, cache recipes, transform artifacts) but offer no dry-run, preview, or confirmation step. Agents cannot safely experiment.
idempotency_key parameter in execute_code, run_cached_code, and submit_code is described as 'Optional owner-scoped deduplication key, not write approval.' This is confusing, does it prevent duplicate execution, or just tag calls for tracking? LLMs need to know if retrying with the same idempotency_key is safe.