Gryphon — governed API access and bounded Python execution for AI agents
Gryphon Runtime demonstrates solid definition quality with 11 well-named tools, consistent parameter documentation, and clear schema definitions. Tool names follow verb_noun pattern and are action-oriented (list_*, search_*, execute_*, get_*, cancel_*). Most parameters include type definitions and descriptions. However, there are gaps: some parameter descriptions lack detail about constraints and valid ranges, output schemas are referenced but not fully documented inline, and error handling guidance is minimal. The server shows clear separation of concerns across discovery, execution, and artifact management tools. Parameter documentation is above average but could be more prescriptive about constraints, formats, and dependencies.
Request cancellation of a pending or running execution.
Run restricted Python with inputs, result, and await call_tool("server.function", args).
Inspect real parameter, request-body, and response metadata for 1–5 functions.
Poll a persistent run receipt by handle.
List cached code recipes available for reuse.
Discover servers in the operator-configured compact or function-summary mode.
Output schemas not documented in tool definitions. While input schemas are well-defined, return types and output structures are not explicitly documented. LLMs cannot plan downstream tool calls or extract required fields without knowing what execute_code, run_cached_code, and submit_code return.
Parameter descriptions lack constraint details. Parameters like 'limit' (accepted range 1-100) and 'code' (what imports/network access are forbidden) have minimal constraint documentation. The descriptions state 'capped by discovery_limit' but don't explain what that limit is or why it exists.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 79 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 32 | - | v1 |
Read a stored JSON artifact by owned handle.
Reuse unchanged cached code with a new structured inputs object.
Search the compiled registry before requesting full function schemas.
Submit execution and return a persistent receipt, not an MCP Tasks promise.
Reduce stored JSON offline, without refetching or reading chunks.
Execution tools (execute_code, run_cached_code, transform_artifact) lack error recovery guidance. Descriptions do not explain what errors might occur, what they mean, or what the LLM should do next (retry, ask user, abandon). No indication of timeout behavior or partial failure modes.
Tool descriptions for execution operations do not explicitly state they have side effects. 'execute_code' and 'submit_code' are marked WRITE/WRITE risk but descriptions don't clearly say 'This creates persistent state' or 'This is non-reversible'. Agents need explicit language to understand retry behavior.
Pagination parameters present but guidance incomplete. Tools like list_servers accept cursor/limit but descriptions don't explain how to iterate (follow next_cursor until it's absent?) or what happens if limit exceeds discovery_limit (silently capped? error?).
Cached code tool (run_cached_code) description does not explain what 'cache_id' is, where it comes from, or how to discover valid cache IDs. Parameter description says 'Recipe identifier returned by execute_code or list_recipes' but doesn't explain the relationship between execution and recipe caching.
Confirmation/dry-run pattern absent for destructive tools. cancel_run can terminate running operations but there's no dry-run, confirmation step, or reversal mechanism documented.