GraphOS runtime for governed AI agents, orchestration, and the epistemic-graph knowledge engine. Exposes MCP tools for quant trading, fleet tool discovery, orchestration, and governance.
The agent-utilities MCP server exposes a single tool 'quant' with significant definition quality issues. While the tool has a description and input schema, the description is vague and domain-specific marketing language rather than LLM-actionable guidance. The input schema provides types and defaults but lacks critical descriptions for most parameters, making it difficult for LLMs to understand when and how to use each parameter. The schema is overloaded with 13 parameters covering four distinct domains (orchestrate, data, execute, portfolio), violating the single-responsibility principle. Parameter naming uses inconsistent conventions (some snake_case, some generic like 'side', 'mode', 'period'). No output schema is documented, so LLMs cannot plan downstream tool calls or understand what fields to expect. Error handling is not visible in the provided source. The tool appears to be a catch-all gateway rather than a proper composition of focused, single-purpose tools.
The Ultimate Quant System Tool. Domains: - 'orchestrate': Intelligence layer (Actions: debate, analyze, regime, ensemble_predict) - 'data': Telemetry layer (Actions: historical, order_book, fundamentals) - 'execute': Trading layer (Actions: submit_order, cancel_order, status). SAFEGUARD: Defaults to mode="paper". - 'portfolio': Risk layer (Actions: balances, positions, risk_metrics, optimize)
Tool name 'quant' is not verb-based and does not convey action. LLMs cannot infer what calling this tool does from the name alone. Should be split into verb-based tools like 'get_market_data', 'submit_order', 'analyze_portfolio'.
Tool description is marketing copy ('The Ultimate Quant System Tool') rather than LLM-actionable guidance. It lists domains and safeguards but does not explain WHEN to call this tool, WHAT problem it solves, or WHAT the agent should expect back. Descriptions should be 10-1024 chars and answer: What does it do? When use it? What does it return?
Most input parameters lack descriptions. 'ticker' has a description, but 'domain', 'action', 'side', 'quantity', 'order_type', 'price', 'mode', 'portfolio_id', 'rounds' either have missing or insufficient descriptions. Per the rubric, 100% of A+ tool params must have descriptions. LLMs cannot infer that 'side' means buy/sell without explicit text.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 42 | 2025-06-18+ | v2 |
| 2026-03-09 | C | 61 | - | v1 |
Single tool combines four distinct domains (orchestrate, data, execute, portfolio) with overlapping, context-dependent parameters. This violates single-responsibility principle. Parameter 'action' is free-form string (debate, analyze, regime, ensemble_predict, historical, order_book, fundamentals, submit_order, cancel_order, status, balances, positions, risk_metrics, optimize), no enum constraint. LLMs will hallucinate invalid actions.
Action parameter is a free-form string without enum constraint. Valid values are hard-coded in description but not enforceable via schema. LLMs frequently hallucinate values outside the documented set, leading to failures.
No output schema documented. LLMs cannot plan what fields to expect or how to chain this tool's output into downstream calls. Without documented return structure, agents cannot build multi-step workflows reliably.
Parameter dependencies not documented. When 'domain=execute', the 'action' parameter accepts 'submit_order', 'cancel_order', 'status', but when 'domain=orchestrate', it accepts 'debate', 'analyze', etc. LLMs must infer these dependencies from description text, and will often pass invalid combinations.
Parameter 'mode' defaults to 'paper' (safe), but 'quantity' and 'price' default to 0, which may trigger unintended behavior if omitted. Defaults must not cause data loss or unintended side effects. If 0 is invalid for orders, either require the param or document the 0-default behavior clearly.
No validation rules stated for numeric parameters. 'quantity' and 'price' accept any number type with no min/max bounds. LLMs can pass negative quantities, zero prices, or astronomically large orders, all of which should be rejected early with actionable error messages.
Tool is marked WRITE-level risk but no error handling guidance is visible. When a live order submission fails (insufficient funds, invalid symbol, market closed), how does the tool communicate this to the LLM? Are errors retryable? Should the agent ask the user? This guidance is critical for financial tools.