ag2 exposes 2 tools via HTTP MCP. Both tools lack critical definition quality. tool_search has a generic, unhelpful description ('Server-side tool search over a set of deferred tools') that does not explain what a 'deferred tool' is, when to use this vs alternatives, or what it returns. run_code has a better description template but relies on string interpolation ({languages}) that is not visible in the static source, making the actual constraint unknown. Neither tool has a visible input schema in the provided source, only run_code shows a partial schema hint in the metadata (code: string, language: string), but there is no evidence of enum constraints on 'language', parameter min/max bounds, or output schema documentation. tool_search has no visible parameter definitions at all. Error handling is not documented for either tool. The codebase appears to be a framework (ag2) rather than a standalone MCP server implementation, so tool definitions may be dynamically constructed or delegated to client code, which prevents full static analysis. This is a major blocker, if tool registration happens at runtime or via configuration not visible in the source files provided, the tools cannot be properly evaluated.
Execute code in a sandboxed environment. Supported languages: {languages}.
Server-side tool search over a set of deferred tools
tool_search has no visible input schema, parameter types, descriptions, and constraints are not present in source code.
tool_search description is vague and circular ('Server-side tool search over a set of deferred tools'). Does not explain what 'deferred tools' are, when to use this tool, what it returns, or how results differ from built-in tool discovery. Baseline for good descriptions is 50 - 200 chars with explicit WHAT/WHEN/WHAT-IT-RETURNS structure.
run_code description uses string interpolation placeholder {languages} which is not resolved in static source. Actual supported languages, their names, and how to specify them are hidden. LLM cannot infer valid language values from the description alone.
run_code parameter 'language' lacks enum or format constraint. Free-form string invites hallucinated values like 'javascript' or 'py' when 'python' is expected. Should declare supported values explicitly.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-04-07 | F | 42 | - | v1 |
No output schema documented for either tool. LLMs cannot plan downstream calls or extract required data (e.g., does run_code return stdout, stderr, exit_code, execution_time separately, or as a single string?).
run_code is a WRITE/destructive tool (executes arbitrary code) but has no error handling guidance, no confirmation/dry-run pattern, and no clear documentation of what happens on failure (e.g., partial execution, resource cleanup, retry safety).
run_code 'code' parameter has no description of format, constraints (max length?), or safety guidelines. Agents are unrestricted in what code they can submit, no bounds on execution time, memory, or I/O.