Multi-Agent Scaling System - A powerful framework for collaborative AI with MCP server support for custom and MCP tools
MassGen exposes 6 background job management tools with severely limited schema documentation and generic descriptions. All tool names follow the verb_resource pattern (get_, list_, cancel_, wait_, start_), which is positive. However, descriptions are generic and lack LLM-optimized context ('Get status of a background tool execution job' is 54 chars but provides no guidance on when to call it vs alternatives). Input schemas are present for all tools, but parameter descriptions are minimal ('Job identifier to query' lacks format, constraints, or expected patterns). No output schemas are documented anywhere in the visible code, LLMs cannot infer what fields to expect from response objects. The risk classification (WRITE vs READ_ONLY) is a positive security signal, but tool definitions lack idempotency guarantees, error handling guidance, or integration context. The 'arguments' parameter in start_background_tool is a bare object with no schema validation, inviting hallucinated or invalid tool invocations.
Cancel a running background tool execution
Get result of a completed background tool execution
Get status of a background tool execution job
List all background tool execution jobs
Start a background tool execution job
Wait for a background tool execution to complete
No output schemas documented for any tool. LLMs cannot infer response structure, field types, or downstream chaining data.
Generic, uninformative descriptions (35-40 chars) lacking LLM-optimized context. 'Get status of a background tool execution job' tells the LLM what, but not when/why to use it vs alternatives or what fields are returned.
'arguments' parameter in start_background_tool is untyped object with no validation schema. LLMs cannot infer valid keys/types, inviting hallucinated argument payloads.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Parameter descriptions lack format, constraints, or context. 'Job identifier to query' does not specify: format (UUID, alphanumeric?), length limits, regex patterns, or how to obtain a valid job_id.
No timeout_seconds constraint documentation. timeout_seconds accepts 'number' but provides no min/max bounds, unit clarity, or behavior description (what happens at timeout: exception vs graceful return?).
No idempotency guarantees documented. If an agent retries get_background_tool_status due to network failure, is the call safe? This is critical for agent resilience but not addressed.
No error handling guidance. What error does start_background_tool return if tool_name is invalid? What about invalid arguments? How should the LLM recover?