A multi-agent orchestration platform with A2A (Agent-to-Agent) protocol support, MCP integration, and web-based UI for managing agent servers
AgentWeave exhibits significant definition quality gaps across 18 tools. While all tools have names and brief descriptions, most descriptions are under 50 characters (well below the 194-char baseline for A+ tools). Input schemas are present but lack detail: many parameters use generic 'string' types without enums, format constraints, or valid value documentation. Parameter descriptions are minimal or missing context about expected formats, ranges, or dependencies. Error handling and recovery guidance is absent. The codebase shows tool definitions scattered across frontend (python-server/server.py) and backend (main.py, base_task_manager.py) modules with inconsistent schema patterns. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present. Output schemas are not documented. Several tools accept JSON-encoded strings as parameters (e.g., 'acceptedOutputModes' as JSON string) instead of proper typed objects, forcing LLMs to construct valid JSON manually and increasing error likelihood.
Accepts a JSON payload with agent details (agent_name, agent_description, agent_prompt, mcp_address, mcp_transport_type) and forwards it to the backend discovery server
Accepts agent configuration (agent_name, agent_description, agent_prompt, mcp_address, mcp_transport_type), saves to DB, starts the agent server, and returns the created configuration with assigned port
Creates a new task with agent details, message, and optional push notification configuration. Accepts form data with id, sessionId, acceptedOutputModes, message, agentName, agentCard, and pushNotification
Forwards agent deletion request to the backend discovery server by agent_name
Deletes an agent by agent_name, stops the agent server if running, and removes it from the database
Health check endpoint that forwards request to the backend discovery server at /health and returns status
Weak naming conventions: 8 tools use 'on_' prefix (on_get_task, on_cancel_task, on_send_task_subscribe, etc.), which violates verb_noun pattern and confuses LLMs. 'on_' suggests event handlers, not action verbs. Should rename to get_task, cancel_task, send_task_subscribe, etc.
Most descriptions are under 100 characters and lack WHEN/WHY guidance. Baseline for A+ tools is 194 chars. Examples: discovery_health (98 chars), on_cancel_task (99 chars), on_set_task_push_notification (74 chars). Descriptions must answer: What does it do? When to call instead of similar tool? What is returned?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 11 | - | v1 |
Fetches the server list from the backend API. Returns empty agents list instead of error if backend is unavailable
Returns agent cards for all backends in the server list by fetching from each running backend
Fetch the current LLM provider config and its fields from backend discovery server
Returns all agents from the DB with status field (running, timeout, connection refused, etc.) based on socket connection and MCP tool connectivity checks
Cancels a task by task ID. Returns TaskNotCancelableError if task cannot be cancelled
Retrieves task information by task ID with optional history length limiting
Retrieves push notification configuration for a task
Resubscribe to an existing task's streaming updates (not yet implemented)
Streaming endpoint for sending tasks with real-time status updates via Server-Sent Events
Sets push notification configuration for a task
Forwards refresh request to the backend discovery server for a specific agent, which checks and restarts the agent server if needed
Save the LLM provider config fields to backend discovery server
No output schemas documented for ANY tool. LLMs cannot infer what fields/types are returned, forcing them to guess and breaking downstream tool chains. Every tool must document its output structure (type, required fields, field types, descriptions).
Parameters accepting enums lack enum constraints. Examples: mcp_transport_type in create_agent (should enforce 'sse'|'stdio'), acceptedOutputModes in create_task (passed as JSON string instead of typed array). Free-form strings invite hallucinated values and force LLMs to guess valid options.
JSON-encoded string parameters reduce usability: acceptedOutputModes, message, agentCard, pushNotification in create_task are passed as JSON strings, not typed objects. Forces LLMs to construct JSON manually and increases syntax errors. Use proper typed parameters instead.
Destructive/write operations lack risk annotations and confirmation guidance. delete_agent, delete_agent_backend, on_cancel_task have DESTRUCTIVE/WRITE risk but descriptions do not warn of irreversibility or suggest confirmation patterns. LLMs cannot infer consequences.
Error handling absent: no guidance on recovery. Example: on_cancel_task mentions TaskNotCancelableError but does not explain when/why cancellation fails or suggest recovery steps. get_agent_cards and fetch_agents silently return empty lists on failure instead of signaling errors. LLMs cannot adapt.
Ambiguous tool overlap: get_agent_cards vs. fetch_agents vs. list_agents_with_status all retrieve agent lists but with unclear distinctions. LLM cannot determine which to call. Descriptions do not explain when each is appropriate.
Unimplemented tool exposed: on_resubscribe_to_task description states '(not yet implemented)'. This tool should not be available to LLMs. Remove or defer until implementation is complete.
Parameter validation rules undocumented. Examples: agent_name lacks format/length constraints, sessionId lacks description, historyLength lacks min/max bounds. Descriptions must state ranges, formats, and allowed values so LLMs can self-validate.