A LiveKit voice agent that integrates MCP (Model Context Protocol) tools and A2A (Agent-to-Agent) servers for dynamic tool execution in voice interactions
This server exposes two dynamically-generated tools that lack explicit, verifiable schema definitions and clear separation of concerns. Tool 1 (dynamic_mcp_tools) is a meta-tool that wraps external MCP server tools, its actual parameters and output schema are not defined in this codebase, making it impossible to score the individual tool quality. Tool 2 (a2a_skill_tools) has a minimal schema with only a 'prompt' parameter and a generic string description. Neither tool has documented output schemas, error handling guidance, or parameter validation rules. The code shows attempt at schema conversion and parameter building (lines in mcp_client/agent_tools.py), but the actual tool registration is incomplete: tools are inferred from external servers rather than explicitly defined. This violates the principle that tool definitions must be self-contained and verifiable. The descriptions are too generic to guide LLM tool selection ('Dynamic tools fetched...' and 'Skills from A2A servers...' do not explain WHEN or WHY to use them). No evidence of security checks, permission gates, or audit logging beyond basic Python logging.
Dynamically loaded tools from connected MCP servers via prepare_dynamic_tools
No explicit tool registration: tools are dynamically fetched from external MCP servers rather than defined with full schemas in this server. Tool definitions are inferred at runtime, making them impossible to validate, document, or version. LLMs cannot reliably understand tool semantics when schemas are generated on-the-fly.
Missing output schemas for both tools. No documentation of what fields are returned, their types, or how downstream tools should reference returned data. Agents cannot plan multi-step sequences when output schemas are unknown.
Generic, non-actionable tool descriptions. Both descriptions are meta-explanations of what the tool wrapper does, not what the tool accomplishes for users. Descriptions lack WHEN-to-use guidance, prerequisites, or return type hints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 13 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 19 | - | v1 |
Tool names do not start with action verbs. 'dynamic_mcp_tools' and 'a2a_skill_tools' are nouns/descriptors, not verb_noun patterns (e.g. 'invoke_mcp_tool', 'run_skill'). LLMs cannot infer intent from names alone. Baseline: 90% of A+ tools start with action verbs (get, list, create, search, update).
No error handling guidance. Code catches exceptions and logs them (lines in mcp_client/agent_tools.py) but does not return structured error messages to the agent. When a tool fails, the agent receives no hint about whether to retry, what went wrong, or what to try next.
No input validation or constraint declaration. The 'prompt' parameter in a2a_skill_tools has no min/max length, no format pattern, no enum, LLMs can pass arbitrarily long or malformed prompts.
No security or permission checks visible. Code does not verify caller permissions, does not log audit trails of tool invocations, and does not inject secrets server-side. Dynamic tools from MCP servers may expose sensitive operations without guardrails.
Unclear composition with MCP servers. The server wraps external MCP tools dynamically, but does not document which MCP servers it connects to, what tools they provide, or how naming collisions are resolved. Users cannot predict what tools will be available.