Fastest way to build, prototype and deploy AI Agents with MCP (Model Context Protocol) support
The Agentor MCP server has CRITICAL quality issues that prevent production recommendation. Most significantly: (1) THREE DUPLICATE TOOL DEFINITIONS with identical or near-identical names ('get_weather' appears 3 times, 'greeting' defined once). This violates the composition pattern and will confuse LLMs about which tool to call. (2) Tool descriptions are generic and too short, most under 50 characters, many under 30. (3) Input schemas visible but minimal, most contain only 1-2 parameters with basic type coverage. No output schemas documented anywhere. (4) No evidence of parameter validation descriptions, format constraints, enums, or error guidance. (5) Duplicate tools suggest the server code was not properly reviewed before exposure. Evidence: tool 1 (get_weather, 38 chars), tool 2 (tool_search, 39 chars), tool 3 (get_weather, 43 chars, DUPLICATE), tool 4 (greeting, 33 chars), tool 5 (get_settings, 28 chars), tool 6 (get_weather, 57 chars, DUPLICATE), tool 7 (get_time, 40 chars), tool 8 (code_review, 33 chars). The code excerpt shows a tool decorator system (BaseTool.from_function) and schema generation logic, but tool registration itself is not visible in the provided source, suggesting inferred definitions.
Generate a code review prompt
Get application settings
Get current time - no context needed
Get current weather information
Get current weather for a location
Get current weather information with user context
Generate a personalized greeting
Search for a tool based on a query.
THREE DUPLICATE TOOL DEFINITIONS: 'get_weather' defined three times (instances 1, 2, 6) with slightly different descriptions. LLMs will not know which to call and cannot distinguish between them. This violates the composition pattern requiring one tool per concern and clear naming distinction.
MOST TOOL DESCRIPTIONS UNDER 50 CHARACTERS: Rubric baseline is 194 chars average (p10=34, p90=392). Current descriptions: get_weather 38-57 chars, tool_search 39 chars, greeting 33 chars, get_settings 28 chars, get_time 40 chars, code_review 33 chars. These are too generic for LLM tool selection. Per pattern:tool-description, descriptions must be 10-1024 chars AND answer WHAT, WHEN, and what it returns.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
NO OUTPUT SCHEMAS DOCUMENTED: Input schemas are visible (with 'type', 'properties', 'required'), but NO output schema or return type documentation exists for ANY tool. Rubric baseline: 100% of A+ tools have documented return types. LLMs cannot plan downstream calls or extract correct fields without knowing what each tool returns.
MISSING PARAMETER DESCRIPTIONS: Tool inputs show JSON Schema with properties but no descriptions for many parameters. E.g., 'score_threshold' in tool_search lacks guidance on valid range (0-1? 0-100?). 'style' in greeting has a hint ('formal or casual') but no description of what each produces. Per pattern:tool-description, EVERY parameter must have a non-empty description explaining what it controls.
NO ENUM CONSTRAINTS WHERE NEEDED: 'style' parameter in greeting accepts 'formal or casual' but is typed as a free-form string, not an enum. Per pattern:constrained-input, known sets of values must be enums to prevent hallucination. Same issue with any status/type/mode parameters across the tool suite.
NO ERROR HANDLING OR RECOVERY GUIDANCE: Tool definitions and code excerpt provide no evidence of error responses, failure modes, or guidance for LLM recovery. Per pattern:recovery-guide, error responses must tell the LLM what to do next (e.g., 'User not found. Try search_users() with a partial name.'). No evidence this is implemented.
TOOL DEFINITIONS NOT EXPLICITLY VISIBLE IN SOURCE: The provided source excerpt shows tool decorator logic (BaseTool.from_function) and schema generation (build_schema function), but actual tool registration and invocation paths are not visible. Tools are inferred from the specification metadata provided, not directly observed in the code. Per hard scoring rules, this caps per-tool scores at 50.
VAGUE NAMING: 'tool_search' is self-referential and unclear, does it search for tools by name, capability, or category? 'get_settings' with a 'uri' parameter is ambiguous, is it fetching app settings, user settings, or endpoint-specific configuration? Per pattern:tool, tool names must convey action + clear object. Recommend 'search_available_tools' and 'get_app_settings' or 'get_resource_settings'.