A Flask-based orchestrator that manages agent registration, search routing, and LLM-powered query synthesis across multiple agents with category classification.
Oqtopus exhibits significant definition quality gaps typical of early-stage MCP servers. While 12 tools are defined with basic schemas and descriptions, the descriptions lack actionable detail, parameter constraints are sparse, output schemas are not documented, and error handling guidance is missing. Tool naming follows some conventions (verb_noun) but several tools have unclear responsibilities. The server mixes authentication, content management, and agent registry operations without clear separation of concerns. Schemas are present but incomplete, parameters often lack granular constraints (enums for categories exist, but many string fields accept open-ended input). No tool has documented return types, pagination patterns, or error recovery guidance. Security concerns exist (email/password as parameters, no visible input sanitization docs). This server would benefit from LLM-optimized descriptions (50 - 200 chars), comprehensive parameter annotations, and structured error responses.
Render the contact form page.
Edit an existing agent's configuration including name, URL, categories, description, and visibility.
Display the main index page with registered agents, domains, and request quota information.
Handle user login with email and password validation. Rate limited to 50 per minute.
Log out the current authenticated user.
Display list of agents owned by the authenticated user.
Register a new user with username, email, and password validation. Rate limited to 10 per minute.
Missing output/return schemas for all 12 tools. LLMs cannot infer downstream field names, types, or pagination structure. No tool documents what fields are returned, their types, or references needed for chaining (e.g., agent_id, user_id). This prevents agents from composing tools and wastes round-trips on discovery calls.
Vague tool descriptions under 100 characters lack LLM selection context. 'Render the contact form page' (contact_page), 'Display the main index page' (index), 'Serve the sitemap.xml file' (sitemap), 'Serve the robots.txt file' (robots) do not explain WHEN to call them, WHAT they return, or their purpose in agent workflows. Baselines show A+ tools average 194 chars; these range 24 - 80 chars.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 42 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 49 | - | v1 |
Register a new agent with URL, categories, name, orchestrator ID, description, and visibility settings. Includes HTTPS enforcement and SSRF protection.
Serve the robots.txt file.
Submit a search query that is classified by category using LLM, matched against registered agents, fetched asynchronously, and synthesized into a final answer. Implements rate limiting (5 requests/24h for authenticated users, 1 request/24h for guests).
Handle contact form submission and send email via SMTP with validation.
Serve the sitemap.xml file.
Generic tool names reduce clarity. 'index' does not start with an action verb and conflicts with HTTP conventions, does it query, render, or return metadata? 'my_agents' is possessive and unclear, List? Get? Better: 'list_my_agents' or 'get_user_agents'. Pattern baselines: 90% of A+ tools start with an action verb.
Input parameters lack granular constraints and descriptions in several tools. Example: login/register accept 'email' and 'password' with only type:string, no format, min/max length, or validation rules documented. Edit_agent accepts 'categories' array with type:string enum, but edit_agent's category items lack the enum constraint that register_agent provides. Inconsistent and vague.
No error handling or recovery guidance visible in any tool description. Tools like 'login' offer no context on what happens on auth failure, should the LLM retry, ask for different credentials, or suggest password recovery? Pattern:recovery-guide requires errors to tell the agent what to do next.
Credentials (email, password) appear as tool parameters in login/register. This violates pattern:secret-injection, secrets logged in agent traces expose user credentials. Implement server-side session handling or OAuth instead.
Tools combining multiple concerns signal poor composition. 'register_agent' handles URL validation, SSRF protection, and metadata storage in one call, good. But 'index' displays agents, domains, AND quota in one tool, should be split. 'search' classifies via LLM, matches agents, fetches asynchronously, AND synthesizes results, four responsibilities. Pattern:tool requires one clear job per tool.
No documented idempotency or retry semantics. Tools like 'send_contact_email', 'register_agent', and 'register' are WRITE operations (irreversible), agents need to know if it's safe to retry on timeout or if retry risks duplicates. Pattern:idempotent-operation.
Missing rate-limit error context. 'register' and 'login' mention rate limits in descriptions ('50 per minute', '10 per minute') but do not document what error the agent receives when limits are hit or how to recover. Agents need actionable guidance, not just a statistic.
Pagination and result limits not documented for 'index' (which returns agents, domains, quota), 'my_agents', and 'search'. If 'index' returns 1000 agents, LLMs cannot reason over all of them. Pattern:paginated-result requires limit/offset/cursor and total counts.