Intelligent Agent Platform Backend Service providing agent orchestration, tool management, graph workflows, and MCP server integration
Server exposes 3 tools with basic schemas and descriptions, but significant quality gaps prevent higher scores. Tool names follow verb conventions (tavily_search, think_tool, preview_skill), but descriptions lack depth and parameter documentation is sparse. No input validation guidance, error handling is minimal, and output schemas are not documented. The server appears to be a research/agent platform backend, but tool integration feels incomplete for production use.
Preview a skill generated in the sandbox. Reads all files and returns JSON with contents and validation.
Search the web for information on a given query. Uses Tavily to discover relevant URLs, then fetches and returns full webpage content as markdown.
Tool for strategic reflection on research progress and decision-making. Use this tool after each search to analyze results and plan next steps systematically.
No output schemas documented for any tool. Descriptions mention what is returned (e.g., 'returns JSON', 'returns markdown') but provide no structured definition of response fields. LLMs cannot plan downstream tool calls without knowing what fields are available.
Minimal error handling guidance. No tool describes what errors can occur, how to recover, or which failures are retryable. For tavily_search, what happens if the API fails? Is 'max_results' capped? For preview_skill, what if the skill file is missing or corrupt?
Parameter constraints are incomplete. tavily_search lacks min/max for max_results. preview_skill lacks validation rules for skill_name (is it alphanumeric? can it contain paths?). think_tool reflection parameter has no length guidance or structure hints.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 47 | - | v1 |
think_tool's semantic role is unclear. Is it a logging/no-op tool, or does it actively shape subsequent agent planning? The description doesn't clarify what the tool returns or how it influences agent behavior. This ambiguity will cause LLMs to misuse it.
tavily_search exposes implementation details (provider name 'Tavily'). If the backend switches providers, the tool name becomes misleading. Use a generic name like 'search_web' or 'search_internet' and hide the provider inside.
preview_skill has potential path traversal vulnerability. The 'skills_subdir' parameter defaults to 'skills' but is user-controlled. No validation is evident in the code snippet. An LLM could be tricked into reading /etc/passwd or other sensitive files.