A multi-modal AI agent system with document processing, RAG (Retrieval-Augmented Generation), web scraping, and code execution capabilities. Provides MCP tools for document vectorization, context retrieval, URL requests, and JavaScript execution.
Omni-Agent exhibits significant quality gaps across tool definitions. While all 6 tools have names and descriptions, parameter documentation is incomplete, error handling lacks recovery guidance, and the javaScriptTool presents a critical security vulnerability. The server mixes read-only RAG/document tools with an irreversible code execution tool without clear composition patterns. Output schemas are not documented. Naming follows verb_noun conventions but several tools have overlapping responsibility (rag_search and traditional_rag both answer questions from documents; process_document, rag_search, and traditional_rag all accept document_url). Parameter descriptions are present but often lack format specifications, constraints, and examples. No parameters use enums for constrained inputs. Error messages are not shown in the sample code. The javaScriptTool allows arbitrary Node.js execution without sandboxing, confirmation, or rollback, a critical security and safety issue for agentic systems.
Execute JavaScript code and return the output. The code is executed with Node.js. Use this tool to solve computational problems by writing and running JavaScript code. Always call this tool when you need to generate code to solve a problem.
Download & vectorise a document once and return a document_id for later retrieval
Process a document from URL and retrieve relevant context/chunks based on questions. Returns the actual document content chunks rather than generated answers, allowing you to see what information is available in the document.
Retrieve relevant chunks from a previously processed document
Run a full non-agentic RAG pipeline on a document URL and return direct answers to questions.
Make GET requests to URLs and retrieve content. Simple tool for fetching data from web pages or APIs.
javaScriptTool allows arbitrary Node.js code execution without sandboxing, permission gating, confirmation step, or rollback capability. This is an irreversible destructive tool (can modify filesystem, environment, spawn processes) exposed directly to an LLM agent without safety guardrails. Critical security vulnerability.
Three tools (rag_search, traditional_rag, process_document) accept 'document_url' and return/answer questions about documents. No clear distinction of when to use each. Agents will waste reasoning cycles deciding between similar tools. Violates tool uniqueness principle.
Parameter 'k' (number of chunks to retrieve) in rag_search and retrieve_context lacks numeric bounds. No min/max constraints documented. LLMs can pass arbitrary values causing performance degradation or API overload.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | 2026-07-28+ | v2 |
| 2026-04-07 | F | 25 | - | v1 |
No output schema documentation for any tool. LLMs cannot predict what fields to expect (e.g., does rag_search return 'chunks' or 'results'? Does retrieve_context return 'document_id'?). Downstream tool chaining and field extraction are guesswork.
Error handling and recovery guidance are not visible in tool implementations. No indication of what errors can occur, whether they are retryable, or what the agent should do next (e.g., 'URL not found, try searching for the document first').
retrieve_context requires 'questions' parameter but does not document how to obtain the document_id it implicitly relies on. The tool description says 'retrieve from a previously processed document' but does not explain the prerequisite: process_document must be called first to get the document_id, and that ID must be passed somewhere (not visible in parameters).
url_request accepts a single 'url' string with no format validation, timeout specification, or redirect/security constraints (e.g., can it access internal/private IP ranges? What happens on infinite redirects?). Agents can be tricked into SSRF attacks.
use_ocr parameter in rag_search and use_cache in multiple tools have generic descriptions ('slower' extraction, 'cache if available'). No guidance on when these trade-offs matter or what they cost (time, memory, accuracy). Agents cannot make informed decisions.
No input validation rules documented for questions parameter (array of strings in rag_search, retrieve_context, traditional_rag). Can questions be empty? What is max length? What if an array contains 100 questions? Agents may pass invalid input.
javaScriptTool parameters 'filename' and 'description' are optional but their purpose is unclear. Does 'filename' control where the code is saved? Can it be exploited for path traversal? Is 'description' purely for logging or does it affect behavior? Ambiguity invites misuse.