An agent framework and development platform with MCP server integration, chat capabilities, and plugin system for building agents with custom tools and workflows
Selene exhibits pervasive definition quality deficiencies across all 19 tools. While tool names follow basic verb-noun conventions and descriptions exist, they are severely underdeveloped. Most descriptions lack actionable detail about when to use each tool, what it returns, prerequisites, and error recovery paths. Input schemas are present but parameter descriptions are minimal or absent entirely. Output schemas are not documented. Parameter constraints (enums, ranges, formats) are largely missing. No evidence of idempotency guarantees, error categorization, or LLM-friendly recovery guidance. The tool definitions appear to be skeletal wrappers around internal functions rather than LLM-optimized interfaces. For example, 'editFile' has a description 'Edit specific sections of a file' but does not explain the format of the 'edits' array parameter, what modifications are supported, or how failures are reported. 'executeCommand' accepts 'args' as an array with description 'Command arguments' but provides no guidance on escaping, injection prevention, or timeout behavior. 'skill' is entirely generic ('Execute a registered skill or plugin capability') with no indication of how to discover available skills or what parameters they expect. These gaps force LLMs to reason from minimal information, increasing error rates and wasted tool calls.
Prompt the user with multiple choice or free-form questions
Execute bash commands with stdin support and output streaming
Perform mathematical calculations and evaluate expressions
Compress and compact the current chat session
Interact with the design workspace for UI/UX design
Edit specific sections of a file
Execute shell commands with streaming output and subprocess management
Search files locally using ripgrep for pattern matching
Parameter descriptions are missing or minimal across all tools. The 'edits' param in editFile, 'patch' in patchFile, 'action' in workspace, 'skillName' and 'input' in skill, 'questions' in askFollowupQuestion have only 1-sentence descriptions that do not explain format, constraints, or valid values. LLMs cannot infer parameter semantics from one sentence.
Output schemas are undocumented for all 19 tools. Tool descriptions do not state what fields are returned, what data types they contain, or how to chain results into downstream tool calls. This forces LLMs to guess at response structure and breaks multi-step reasoning.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 55 | 1.25.2+ | v1 |
Apply unified diff patches to files
Read file contents from the workspace
Retrieve full content and context for a specific item
Search and discover available tools in the tool registry
Send a message to a chat channel
Execute a registered skill or plugin capability
Update the agent's execution plan or task list
Search indexed documents using vector embeddings
Search the web for information
Manage the development workspace and files
Write content to a file in the workspace
Generic and ambiguous parameter names invite LLM confusion. 'action' (workspace, designWorkspace), 'input' (skill), 'query' (vectorSearch, webSearch, searchTools), 'theme' (designWorkspace) lack type information and valid value enums. For example, 'action' in workspace could mean 'create', 'delete', 'list', 'configure', the LLM must guess which are valid.
Destructive operations (writeFile, editFile, patchFile, deleteWorkspace operations via workspace tool, bash/executeCommand) lack explicit confirmation or dry-run support. No tool description indicates whether the operation is reversible or what state it modifies. This violates the confirmation-request pattern and risks silent data loss.
No error handling guidance. Tool descriptions do not explain what errors are possible, whether they are retryable, or what the LLM should do if a call fails. For example, executeCommand and bash do not mention timeout handling, command not found errors, permission denied, or exit code semantics.
Security parameters are exposed. 'executeCommand' and 'bash' accept arbitrary shell commands with no mention of injection prevention, allowlisting, or sandboxing. Command arguments passed via 'args' array have no escaping guidance. Credentials or sensitive env vars passed to these tools could leak into logs.
Tools with overloaded semantics: 'skill' is a catch-all that delegates to an external registry with no discovery mechanism exposed in the MCP interface. 'workspace' and 'designWorkspace' accept free-form 'action' strings with no enum or validation. These force LLMs to call searchTools() first or guess at valid values, wasting round-trips.
Tool composition is unclear. It is not documented whether tools are idempotent (e.g., writing the same content twice to the same file), whether concurrent calls are safe, or what the semantics are if a tool is called in the middle of a long operation. No tool description answers these questions.
Pagination and result limits are not mentioned for list-like tools. 'webSearch', 'vectorSearch', 'localGrep', 'searchTools' accept a 'limit' or 'query' but do not document maximum result counts, whether pagination cursors are supported, or how to retrieve 'next' results. This risks context window overload.
Parameter naming inconsistencies reduce clarity. 'filePath' (readFile, writeFile, editFile, patchFile) vs unspecified path format for 'paths' (localGrep); 'query' used in multiple contexts (webSearch, vectorSearch, searchTools, localGrep) without disambiguating the domain (free text vs vector vs pattern); 'channel' in sendMessageToChannel, is this a channel name, ID, or email?