A code indexing and refactoring MCP server that provides structural inspection, semantic navigation, and edit capabilities for codebases across multiple languages
ARNO presents 19 tools with complete JSON Schema input definitions and descriptions for all tools. However, most descriptions are brief (30-90 chars), falling below the production baseline of 194 chars average. Parameter descriptions are present and specific (e.g., 'Commit hash for freshness comparison'), but parameter validation constraints (min/max for numeric fields, enum alternatives for mode parameters) are sparse. Output schemas are not documented in the provided source, the response shape for each tool is unknown to the LLM at planning time. No tool demonstrates error handling guidance, recovery patterns, or actionable error messages. The tool naming is functional but generic: verb_noun is present (e.g., workspace_tree, read_symbol, replace_range), but names like 'arno.outline' and 'arno.search_nudge' lack clarity for LLM intent inference. The 'arno.rename' and 'arno.replace_symbol' tools accept an expectedRevision parameter for concurrency control, indicating thoughtful design, but this feature is not leveraged across all write operations (e.g., arno.replace_range accepts it, but arno.revert does not). Overall, the definitions are functional but underoptimized for LLM agentic reasoning, descriptions are too terse, output schemas are invisible, and error recovery paths are absent.
Returns a summary of changes in the workspace
Creates a checkpoint of the current workspace state with an optional note
Computes a unified diff between two file revisions
Retrieves workspace events (changes, indexing, etc.) after a cursor
Searches for symbols or content matching a query pattern
Retrieves the output of a completed background job
Queries the status of a background job
Returns the grouped outline and flat symbol list for a file with parser information
Output schemas are not documented. The response shape for each tool is invisible to the LLM at planning time. Tools like arno.retrieve, arno.repository_map, and arno.references return complex results (ranked files, symbol contexts, reference lists) but the LLM must infer the response structure. This violates pattern:tool and pattern:response-shaper, documented return types are non-negotiable for agentic planning.
Tool descriptions are too brief (30 - 90 chars vs. production baseline of 194 chars). Examples: 'arno.outline' is 77 chars, 'arno.changes' is 46 chars. Descriptions lack context about WHEN to use the tool or its relationship to similar tools. LLMs cannot distinguish between arno.find (query-based search), arno.retrieve (context retrieval), and arno.repository_map (relevance ranking) from their current descriptions alone. Each must explain its unique purpose and when to call it instead of alternatives.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 48 | 2026-07-28+ | v2 |
Reads a line range from a file, paging by token budget with continuation handles
Reads a symbol's body by symbol ID or name, returning its source and metadata
Finds all references to a symbol across the workspace
Renames a symbol across the workspace and returns the required edits
Replaces a line range in a file with new code
Replaces a symbol's entire body with new code
Ranks files in the workspace by relevance to a query within a token budget
Retrieves the most relevant code symbols and their context for a query
Reverts the workspace to a previously created checkpoint
Provides contextual search suggestions based on recent activity
Returns a plain, bounded structural listing of the workspace for orientation
No error handling guidance or recovery patterns across any tool. None of the 19 tools include descriptions of failure modes or recovery steps. Example: arno.rename and arno.replace_range accept an expectedRevision parameter for concurrency control, but the description does not explain what happens if the revision is stale or what the LLM should do in response. Violates pattern:recovery-guide and pattern:error-classification.
Numeric parameters lack validation constraints. arno.read_range accepts startLine and endLine (integers) with no documented minimum, maximum, or relationship constraints. arno.workspace_tree accepts maxEntries (integer) with no min/max bounds, LLMs may pass 0, negative, or absurdly large values. arno.find accepts maxResults with no documented upper limit. Violates pattern:constrained-input.
Enum parameters lack explicit constraint documentation. arno.find accepts a 'mode' parameter with enum ['auto', 'symbol', 'semantic'], but the description does not explain the difference between modes or when to use each. An LLM reading the description cannot determine whether mode='auto' is appropriate for a symbol search or if mode='semantic' is required. The enum values are in JSON Schema, not in the description text, LLMs often miss or misinterpret JSON Schema enums.
Inconsistent concurrency control. arno.rename, arno.replace_symbol, and arno.replace_range accept expectedRevision for optimistic locking, but arno.revert does not. This inconsistency suggests that arno.revert may not be safe to retry or may allow accidental overwrites if called twice. Either all write operations should support expectedRevision, or the lack of it should be documented with explicit guidance on idempotency and retry behavior.
Destructive operations lack confirmation or dry-run patterns. arno.rename, arno.replace_symbol, arno.replace_range, and arno.revert are all write operations that alter code. None of them offer a dry_run flag, confirmation_token, or preview-before-execute pattern. Agents can accidentally corrupt entire codebases. Violates pattern:confirmation-request.
Parameter interdependencies undocumented. arno.read_symbol accepts both symbolID and symbolName, but the description does not state whether they are mutually exclusive, which one takes precedence, or what happens if both are provided. arno.references similarly accepts both symbolID and symbolName. LLMs will likely pass both, creating ambiguity. Violates review:param-relationships.
Missing pagination and result limits. Tools like arno.find, arno.repository_map, and arno.retrieve accept 'maxResults', 'maxTokens', or 'budget' parameters, but descriptions do not specify defaults, mention pagination mechanisms (cursor, offset, page), or explain what happens if results exceed the limit (truncated, paginated, error?). Violates pattern:paginated-result and mxe:enforce-result-limits.