Local semantic filesystem layer with MCP server integration. Bundles Rust binary (core + CLI + MCP server) and Python AI layer (embeddings, extraction, relation indexing) for semantic search, indexing, and graph analysis of local codebases.
Organon exhibits significant gaps in definition quality. While most tools have basic descriptions (10 - 200 chars), parameter schemas are inconsistent, some tools like 'search' and 'context' expose detailed JSON schemas with type and description fields, but others like 'stats' and 'health' lack visible input schemas entirely. Parameter descriptions are often minimal (e.g., 'Output as JSON' for the 'json' parameter across multiple tools). No output schemas are documented for any tool, forcing LLMs to infer result structures. Error handling is absent, no recovery guidance, retryability classification, or actionable error messages. Tool naming follows verb_noun convention (search, extract, index, export), which is correct, but several tools combine multiple concerns: 'clean' performs both dead-entity removal and stale-relation cleanup; 'archive' both moves files and updates metadata. No dry-run or confirmation patterns for destructive operations (clean, archive). Security gaps: the 'index' and 'archive' tools accept arbitrary filesystem paths without visible input validation or sanitization against path traversal. Composition: tools are reasonably atomic, but 'plan' requires file arrays and produces task plans without documented downstream integration points, if an agent uses 'plan' output to select files for 'context', there's no guarantee 'plan' returns file paths in a format 'context' accepts.
Archive files by moving them and marking them as archived in the graph.
Clean up dead or stale entities from the graph database.
Retrieve context around a query or file within a scoped directory, limited by token budget.
Show differences in entity metadata and content for a path.
Find duplicate files based on content hash.
Export the entity graph in various formats (JSON, etc.).
Extract text content from a file. Supports code files, PDFs, and plain text formats.
Find files with filters for state, extension, creation date, modification date, and size.
No output schemas documented for any of the 20 tools. LLMs cannot infer response structure and must guess which fields exist and their types. This forces unreliable downstream integration and wastes tokens on clarification calls.
Multiple tools with 'dry_run' and 'apply' boolean parameters lack confirmation-request patterns. 'clean' (DESTRUCTIVE) and 'archive' (IRREVERSIBLE) accept bare boolean flags without requiring explicit user confirmation. Agents can accidentally trigger irreversible data loss.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
Analyze the dependency graph of a file, showing imports and dependents up to a specified depth.
Check the health of the indexer daemon and database connectivity.
Show the edit history of a file, including timestamps and previous states.
Analyze the impact of changes to a file, showing which other files depend on it.
Index files and extract embeddings for semantic search. Watches filesystem changes and updates the vector store.
List entities in the graph with optional state filtering.
Generate an execution plan for a task based on related files.
Find test files related to given source file paths.
Semantic search across indexed entities. Embeds the query and finds semantically similar files using vector similarity.
Find files semantically similar to a given file by looking up its existing vector embedding.
Display statistics about the indexed entities and graph.
Show the current status of a path or the entire workspace.
No error handling guidance in any tool description. Tools like 'index' and 'search' may fail (index: watch filesystem, parse files; search: embedding service down), but no description explains what can go wrong, whether errors are retryable, or what the agent should do next.
'stats' tool has empty input schema ({}) but no description explaining what it returns or when to call it. Cannot score schema as visible but content is missing.
'ls' is a generic acronym, non-obvious to LLMs unfamiliar with Unix. Should be renamed to 'list_entities' to follow verb_noun convention and clarify intent. Current name forces the LLM to reason about abbreviations.
Tools 'clean' and 'archive' accept paths/directories without documented constraints or validation. No description specifies whether path must be absolute, relative, or scoped to workspace. Path traversal risks invisible. E.g., can 'archive dir=../../sensitive' succeed?
'clean' combines two distinct operations (dead-entity cleanup + stale-relation pruning) under one tool. Should split into 'clean_dead_entities' and 'clean_stale_relations' so agents can compose them independently and retry individual steps on partial failure.
Parameter descriptions across tools are minimal. E.g., 'json' param described simply as 'Output as JSON' across 6 tools (health, history, impact, duplicates, diff), no explanation of why an agent would choose JSON vs default format, or what structure JSON implies.
'mode' parameter in 'search' and 'context' described as 'Search mode (semantic or other)' but no explanation of what 'other' means, what triggers each mode, or whether enum values exist. LLMs will guess.
No pagination guidance in 'search', 'find', 'related_tests'. These tools have 'limit' parameters, but no 'offset', 'page', or 'next_cursor' to support continuation. If results exceed limit, agent cannot fetch the rest.