The open source standard for secure, private AI: operating-system enforcement is live on macOS today; Linux and Windows are not live enforcement yet, and your data stays under your own keys with current portability bounds called out in the Assurance Matrix. Any agent, local or cloud, one machine or many.
Critical failures across all 13 tools. The source code provided contains only a Python requirements file and tool name declarations in tool-names.ts, but NO actual tool definitions, input schemas, parameter descriptions, or output documentation are visible. All tools are listed with generic placeholder descriptions like 'Store data in agent memory (cooperative tool)' that lack any actionable detail for LLM reasoning. No input schemas can be verified. No parameter documentation exists. The server declares 13 tools but provides zero evidence of proper schema registration or parameter typing. This is an F-grade implementation masquerading as a framework.
CRITICAL: No input schemas visible for any of 13 tools. Cannot verify parameter types, constraints, or validation. By rubric rule, schema score must be 0 when schemas are not visible.
CRITICAL: All tool descriptions are generic placeholders ('Store data in agent memory (cooperative tool)', 'Retrieve data from agent memory (cooperative tool)', etc.). Descriptions are 10-50 characters, well below the 10-1024 baseline. By rubric rule, descriptions under 20 chars score 0-20. These generic phrases do not explain WHAT the tool does precisely, WHEN to use it, or what it returns.
Create comprehensive input schemas for all 13 tools using JSON Schema with explicit type definitions, required fields, and parameter constraints. Every parameter must have a 'type' and 'description' field. Reference: https://arcade.dev/patterns/constrained-input
Rewrite all tool descriptions to follow the LLM-optimized pattern: 'Action: [verb + noun]. Context: [when to use]. Returns: [what fields, in what format]. Example use case: [concrete scenario].' Target 50 - 200 chars per description. Current placeholders are non-functional.
For each parameter, add a description explaining its purpose, valid range, format, and constraints. Example: 'The memory key (1 - 256 alphanumeric characters and underscores, must be unique per agent). Used to retrieve the stored value.' Do NOT rely on parameter names alone.
Document output schemas for each tool. Show what fields the response contains, their types, and what the agent can do next. Example: sanctuary_recall should document whether it returns {success: boolean, value: any, expires_at: string} or some other structure.
Add error handling guidance to every tool description. Example: 'If the key is not found, returns error code NOT_FOUND with message "Key '{key}' does not exist in agent memory.". Try sanctuary_recall first to check if the key exists, or use sanctuary_capabilities to list available keys.' Reference: https://arcade.dev/patterns/recovery-guide
For stateful tools (cursor-based sanctuary_events_*), document the state machine: 'Call open_cursor first to initialize a cursor ID. Pass the cursor ID to each read call. Call close_cursor when done.' Include error scenarios (cursor timeout, invalid cursor ID).
Score history
Overall score trend
First recorded score · v2 rubric
27/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-23
F
27
2026-07-28+
v2
irreversiblesource verified13/100
Delete data from agent memory (cooperative tool)
sanctuary_helpread onlysource verified13/100
Get help with Sanctuary Framework capabilities (cooperative tool)
sanctuary_hidewritesource verified13/100
Hide sensitive data from logs and memory (cooperative tool)
sanctuary_recallread onlysource verified13/100
Retrieve data from agent memory (cooperative tool)
sanctuary_rememberwritesource verified13/100
Store data in agent memory (cooperative tool)
sanctuary_who_am_iread onlysource verified13/100
Get current agent identity information (cooperative tool)
CRITICAL: No parameter documentation visible. Cannot verify parameter names, types, descriptions, constraints, or defaults for any tool. An LLM cannot safely call tools without knowing what parameters exist and what they mean.
CRITICAL: No output schemas documented. Cannot verify what fields tools return, what types they are, or what data the LLM should expect. Agents cannot plan downstream calls without knowing response structure.
HIGH: Tool names use underscores and generic patterns (sanctuary_remember, sanctuary_recall) but lack clarity on semantics. 'sanctuary_events_open_cursor' vs 'sanctuary_events_read' vs 'sanctuary_events_close_cursor' form a cursor-based pattern that could be simplified. Naming does start with action verbs (good), but function relationships are unclear.
HIGH: No evidence of error handling documentation. Tools like 'sanctuary_forget' (IRREVERSIBLE risk) and 'sanctuary_compound_execute' (WRITE risk) should document what errors can occur, whether they are retryable, and what the agent should do. No recovery guidance found.
HIGH: No evidence of input validation or sanitization documentation. Tools that store or manipulate agent memory (sanctuary_remember, sanctuary_compound_execute) should document maximum input sizes, allowed character sets, and what happens on invalid input.
HIGH: No documented pagination or result limits for tools that could return large datasets (sanctuary_events_read, sanctuary_audit_search). Without explicit limits and pagination support, agents risk context window exhaustion.
MEDIUM: Tool composition and atomicity unclear. 'sanctuary_compound_execute' hints at multi-step operations, but without documentation on what operations it supports, what guarantees it provides (ACID?), and how it relates to individual tools, agents cannot reliably use it.
MEDIUM: No evidence of idempotency guarantees documented. Agents retry on failures, tools like sanctuary_remember and sanctuary_hide should state whether repeated calls with the same input produce the same result (idempotent) or have cumulative side effects.
sanctuary_remembersanctuary_hidesanctuary_forget
For destructive tools (sanctuary_forget, sanctuary_compound_execute with delete ops), implement a dry-run or confirmation step. Document this in the tool description so agents know they can preview what will happen before committing. Reference: https://arcade.dev/patterns/confirmation-request
Add pagination support to sanctuary_events_read and sanctuary_audit_search. Document the parameters (limit, offset/cursor) and response structure (results array, next_cursor, total_count). Cap default limit at 20 and maximum limit at 100.
Clarify the relationship between sanctuary_compound_execute and individual memory tools. Document which operations it supports (remember + forget atomically?), what guarantees it provides, and what errors can occur in a partial failure scenario.
Document idempotency for each tool. Example: 'sanctuary_remember is idempotent, calling it twice with the same key and value produces the same result.' Or: 'sanctuary_forget is idempotent, deleting an already-deleted key returns success without error.' Reference: https://arcade.dev/patterns/idempotent-operation
Add input validation rules to descriptions. Example for sanctuary_remember: 'The value parameter accepts any JSON-serializable object up to 1 MB. Non-serializable types (functions, circular references) return a validation error with a hint to use JSON.stringify or similar.'
Implement and document per-tool permission checks. Example: 'sanctuary_forget requires agent:memory:delete permission. If called by an agent without this scope, returns error PERMISSION_DENIED with the required scope in the message.' Reference: https://arcade.dev/patterns/permission-gate