Deterministic MCP security middleware for enforcing policy-based tool access control and data flow analysis
Lilith Zero presents as a policy-enforcement middleware for MCP servers, not itself a production MCP server. The tool definitions visible in the codebase are **mock/test tools** (sdk/tests/resources/mock_tools.py, vulnerable_tools.py, and example servers). These are intentionally unsafe, poorly described, and designed to demonstrate policy violation scenarios, not production quality. The actual Lilith Zero binary is a Rust policy validator/webhook server (Dockerfile shows it runs as a webhook endpoint on port 8080). No real MCP-compliant tool definitions are registered by Lilith Zero itself. The 38 'tools' listed are from test fixtures and Python example servers embedded in the repository, they are not Lilith Zero's own tools. Evaluating these test fixtures as if they were production tools would be misleading; they are deliberately vulnerable examples used to validate Lilith Zero's policy checking capability. The real value of Lilith Zero is policy enforcement, not tool quality. That said, evaluating the visible tool definitions against the rubric: most lack substantive descriptions, many have trivial or missing schemas, parameter types are inconsistently declared, and error handling is absent. This is appropriate for mock/demo tools but means the score reflects the actual definition quality of what is visible in the codebase.
Add two numbers together.
Permanently archive (delete) a document. Requires confirmed=true.
Performs a mathematical calculation.
Evaluate a simple arithmetic expression (e.g. '2 + 2').
Query the internal knowledge database.
Permanently delete a record from the database.
Deletes records from a database table.
Divide a by b. Blocked by policy — use multiply with reciprocal.
Tool definitions are from test fixtures and example servers, not production Lilith Zero tools. These are intentionally vulnerable mock tools designed to demonstrate policy violations. Evaluating them as Lilith Zero's own tools is conceptually incorrect, Lilith Zero is a policy validator, not a tool provider.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 59 | 2024-11-05+ | v1 |
Executes a shell command (Extremely Dangerous).
Executes a raw SQL query against the production database. WARNING: No RLS or authorization is performed.
Execute system command (Logic Rules).
Exports data to an external cloud sink.
Export data to cloud (Sink).
Fetches content from an external URL. WARNING: Unrestricted outbound access.
Returns the current system time.
Return the current UTC time.
Current UTC time
Returns sensitive user profile data (PII).
Get user profile (Source of Taint).
Multiply two numbers.
Simple health check.
Post a message to the public Slack channel #general.
Posts a message to a Slack channel.
Execute a raw SQL query against the production database.
Raw SQL query
Mock database read for basic flow tests.
Read a document report from the document store.
Reads sensitive user data (PII).
Redact personally identifiable and confidential information from text.
Sanitize data (Removes Taint).
Search the web for information on a topic.
Search the web
Sends an email to an external recipient.
Mock slack send.
Simulates a long-running process.
Compute the square root of a non-negative number.
Summarize a block of text into bullet points.
Search the public internet for information.
Duplicate tool names detected (get_user_profile, post_to_slack, search_web, get_time, query_database appear 2+ times). This suggests cross-copying from multiple example servers without de-duplication. LLMs will be confused about which variant to call.
Descriptions are consistently shallow (20 - 55 characters). Most lack context for when to use the tool vs similar ones. E.g., 'Search the web' (13 chars) does not explain query format, rate limits, or result structure. Rubric baseline for A-grade tools: 50 - 200 chars with context.
Many tools with destructive or sensitive actions (execute_shell, delete_records, delete_record, execute_sql, archive) lack error handling guidance or confirmation requirements. Descriptions do not explain prerequisites, dry-run options, or what the LLM should do on failure.
Parameter descriptions are minimal or missing for many tools. E.g., 'send_slack' has msg param described as 'Message to send', no mention of length limits, encoding, or formatting. 'execute_shell' command param lacks syntax hints, allowed commands, or injection warnings. Rubric requires all params to have substantive descriptions.
No tool declares its risk level or required permissions (e.g., read:email, write:database). This makes it impossible for agents to be configured with least-privilege scope. Risk annotations are noted in the input list (READ_ONLY, WRITE, DESTRUCTIVE) but not visible in the tool definitions themselves.
No output schemas are documented. Tools return responses but LLMs cannot plan downstream calls without knowing what fields to expect. E.g., does 'get_user_profile' return user_id, name, email, all three, or something else? Pattern requires documented return types for ALL tools.
No error handling guidance. Tools provide no recovery instructions. E.g., if 'read_db' fails with a SQL syntax error, what should the LLM do next? Pattern requires errors to guide recovery and categorize as retryable, user-fixable, or fatal.