Open security engineering for code written by humans and AI agents — every repair carries its own offline-verifiable proof. Provides security scanning, code remediation, fix verification, and agent-based task execution.
DvalinCode presents a well-designed security-focused MCP server with clear tool responsibilities and comprehensive parameter schemas. All 9 tools have explicit descriptions and well-formed JSON Schema inputs with proper type constraints and enums. Naming is consistent and action-oriented (scan, verify, list, get, run). However, output schemas are not documented in the provided source, only input schemas are visible. Descriptions are clear but vary in depth (80-250 chars). Error handling guidance is implied by the tool descriptions but not formally specified in schema definitions. The server demonstrates strong composition (each tool has singular responsibility) and good parameter design (enums for permission_mode, scope, executor; required fields properly marked). No security issues detected with credentials in parameters.
Scan a workspace and persist a compact local verification workflow. Call this only after dvalin_scan reports a finding that will be repaired; it returns a workflow ID for dvalin_get_finding and dvalin_verify_findings. It never edits project files.
Render the Markdown audit evidence for a DvalinCode run (defaults to the latest run).
Read one compact finding from a persisted security workflow by fingerprint.
Get a durable DvalinCode session summary and its latest audit anchor.
List built-in and optional scanner engines, availability, and reviewable install commands.
Run a complete governed coding task inside DvalinCode. Long calls are expected; callers should use a generous timeout.
Output schemas not documented. The input schemas are complete and well-formed, but there is no visible documentation of what each tool returns. LLMs need to know the structure of responses to plan downstream tool chains and extract relevant fields.
dvalin_get_session and dvalin_get_evidence have sparse descriptions (70 - 75 chars). They explain WHAT but not clearly WHEN to use them or what differentiates them. 'Get a durable DvalinCode session summary' does not explain when you'd call this vs dvalin_run_task. Expand to 100 - 150 chars with context about use cases.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2025-06-18+ | v2 |
Scan a workspace for injection, hardcoded secrets, XSS, dynamic code execution, and unsafe shell use. Deterministic and read-only: it runs no model, needs no credentials, never edits project files, and does not persist Dvalin state — safe to call by default after writing code. Returns findings with file, line, severity, and rule reference. Call dvalin_begin_verification only when a finding will be repaired and independently re-verified.
Re-scan a persisted workflow and independently verify that the blocked targets are gone and no new severe findings were introduced. Unlike dvalin_scan, this DOES execute the project's own checks (test, typecheck, build) so the verdict rests on exit codes Dvalin observed rather than on any report it was given — every command still passes the org policy gate before it runs. Returns a Verified Fix Record: a portable, offline re-verifiable statement of what was checked and what was found.
Re-derive a Verified Fix Record offline. Recomputes the record hash and re-checks that its verdict follows from its own evidence, so a record that was edited after it was issued fails here. Needs no workspace, no network, and no Dvalin state — call it to confirm that a fix record handed to you is sound before relying on it.
dvalin_run_task's description ('Long calls are expected; callers should use a generous timeout') does not explain WHAT the task does or what the 'prompt' parameter should contain. Does it execute code? Apply fixes? Generate a plan? Expand to clarify the execution model and what 'governed coding task' means.
No parameter descriptions provided in the schema for any tool. While parameter names are clear (cwd, scanners, scope, finding_limit, timeout_seconds, prompt, session_id, workflow_id, fingerprint, executor, record), LLMs benefit from explicit descriptions in the schema, e.g., 'scanners: comma-separated list of scanner engine names (builtin, semgrep, trivy, osv-scanner)' or 'finding_limit: maximum number of findings to report before stopping scan (limits context window usage)'.
Error recovery guidance not visible in schemas. If dvalin_get_session fails with 'session_id not found', what should the LLM do? List sessions? Retry? Create a new one? Add descriptions like 'If session not found, try dvalin_run_task to create a new session' to guide recovery.