A hardened credential proxy and MCP server that gives AI agents and tools governed access to vaulted API keys with real-time cost tracking, budget caps, rate limiting, and audit trails. Supports universal OpenAI-compatible chat completion endpoint, MCP credential broker, and HTTP egress firewall.
BlackVault exposes 2 tools via HTTP MCP endpoint. Both tools have descriptions and input schemas present, but descriptions lack context for when/why LLMs should call them, parameter descriptions are minimal or missing, and output schemas are not documented. The 'chat' tool has a well-structured input schema with enums and types, but the 'list_models' tool is trivial (no parameters). Error handling is present but generic. Overall structure is functional but falls short of production-grade documentation standards.
Run a chat completion on any allowed model (gpt-*, claude-*, gemini-*, open-source via Nebius). BlackVault injects the real provider key, enforces budget/rate/model limits, and audits the call.
List the AI models this BlackVault token is allowed to use (across all vaulted provider keys).
Output schemas not documented. 'list_models' and 'chat' tool descriptions do not specify what fields/structure the response contains. LLMs cannot plan downstream steps or extract required data without knowing the response shape.
Parameter descriptions incomplete or missing. 'messages' parameter in 'chat' tool has a schema description but lacks domain context: 'OpenAI-style chat messages' does not explain what role values mean in BlackVault's context, whether system prompts are supported differently, or constraints on message length/count.
No pagination or result limiting documented for 'list_models'. If many models are available, a large unfiltered response could exhaust context. No indication of whether results are paginated, capped, or how to filter by provider.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 56 | 2026-07-28+ | v2 |
Error handling lacks recovery guidance. Code shows generic error responses ('error: err.message') without classification (retryable vs user-fixable) or actionable next steps. LLMs have no guidance on what to do on failure.
No tool annotations. Neither 'list_models' nor 'chat' declares readOnlyHint, destructiveHint, or idempotentHint. 'list_models' is clearly read-only but not marked; 'chat' is idempotent (same input → same output if model/messages unchanged) but unmarked.
'list_models' description lacks context for when to call it. Does it list all models user has access to, or all models the vault supports? When should an LLM call this vs. just try a model in 'chat'? Ambiguous.