Credential audit pipeline — validates API keys against live provider endpoints
The server has 4 tools with basic definitions but significant quality gaps. All tools have names and descriptions present, but descriptions are generic and lack actionable detail. Input schemas are visible in the specification but appear minimal, most parameters lack detailed type constraints and validation guidance. No output schemas are documented in the provided code. The credential management domain is clear, but the tool definitions do not follow production-grade patterns for LLM-optimized descriptions, parameter constraints, or error guidance. This is typical of early-stage community servers (C-D range).
Run credential audit pipeline to validate API keys against live provider endpoints
Retrieve a single credential by key name from allowed list
Get usage statistics for credentials including request count, token usage, and RPM
List all allowed credential keys for the current agent
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract fields without knowing the response structure. Responses should document returned fields, types, and when to call each tool next.
Descriptions lack actionable detail on WHEN to use each tool and what to do with results. 'Retrieve a single credential by key name from allowed list' does not explain the purpose (why retrieve credentials?), when to call list_credentials first, or what the response contains. Descriptions should be 50-200 chars and answer: what does it do, when use it, what does it return?
Parameter descriptions are minimal or missing context. 'key' in get_credential says 'Credential key name to retrieve' but does not explain: is the key case-sensitive? What format? What keys are allowed? get_usage says 'Optional credential key name; if empty, returns all', but does 'empty' mean null, empty string, or omit the param? Ambiguity forces LLM guessing.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 8 | - | v1 |
audit_credentials 'providers' parameter is an array but lacks enum constraint or discovery guidance. Description says 'Optional list of specific providers to check' but does not enumerate valid provider names. Agents cannot know which providers exist without calling list_providers first (if such a tool exists), or will hallucinate invalid provider names.
No error handling guidance. If audit_credentials times out, returns invalid provider, or finds a malformed .env file, how should the LLM recover? No indication of retryable vs. fatal errors, or what the LLM should try next. Error responses should guide recovery: 'Provider not found. Run list_providers() to see available options.'
No pagination or result limits documented. get_usage says it 'returns all' if key is empty, but how many credentials could exist? Returning 1000 usage records in a single response risks exhausting the context window. Results should be capped (e.g., max 50 records) with pagination support (limit, offset, next_cursor) and documented in the description.
Tool naming is functional but generic. 'get_credential' and 'list_credentials' follow the verb_noun pattern, but 'audit_credentials' is vague, does it validate, test, report, or fix? Clearer names like 'validate_api_keys' or 'test_credential_endpoints' would better signal intent. Naming is baseline; not a critical failure but could be sharper.