Verifies code through external AI models with Zero-Trust approach. Supports single file review, git diff analysis, and multiple file cross-file dependency checking. Includes retry with exponential backoff, automatic fallback to other models on error, result caching, and language-aware checks for 10 programming languages.
The server has 6 tools with varied definition quality. The primary tool (verify_code) has a comprehensive schema with good coverage of input parameters and detailed descriptions, placing it in the B range. However, several tools lack proper descriptions and schemas are minimal. Most tools have descriptions present but they are verbose and not LLM-optimized. Parameter descriptions are generally adequate but some tools are missing required parameter details. No output schemas are documented anywhere in the provided source.
Shows cache statistics for code verification results. INFORMATION: - Cache status (enabled/disabled) - Current size / max size - TTL (entry lifetime in seconds) - Fill percentage PURPOSE: Cache stores verification results to speed up repeated requests with the same code. If code hasn't changed, result is taken from cache (~50ms instead of 2-5 sec). USAGE: - "Show cache stats" - "How many results in cache?" - "Check cache status" RESULT: Detailed information about cache state
Diagnose API connectivity and show recent errors. PURPOSE: Helps troubleshoot when code verification fails. Tests connection to each AI provider and shows recent error log. CHECKS: - API key presence for each model - Connection test to z.ai and OpenRouter - Recent error history with timestamps - Recommendations for fixing issues USAGE: - "Diagnose Argus" - "Why is verification failing?" - "Check API status" - "Show recent errors" RESULT: Diagnostic report with connection status and error analysis
Shows list of all available AI models for code verification. INFORMATION: - Model name and key - Provider (z.ai, OpenRouter) - Status (✅ available / ❌ unavailable) - Cost per 1K tokens - Max tokens - Current default model USAGE: - "Show available models" - "What models can I use?" - "List models for code review" RESULT: Table with full information about each model
Retries the last failed code verification with fallback models. PURPOSE: If the primary model failed during code verification, this tool allows you to retry the verification using fallback models. USAGE: - "Retry with fallback models" - "Try other models" - "Use fallback for last check" NOTE: This will use the exact same code and parameters from the last failed verification.
Output schemas are completely undocumented. None of the 6 tools include a documented return schema. LLMs cannot plan downstream operations or extract required fields without knowing what the tool returns. This violates pattern:tool and mxe:response-field-naming.
Tools list_models, cache_stats, retry_with_fallback, and diagnose have empty or minimal input schemas (empty properties {}). While they may not require parameters, the schema should still be properly structured, and the tool descriptions should explicitly state 'This tool requires no parameters.'
Tool descriptions are verbose and not optimized for LLM parsing. Descriptions range from 200-400+ characters with excessive use of all-caps headers (PURPOSE, CHECKS, USAGE, RESULT, INFORMATION, NOTE). The baseline for A+ tools is 50-200 chars. This wastes tokens and makes it harder for LLMs to extract key intent.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 50 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 24 | 2024-11-05+ | v1 |
Sets default model for current session. PURPOSE: Changes the base model that will be used for all subsequent code checks if model is not specified explicitly. AVAILABLE MODELS: - glm-4.7 - fast, $0.40/M input (default) - gemini-flash - very fast, $0.50/M input - minimax - medium speed, $0.30/M input USAGE: - "Set Gemini as default model" - "Use MiniMax for all checks" - "Switch to GLM 4.7" NOTE: Change applies only to current session
Verifies code through external AI model with Zero-Trust approach. MODES: 1. Single File - review one file (params: code + file_path) 2. Git Diff - review changes via git diff (param: diff) 3. Multiple Files - review multiple files with cross-file dependencies (param: files[]) FEATURES: - Retry with exponential backoff (3 attempts) - Automatic fallback to other models on error - Result caching (TTL: 1 hour) - Language-aware checks for 10 languages (Python, JS, TS, Vue, React, Go, Rust, Java, PHP) - Security (OWASP), performance, and architecture checks MODELS: - glm-4.7 (z.ai) - $0.40/M input, fast - gemini-flash (OpenRouter) - $0.50/M input, very fast - minimax (OpenRouter) - $0.30/M input USAGE: - "Review my code" - basic check - "Check code with Gemini" - model selection - "Verify changes in multiple files" - cross-file review
Parameter verify_code.project_stack is an optional nested object but has no description of when to use it or what impact it has on verification results. If it is optional, the description should state the default behavior when omitted.
No error handling patterns documented. The verify_code tool mentions 'Automatic fallback to other models on error' and 'Retry with exponential backoff' but provides no guidance on what errors are retryable, what errors should be escalated to the user, or what the LLM should do if all fallback models fail. This violates pattern:recovery-guide.
verify_code accepts a 'model' parameter with enum constraint, but the description references 'Available: {glm-4.7, gemini-flash, minimax}', this is brittle and not kept in sync. The enum declaration should be the source of truth; the description should reference the enum or state 'See model parameter enum for available options.'