A CLI tool and MCP server to verify and fix AI/ML environment compatibility (Driver <-> CUDA <-> Wheels) with platform-specific installation guides
env-doctor offers 12 well-named, read-only diagnostic tools with clear descriptions and proper JSON Schema input validation. All tools start with action verbs (env_check, model_check, cuda_info, etc.) and descriptions are in the 60-200 character range, meeting LLM-optimization baselines. However, CRITICAL GAPS exist: (1) NO OUTPUT SCHEMAS are documented anywhere in the codebase, tools return results but the structure is completely undocumented for LLM planning; (2) LIMITED ERROR HANDLING guidance, no recovery instructions, categorization, or self-correction hints; (3) PARAMETER DESCRIPTIONS are present but minimal (10-60 chars), lack detail on formats, constraints, and expected ranges; (4) NO EXAMPLES of tool outputs provided in code or docs, making it difficult to verify what fields agents should expect. These are significant gaps for production agent use. The tool naming and basic schema quality are solid (baseline ~70), but missing output documentation and weak error handling drop the overall score to 68.
Get detailed CUDA toolkit information including: nvcc version and path, all CUDA installations, CUDA_HOME configuration, PATH/LD_LIBRARY_PATH status, libcudart runtime library, and driver compatibility analysis.
Get step-by-step CUDA Toolkit installation instructions tailored to the user's platform. Detects OS/distro, recommends the best CUDA version based on GPU driver, and provides copy-paste installation commands for Ubuntu, Debian, RHEL, Fedora, WSL2, Windows, and Conda.
Get detailed cuDNN library information including: version, library paths, symlink status (Linux), PATH configuration (Windows), multiple version detection, and CUDA compatibility.
Validate docker-compose.yml content for GPU configuration issues. Checks for: missing deploy.resources.reservations.devices, incorrect GPU driver settings, missing runtime: nvidia, and other GPU passthrough configuration problems.
Validate Dockerfile content for GPU/CUDA configuration issues. Checks for: CPU-only base images, missing PyTorch --index-url flags, CUDA version mismatches, driver installations in containers, and deprecated package usage.
No output schemas documented for any tool. Tools return complex diagnostic results (versions, paths, recommendations, metadata) but the response structure is completely invisible to LLMs. This prevents agents from planning downstream operations, extracting specific fields, and chaining tools. Critical for agentic patterns.
Error handling lacks recovery guidance and categorization. No evidence that tools return actionable error messages like 'GPU driver not found, try running nvidia-smi to verify installation.' Errors appear to be simple success/failure, not categorized as retryable, user-fixable, or fatal.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 67 | - | v1 |
Run full GPU/CUDA environment diagnostics. Checks NVIDIA driver, CUDA toolkit, cuDNN, Python AI libraries (PyTorch, TensorFlow, JAX), and WSL2 configuration. Returns status, versions, and recommendations for each component.
Run diagnostics for a specific component. Available components: nvidia_driver, cuda_toolkit, cudnn, python_library, wsl2.
Get the safe pip install command for an AI library based on the detected GPU driver. Automatically determines the correct CUDA version and wheel URL for libraries like PyTorch, TensorFlow, and JAX.
Check if an AI model fits on available GPU hardware. Analyzes VRAM requirements across precisions (fp32, fp16, bf16, int8, int4) and provides recommendations for running the model.
List all available AI models in the database, grouped by category (LLM, diffusion, audio, VLM). Includes parameter counts and HuggingFace IDs.
Check Python version compatibility with installed AI libraries. Detects version conflicts where the current Python version is outside a library's supported range, and identifies dependency cascades where one library's constraint forces version limits on downstream packages.
Get detailed WSL (Windows Subsystem for Linux) environment information including: environment type (Native Linux / WSL1 / WSL2), kernel version, GPU forwarding status, and diagnostic checklist for WSL2 CUDA support.
Parameter descriptions are minimal (10-60 chars) and lack detail on formats, constraints, and expected values. Example: 'component' param in env_check_component has enum values but no description of what each component checks. 'model_name' in model_check lacks format guidance (is 'llama-3-8b' or 'meta-llama/Llama-2-7b-hf' preferred?).
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in tool definitions. All 12 tools are read-only and idempotent, but this is not explicitly declared in the schema, agents must infer it from descriptions.
model_list tool lacks pagination parameters (limit, offset, page_size). No evidence of pagination support, which is critical when listing 'all available AI models', could return hundreds of results that blow context windows.