MCP server that translates a lockfile diff into a human-readable upgrade plan.
dep-diff provides two well-defined tools with comprehensive schemas, detailed descriptions, and clear use-case guidance. Both tools follow the verb_noun naming pattern (analyze_*), have detailed descriptions explaining when and why to use each tool, and include complete input schemas with typed, described parameters. Output schemas are explicitly defined with Zod types covering all response fields. Tool composition is sound: analyze_package_change handles single upgrades while analyze_packages_bulk handles batch scenarios, with a clear division of responsibility. The main gaps are: (1) no documented error recovery guidance in tool descriptions, (2) no explicit mention of rate limiting or timeout behavior despite external API calls, (3) tool descriptions contain embedded examples that could be literal values LLMs reuse (e.g., 'react 18 and 19', 'actions/checkout'), and (4) security handling of GitHub tokens is server-side but not documented in tool descriptions. Overall, this is a B+ implementation with strong definition quality undermined by missing error handling guidance and some best-practice oversights.
Given one package and two versions (from -> to), returns a structured upgrade analysis: semver classification, GitHub release notes summary, detected breaking changes, security advisories fixed in the range, migration guide links, and a clear recommendation. Use when the user asks about a specific package upgrade ('what changed between react 18 and 19', 'is it safe to bump axios from 0.27 to 1.0', 'what does upgrading lodash 4.17.20 to 4.17.21 fix'). Supports npm, pypi, and github-actions (use the action reference as the name, e.g. actions/checkout). For analyzing many packages at once or a Dependabot batch, use analyze_packages_bulk instead.
Analyzes a list of package upgrades in parallel and returns a unified risk report with packages ranked by recommendation level (security > caution > review > likely-safe > safe). Use when the user provides many dependency changes from a Dependabot PR, npm outdated output, lockfile diff, or batch upgrade. Returns: total count, breakdown by semver class, total security fixes found, packages with breaking changes, and per-package details. Limit 50 packages per call (chunk larger lists).
Tool descriptions embed example values ('react 18 and 19', 'axios from 0.27 to 1.0', 'actions/checkout') that LLMs may reuse literally instead of adapting to the actual context. Should use enums and format constraints instead.
No error recovery guidance in tool descriptions. Tools call external APIs (GitHub, OSV, package registries) but don't document what happens on timeout, rate limiting, or network failure, or what the LLM should do next. Missing recovery guidance violates pattern:recovery-guide.
Parameter 'fromVersion' and 'toVersion' lack format constraints (e.g., semver pattern, length bounds). LLMs cannot validate range relationships or detect invalid version strings. Should document as 'SemVer string (e.g., 1.2.3)' with validation inside the tool.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2025-06-18+ | v2 |
analyze_packages_bulk accepts minItems=1, maxItems=50 in schema but does not document what happens if the LLM attempts 51+ packages. Should return a clear error with retry guidance (e.g., 'Too many packages (51). Chunk into groups of 50 and retry.').
GitHub token handling is documented in server.ts but not surfaced in tool descriptions. Users/agents may not understand that rate limiting improves with GITHUB_TOKEN or fallback to 60 req/hr. Tool descriptions should mention this.
Output schemas are well-defined but failedAnalysisSchema always sets recommendationLevel to 'review' without explanation. Should document why failed analyses default to 'review' (conservative, requires human review).