MCP server that performs security analysis on GitHub repositories, including dependency risk assessment, vulnerability scanning via OSV, GitHub Actions workflow security checks, and risk-based governance recommendations.
This MCP server implements 4 GitHub security scanning tools with HTTP transport. Tool definitions are visible and schemas are present, but descriptions lack depth and parameters are minimally documented. Tool names are clear and action-oriented (github_scan, risk_score, recommend_actions, actions_security_scan), but descriptions are overly technical and assume deep domain knowledge. Most critically, parameter descriptions are sparse or missing entirely, forcing LLMs to infer intent. Output schemas are not explicitly documented, responses are shown via code but not as formal schema definitions. The server handles real security concerns (GitHub scanning, vulnerability detection) but the tool interface does not meet production quality standards for agentic use.
Scans GitHub Actions workflow files (.github/workflows/*.yml) for security issues, insecure patterns, permissions misconfigurations, and suspicious actions. Returns findings with severity levels and recommendations.
Scans a GitHub repository for dependency files, secret patterns, and CI/CD workflow anomalies. Extracts package.json, requirements.txt, pyproject.toml, Pipfile, poetry.lock, pom.xml, build.gradle files and identifies potential secrets (matching .env, id_rsa, .pem, secret, key patterns) and GitHub/GitLab workflow files.
Generates governance-based action recommendations based on overall risk level and findings. Returns actions such as ALERT (if secrets found), BLOCK_PR (if high risk), COMMENT (if medium risk), or PASS (if low risk).
Analyzes dependency files to extract package names and calculate risk scores for each package. Uses heuristic-based scoring (9=highest risk patterns like eval/exec/unsafe, 8=deprecated packages, 5=medium-risk packages, 2=baseline) and performs live OSV (Open Source Vulnerabilities) database lookups to identify CVEs and adjust scores based on vulnerability count and severity.
Parameter descriptions are missing or trivial. 'deps' in risk_score says 'Array of dependency file objects with path and content properties', but what formats are path/content? How large can content be? What encodings? LLMs cannot infer valid input without explicit constraints.
Output schemas are not formally documented. The code shows responses via JSON examples (e.g., res.json({ output: { files, deps, findings } })), but LLMs have no formal schema to know what fields to expect or their types. For 'findings' object in recommend_actions, is it { secrets: [], anomalies: [] } or something else?
Tool descriptions are overly technical and domain-specific. 'github_scan' description lists file types and patterns, useful for documentation, but an LLM choosing tools needs context: WHEN to use this vs. actions_security_scan? WHAT makes it different? How does output feed downstream tools?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | - | v1 |
No pagination or result limiting documented. If a repo has 1000 dependency files or 500 OSV vulnerabilities, does risk_score return all? The rubric baseline requires paginated results accept limit/offset and return total_count. The code does not enforce or advertise limits.
Error handling is generic. Code shows 400 for 'Invalid GitHub repo URL format' and 500 for exceptions, but responses do not guide recovery. If github_scan fails with 'Invalid branch', should LLM retry with 'main'? Try a list-branches tool first? The error response gives no hint.
GITHUB_TOKEN is injected via environment variable, which is correct [pattern:secret-injection], but the tool accepts 'repoUrl' as a parameter, no scope or permission declaration. If the token has read-only access, the tool should advertise 'read:repos' scope. If it lacks access to private repos, that should be documented.
'overall' parameter in recommend_actions is an enum {Low, Medium, High}, which is good, but the description does not explain what input range triggers each level. Is Low <10 risk score? Medium 10-50? This forced LLMs to guess the thresholds.
'branch' parameter in github_scan defaults to 'main', which is idiomatic, but the description should note: 'If branch does not exist, the tool will fail, call get_branches first if unsure.' This dependency is undocumented.