Automated registry and trust-scoring system for Model Context Protocol (MCP) servers with AI-driven evaluation, metrics collection, and release automation
This is a metrics/trust scoring utility library masquerading as an MCP server. The 25 tools are well-named Python functions with clear verb-noun patterns (fetch_*, compute_*, detect_*, load_*, save_*, etc.), but the descriptions and schemas have significant gaps. Most tools lack formally-defined input schemas in the MCP sense, parameters are documented in docstrings and function signatures, not in JSON Schema with type/description pairs. Output schemas are completely undocumented; callers must infer structure from code. Descriptions are technical and precise but often exceed 200 chars, burying key details. No error handling guidance, no recovery suggestions, no permission gates. The codebase is clearly engineered for internal batch processing (star_history maintenance, repository metrics collection), not as a production MCP server for agent consumption. Transport mechanism is entirely unknown, no indication of STDIO, HTTP, or SSE exposure. Tool compositions are single-responsibility (good), but many require sequential calls with data threading (e.g., fetch_repo_details → readme_stats → detect_red_flags) that could be orchestrated server-side into higher-level tools.
Returns (score 0-100 or None, source). Rubric-based for fresh analyses; falls back to legacy 1-10 quality_score for entries not yet re-judged.
Converts contributor count to a 0-100 community score using logarithmic saturation at 40 contributors.
Computes final trust score (0-100) from raw metrics and AI analysis. Weights: AI 35%, maintenance 20%, popularity 15%, docs 15%, security 10%, community 5%. Applies red-flag penalties and renormalizes missing components.
Returns (flags, penalty). Each flag: {id, label, penalty, fatal}. Fatal flags (archived/disabled/gone) bar listing regardless of score. Detects: repo gone, archived, disabled, new repo, young owner, pipe-to-shell, injection markers, typosquat, injection attempts, security concerns.
Computes documentation quality score (0-100) based on health percentage, README character/heading counts, code blocks, license, and security policy.
No documented output schemas for any tool. Callers must infer structure from code or trial-and-error. Causes agent hallucinations about return types and field names.
Parameter descriptions are technical and lack guidance on when to use vs. similar tools. fetch_commits_90d vs. fetch_scorecard both indicate recency/activity; descriptions don't explain why an agent would pick one over the other.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Number of commits in the last 90 days, capped at 30. None on error.
GitHub community profile health percentage (0-100). None when unavailable.
Contributor count via the Link-header trick (per_page=1, read rel=last). Returns contributor count or 100 for very large communities.
Account creation date of the repo owner. None on error.
Core repo fields. Returns dict, or {'gone': True} on 404, or None on error. Fetches repository metadata from GitHub API including stars, pushed_at, created_at, archived status, disabled status, license_spdx, default_branch, owner info, description, language, topics, and last_update.
OpenSSF Scorecard result. None when the repo isn't indexed (very common for small repos — never treated as a negative signal). Returns score, date, and individual check scores.
True if SECURITY.md exists at the root or under .github/.
Converts final trust score (0-100) to letter grade: A (80+), B (65+), C (50+), F (<50).
Load the known servers cache from disk at given path. Returns dict with 'servers' and 'last_scan' keys.
Parse the exclusion file into a set of lowercase 'owner/repo' slugs. Skips blank lines and comments. Returns empty set if file does not exist.
Load history as {full_name: [(date_str, stars), ...]} sorted by date.
Computes 0.7 x push recency + 0.3 x commit cadence. Returns 0-100 score or None when data unavailable.
Parse AI response text into a dict. Strips markdown fences, validates required fields, sanitizes category (unknown -> 'other'), clamps rubric dimensions to 0-4, synthesizes rubric from legacy quality_score if needed.
Converts star count to a 0-100 popularity score using logarithmic saturation at 30,000 stars.
Plain-text statistics over the FULL (untruncated) README. Returns character count, heading count, code block presence, and pipe-to-shell detection.
Save the known servers cache to disk. Adds schema_version and last_scan timestamp to data before writing.
OpenSSF Scorecard overall score (0-10) scaled to 0-100. Returns None when not indexed.
Star growth vs the nearest snapshot 6-10 days back (d7) and 25-40 (d30). Returns {d7: delta or None, d30: delta or None}.
Truncate text to max_chars, appending '...' if cut.
Append today's snapshot for each server, prune old data, rewrite the file. Retention: every snapshot for ~53 weeks, then first-of-month only; repos absent from servers list with newest snapshot >180 days old are dropped.
No error handling guidance. Functions silently return None on API failures (e.g., fetch_repo_details returns None on RequestException, detect_red_flags returns (None, None)). Agents cannot distinguish 'API down' from 'repo not found' from 'malformed input'.
Many tool descriptions exceed 200 chars and contain implementation detail (API endpoint names, regex patterns, response field names). Long descriptions waste tokens; implementation details belong in code comments, not user-facing docstrings.
No input validation constraints documented. Numeric parameters (days_since_push, commits_90d, final score) lack min/max bounds. String parameters (full_name, full_text) lack length/format constraints. Agents cannot validate inputs before calling.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). save_cache and update_star_history are clearly write operations; agents cannot distinguish them from read tools without consulting descriptions.
Sequential tool chains not pre-composed. Agents must call fetch_repo_details → readme_stats → detect_red_flags → compute_trust in sequence. Each thread breaks if one tool fails. Could be abstracted into higher-level 'evaluate_repository' or 'rank_repository' composites.
No pagination support for tools that could return large results (e.g., fetch_contributor_count caps at 100, but no limit/offset parameters). Tools returning arrays (e.g., detect_red_flags.flags) should support limit and cursor/offset.
Tools like ai_score and parse_ai_response accept/return heterogeneous objects with optional fields (rubric vs. quality_score, 'gone' vs. 'repo_id'). No formal schema definition forces agents to attempt multiple field interpretations, increasing hallucination risk.
No permission gates or scope declarations. Tools expose raw GitHub API responses with sensitive owner/account metadata (owner_login, owner_type, created_at). No verification that calling agent/user has authorization to fetch or act on this data.