A collection of MCP servers including example tools, podcast maker, and GitHub PR reviewer
This collection averages 48/100 across 6 tools, landing in the 'Poor' (D) tier. While schemas are present and descriptions exist, they fall short of production quality. Most tools lack rich error guidance, parameter constraints are minimal, and several descriptions are generic or incomplete. Tool composition is reasonable (single responsibility mostly honored), but output schemas are not documented, and error handling is absent or minimal. The collection would benefit from: (1) expanded descriptions with WHEN/WHY context; (2) documented output schemas; (3) actionable error messages; (4) parameter constraints (enums, ranges, formats). calculate_bmi and fetch_weather are the simplest and cleanest; the PR reviewer tools are the most complex but lack error detail.
Calculate BMI given weight in kg and height in meters
Get file and patch content for the GitHub Pull Request (PR). Only use this tool when you need source materials analyzing or changing the PR content.
Give a non-technical summary and explanation for the GitHub Pull Request (PR). The review is intended for project managers who needs to under the PR at the high level.
Fetch current weather for a city
Create a video podcast for the input title and article. Returns URL links to check the status and download the file video file.
Give a highly technical review for the GitHub Pull Request (PR). The review is intended for software developers. It could give the PR author feedback on how to improve the PR.
Output schemas are not documented. Tools return free-text strings (e.g., review(), explain(), content() return str) without indicating structure, required fields, or downstream chaining opportunities. LLMs cannot plan multi-step workflows or extract structured data from responses.
Descriptions lack actionable context. Most descriptions state WHAT the tool does but omit WHEN to use it and what to do if it fails. E.g., 'Give a highly technical review for the GitHub Pull Request' does not clarify: Is this for code quality, performance, security? When should I choose review() vs explain()? What if the URL is invalid? Current error handling returns plain strings like 'The pr_url is not valid' with no guidance on recovery.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Parameter constraints are absent or minimal. fetch_weather accepts 'city' as free text with no validation (could cause API errors on invalid cities). make_podcast expects 'article' as unbounded text (could exceed LLM token limits or API constraints). review(), explain(), content() accept pr_url but do not constrain format (regex, pattern, or enum). No min/max lengths, no format hints.
Error handling is minimal and non-actionable. Only pr_url parsing returns an error ('The pr_url is not valid'), but offers no recovery path (e.g., 'Did you mean: github.com/owner/repo/pull/123?'). fetch_weather will silently fail or return raw API errors if the API key is missing or the city is invalid. make_podcast has no timeout or failure mode documented. LLMs cannot self-correct from errors.
No support for pagination or result limits. Tools like review(), explain(), and content() may return multi-megabyte responses (full PR diffs, review text) without documenting limits or offering pagination. This risks blowing the context window. The rubric baseline requires capping results and offering pagination for list-like returns.
Secrets are injected via environment variables but not validated or scoped. fetch_weather, make_podcast, review, explain, and content all load credentials from .env files (OPENWEATHERMAP_API_KEY, LLM_APIKEY, etc.). If an LLM or agent logs tool calls, these credentials could leak. No scope declarations (e.g., 'requires read:github') are present to enforce least-privilege.
Tool composition misses chaining context. review(), explain(), and content() all accept a pr_url but do not return structured metadata (owner, repo, pr_number) needed for chaining. An agent that calls review() followed by a hypothetical add_comment(owner, repo, pr_number, text) must re-parse the URL or call a separate lookup tool.
calculate_bmi naming is slightly ambiguous. Should it return the BMI value, a classification (underweight/normal/overweight/obese), or both? Name does not clarify. Description says 'Calculate BMI' but returns a float, no type or range hint in the description.