Example of deploying a MCP server on Cloud Run for resume screening and candidate matching using LlamaCloud and OpenAI
Server has 7 tools with generally clear descriptions and basic schemas. Naming follows verb_noun patterns well (extract_*, find_*, search_*, score_*). However, several critical issues reduce quality: (1) Three math utility tools (add, subtract, multiply) are out of domain and dilute focus; (2) Parameter descriptions lack constraint details (e.g., no mention that top_k max is 50 in descriptions, only in schema defaults); (3) Output schemas are not formally documented, tools return JSON strings with unstated structure; (4) No error handling guidance in descriptions; (5) Tool composition issues: find_matching_candidates and search_candidates_by_skills overlap significantly; (6) Missing idempotency guidance; (7) No permission/scope declarations. Schemas are present and reasonably typed, but descriptions are 10-20 tokens shorter than production baseline (194 chars avg). Baseline comparison: average tool description ~150 chars (below 194 baseline); parameter descriptions present but generic.
Add two numbers together.
Extract structured job requirements from job description text.
Find candidates matching job qualifications from LlamaCloud resume index.
Multiply two numbers.
Score a candidate's resume against specific job qualifications using LLM evaluation.
Search candidates by specific skills or keywords from LlamaCloud resume index.
Subtract two numbers.
Math tools (add, subtract, multiply) are out of domain. This is a resume screening MCP server, utility math functions dilute focus, confuse agents about intent, and should be removed or moved to a separate utilities server.
No documented output schemas. All tools return JSON strings (e.g., find_matching_candidates returns {"candidates": [...], "search_parameters": {...}}), but the response structure is not formally declared. LLMs cannot reliably parse unstated output formats. Document the response schema for each tool.
Parameter descriptions lack constraint clarity. E.g., 'top_k' description says 'default: 10, max: 50' but does NOT state min/max in the description text itself, these are only in schema default/constraints. Descriptions should include 'must be between 1 and 50' explicitly so LLMs understand the constraint without parsing schema.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 74 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Tool overlap and composition issues. find_matching_candidates and search_candidates_by_skills both retrieve candidates from LlamaCloud but differ only in search criteria (qualifications vs skills). This duplicates logic and may confuse agent selection. Consider a single unified search_candidates tool with a 'search_mode' parameter (qualifications | skills).
No error handling guidance in tool descriptions. Tools return error JSON objects (e.g., {"error": "..."}) but descriptions do not document recovery paths. E.g., extract_job_requirements should state: 'Returns error if input is empty or malformed, try with a valid job description text.'
Missing security/scope declarations. No tools declare permissions (e.g., 'read:resumes', 'write:matches'). LlamaCloud service requires authentication but there is no description of what permissions the agent needs or how they are enforced.
No idempotency guarantees. Tools like score_candidate_qualifications and find_matching_candidates do not document whether repeated calls with the same input are safe (idempotent) or may produce different results (non-deterministic). This is critical for agent retry logic.
Description length below production baseline. Tool descriptions average ~150 chars; production baseline is 194 chars (p90=392). Descriptions like 'Add two numbers together' (29 chars) and 'Subtract two numbers' (21 chars) are too terse to guide LLM selection. Expand with WHEN and WHY context.