MCP server for crawling and aggregating scholarship foundation and welfare service information from Korean websites using Gemini search, Playwright rendering, and Ollama filtering
This server has severe definition quality issues across all dimensions. While tools are registered with FastMCP, the schema definitions are minimal or absent, descriptions lack clarity and guidance, and there is no structured output documentation. The server implements complex web scraping and LLM integration logic but exposes these as raw tools without proper agent-friendly wrapping. Most tools operate on unstructured text with no validation guidance or error recovery patterns. The average tool score is 28/100, reflecting systematic gaps in naming clarity, parameter documentation, and schema completeness.
Crawls provided URLs using Playwright to extract rendered content, with configurable depth control
Generates a refined title and categories for welfare service from summary text
Searches for Korean scholarship foundations using Google Search tool via Gemini and returns URLs as JSON array
Generates structured summary of welfare service information from crawled content
Verifies if crawled content is actually valid scholarship or welfare service information using local LLM
search_sites_with_gemini has empty input schema ({}). Tool performs critical search operation but provides no guidance on expected output, pagination, or failure modes.
All tools lack documented output schemas. LLMs cannot predict return structure, forcing them to guess whether output contains pagination, status fields, or error indicators. No baseline output structure visible in code.
Parameter descriptions are minimal or missing. 'urls' in crawl_from_search has basic text but no format guidance (is this http/https only? max items? max URL length?). 'max_depth' has no bounds (1-10? 1-100?). Forces LLM to guess valid ranges.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 31 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 6 | - | v1 |
Tool naming lacks clarity for agent composition. 'verify_crawled_info' suggests binary validation but actually runs inference via Ollama LLM. Name does not convey that this is an LLM-based filter, not a schema validator. Similarly, 'summary_info' is vague, does it extract, summarize, or classify? Should be 'summarize_welfare_service' or 'extract_service_summary'.
No error handling guidance. Code catches exceptions (e.g., Playwright timeouts, Ollama failures) but tool descriptions do not tell LLM what to do on failure: retry? Skip? Request help? Error responses are logged but not structured for agent recovery.
Parameter types are present but incomplete. 'urls' is declared as array of strings, but 'max_depth' integer has no min/max bounds.
Tool descriptions under 50 characters for 'summary_info' and 'generate_title_and_category'. These are too short to guide LLM on when to invoke vs similar tools or what output format to expect.
No schema constraints for free-form text inputs ('title', 'snippet', 'summary'). No max length, no format hints, no guidance on expected language (Korean? English?). Ollama prompt assumes Korean but tool description says nothing.
No pagination or result limits documented. Code caps crawl results at 1500 chars per snippet but tool description does not state this. crawl_from_search can return hundreds of results; no documented limit or next_cursor field for large result sets.
Inconsistent naming across the pipeline. Tools use verb_object pattern ('search_sites', 'crawl_from', 'verify_crawled') but lack consistency in what they operate on. 'crawl_from_search' suggests crawling results of search, but 'verify_crawled_info' suggests verification of crawled content, both are unclear about the data flow.