A web scraping and apartment listing evaluation system for Zurich apartment listings. Includes a FastAPI-based database API for managing apartment listings, a scraper for parsing listings from wgzimmer.ch, and an evaluation workflow triggered via n8n webhooks.
The server exposes 5 tools via FastAPI HTTP endpoints with basic parameter validation. However, critical definition quality issues severely limit LLM usability: (1) Tool names lack action verbs, 'status', 'fetch_unscraped_listings', 'fetch_unevaluated_listings', 'update_listing', 'insert_fresh_listings' are partially verb-forward but inconsistent ('status' is a noun); (2) Descriptions are present but minimal (10-50 chars typically), lacking guidance on WHEN to use each tool, WHAT prerequisites exist, or HOW errors should be handled; (3) Input schemas are visible but incomplete, most tools have empty inputs (status, fetch_unscraped_listings, fetch_unevaluated_listings) with no parameter descriptions; (4) Output schemas are not documented, responses return dict objects but LLMs cannot infer field types or chains; (5) No pagination support despite fetch_* tools potentially returning large result sets; (6) No error recovery guidance, exceptions are caught and logged but error responses do not guide the agent on retry behavior or root cause. The server is HTTP-accessible (good for remote use) but lacks the structured definition quality needed for reliable multi-step agent workflows.
Fetch all listings that have been scraped but not yet evaluated (scraped = 1 and no evaluation exists)
Fetch all listings from the database that have not yet been scraped (scraped = 0)
Insert fresh apartment listings into the database. Duplicates are ignored based on URL uniqueness.
Returns the API status and current timestamp
Update a listing in the database with scraped data and scrape status
Missing output schema documentation. All five tools return dict responses (e.g., {'status': 'ready', 'timestamp': ...}, {'message': '...', 'row_count': ..., 'rows': ...}) but none declare expected return types, field names, or value types. LLMs cannot infer how to chain outputs to downstream tools or extract specific fields.
Descriptions are too brief (10-50 characters). 'Returns the API status and current timestamp' (44 chars) does not explain WHEN to call status vs other tools or what action to take based on response. 'Fetch all listings from the database that have not yet been scraped (scraped = 0)' (81 chars) is better but lacks guidance on pagination limits or expected result size. Baseline for A+ tools is 100+ characters with WHEN/WHY/WHAT prerequisites.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 33 | - | v1 |
Input schemas incomplete. status() and fetch_* tools accept no parameters (empty input dict) but provide no descriptions. update_listing and insert_fresh_listings have typed parameters (listing_id: int, scraped: int, scraped_data: str, etc.) but many lack descriptions. E.g., 'scraped' parameter has description 'Scrape status: 0 = not scraped, 1 = successfully scraped, 2 = scrape failed', good, but 'scraped_data' has description 'JSON string containing the scraped listing data or error message' which does not specify required format or size limits. Baseline: 100% of A+ tools have ALL params described.
No pagination support. fetch_unscraped_listings and fetch_unevaluated_listings return unbounded result sets but provide no limit or offset parameters. If 10,000+ listings exist, returning all rows blows the context window, wastes tokens, and causes LLM reasoning to degrade. Baseline: all list/fetch tools should accept limit (1-100, default 20) and offset/cursor parameters.
No error recovery guidance. fetch_* tools catch exceptions and log them but return raw exceptions to the client ('raise'). Errors do not indicate whether the failure is retryable, user-fixable, or fatal. E.g., if the database is locked, should the LLM retry immediately or wait? If the URL is malformed, should it ask the user to provide a valid one? Baseline: all error responses must classify as retryable|user-fixable|fatal and include actionable next steps.
Tool 'status' is a noun, not an action verb. Naming should be 'check_status' or 'get_status' to clarify that it is a query action. Baseline: 90% of A+ tools start with action verbs (get, list, create, search, update). 'status' alone is ambiguous, does it describe state, require a parameter, or perform an action?
Inconsistent parameter naming between tools. update_listing accepts 'listing_id' (good, unambiguous) but insert_fresh_listings takes 'listings' (an array) without a clear parameter name for individual items. The nested 'web_domain', 'url', 'date_posted', 'scraped' are not aligned with response field names from fetch_* tools (e.g., 'web_domain' vs 'web_domain', OK, but 'rows' returned from fetch_* does not document field names, forcing LLM to infer). Baseline: response field names must match parameter names of downstream tools to avoid mapping confusion.
No input validation error messages. update_listing and insert_fresh_listings accept parameters but do not validate format (e.g., is 'scraped_data' a valid JSON string? what max length?). If LLM passes invalid JSON, the error is a raw exception. Baseline: return 'Invalid scraped_data: must be valid JSON, max 10000 chars. Got: {value}' to enable self-correction.
No composition guidance. fetch_unscraped_listings returns 'id', 'web_domain', 'url'. To call update_listing, the agent needs id, scraped, scraped_data. The fetch response does NOT include a scrape_timestamp or other context the agent might need before scraping. Baseline: ensure fetch_* responses include all IDs and references needed for downstream tools (e.g., include empty 'scraped_data' field with null to signal availability).