CitySense exposes a single tool, query_spatial_context, with a clear verb-noun naming convention and reasonable parameter schema. The tool description is concise (45 chars) and matches the baseline, but lacks actionable context for LLM selection (e.g., WHEN to use vs alternatives, WHAT geospatial operation it performs, dependencies). Parameters are well-typed with enums (output_format) and defaults, but descriptions are minimal (11-31 chars each, below the 72-char baseline). Output schema is not documented in the code, the return type dict[str, Any] is opaque and does not tell the LLM what fields to expect. Error handling is absent; no guidance on recovery paths. The implementation delegates to query_spatial_context_impl but that function is not visible, so we cannot verify the full contract. STDIO transport and lack of logging/error reporting capabilities further limit production readiness.
Output schema not documented. Return type is dict[str, Any], leaving the LLM to guess what fields are returned. A documented schema (e.g., 'Returns {results: [{name, bbox, type, ...}], total_count, source}') is required for tool chaining and downstream processing.
Parameter descriptions are too brief (11 - 31 chars vs 72-char baseline). 'Natural language query string' and 'Optional country code' lack actionable constraints. Examples: no mention of supported query types (POI, boundary, raster data?), no hints on city/country interaction, no guidance on max_results memory implications.
Tool description (45 chars) lacks WHEN and WHY context. Contrast with baseline: 'Query spatial context using natural language' (45 chars) vs 'Search for geographic features (buildings, roads, boundaries) by name or description; returns GeoJSON or summary; requires valid city or country' (150+ chars, actionable). Current description does not differentiate this from hypothetical other geo tools.
query_spatial_context
Recommendations
Document the full output schema with an example: 'Returns {results: [{ id, name, type, bbox, geom, source, confidence }], total_count: int, query_used: str, execution_time_ms: int}.' This allows the LLM to plan downstream tool calls or data extraction.
Expand tool description to 100 - 150 chars: 'Query urban geospatial features (buildings, roads, POIs, boundaries) using natural language; resolves city/country context via vector search over OSM, Mapillary, and Sentinel-2 data; returns GeoJSON or summary results (default 5 per query).'
Rewrite parameter descriptions with actionable constraints: 'query (str, required): Natural language description of features to find (e.g., "hospitals in downtown", "parks near the river"). Supports feature type (building, road, POI), spatial relations (near, within, adjacent), and basic properties (name patterns). Max 500 chars.'; 'country (str, optional): ISO 3166-1 alpha-2 code (e.g., fi, se, az). Required if city is empty; if both provided, city takes precedence for bounding box.',
Add error recovery paths to the tool description or as a separate 'Common errors' section: 'If no results found: the query may be too specific or the region has sparse data. Try broadening location (remove city, search by country) or simplifying feature type. If embedding timeout: the vector service may be under load; retry with a shorter query (under 50 chars).'
Clarify output_format semantics: 'output_format (Literal["geojson", "summary", "both"], default="summary"): geojson = full GeoJSON FeatureCollection (use for mapping); summary = text list with name, type, bounds (use for chat); both = combined output (higher token cost).'
No error handling or recovery guidance. The tool may fail for: invalid country codes, empty results, unsupported queries, vector store unavailability, embedding timeouts. No error schema or guidance (e.g., 'If no results, try broadening the geographic region' or 'If embedding fails, retry with simpler query').
Parameter dependencies undocumented. The 'country' and 'city' parameters have implicit relationships (city is more specific, requires country), but the description does not state this. LLMs may pass both, neither, or in conflicting ways.
output_format enum values lack justification. What is the difference between 'summary' and 'geojson' in practice? When should an LLM choose 'both'? No guidance on token cost or use case (e.g., 'Use summary for chat, geojson for mapping apps').
query_spatial_context
Validate max_results at the tool boundary and document limits: 'max_results (int, default=5, range 1 - 50): Number of results to return. Capped at 50 to avoid token explosion and maintain response latency under 5 seconds. Default 5 is typical for chat; increase only if post-processing results is necessary.'
Expose implementation details needed for LLM planning: 'Note: Queries are embedded and matched against a vector index. Natural language is preferred ("hospitals") over structured queries ("amenity=hospital"). Attribute-based filtering (e.g., open_hours, wheelchair_access) is not yet supported; use summary output and post-filter manually if needed.'
Add a dry-run or explanation parameter if appropriate: 'explain (bool, optional, default=false): If true, returns the parsed query intent and search strategy without executing the spatial index lookup (useful for debugging agent behavior).'
Implement structured error responses: Instead of returning a plain error dict, return {error_code: str, error_message: str, recovery_hint: str, available_cities: [str], available_countries: [str]} so the LLM can offer alternatives or retry intelligently.