Repetitive tool definitions without meaningful differentiation. The 22 103by_* medical specialist tools are nearly identical in structure (city, page, sort_order params) but with minimal variation in descriptions. This violates the single-responsibility and naming clarity patterns, LLMs will struggle to distinguish which tool to call for a specific specialist type.
Consolidate the 22 103by_* medical specialist tools into a single parameterized search_specialist_103by(specialist_type, city, page, sort_order) tool. This reduces API surface, avoids LLM confusion, and follows single-responsibility. Specialist types become an enum: 'oftalmolog', 'lor', 'nevrolog', etc.
Add formal enum constraints to all categorical parameters. For example, in the consolidated 103by tool, define specialist_type enum: ['oftalmolog', 'lor', 'nevrolog', ...]. For avby_search, define engine_type enum: ['petrol', 'diesel', 'hybrid', 'electric'], transmission enum: ['automatic', 'manual', 'robot', 'cvt'], etc. Remove example values from descriptions, use schema enums instead.
Document output schemas for all tools. For each tool, add a section in its code or in API documentation describing the response structure: field names, types, and whether pagination tokens are included. Example: 'web_search returns {results: [{title, url, snippet, date}], total_count, next_cursor}'.
Add error recovery guidance. Update tool descriptions to document expected error scenarios and recovery actions. Example for web_search: 'If rate-limited, wait 60 seconds and retry. If no results found, try broader query terms or search_news for recent events.'
Formalize parameter constraints in descriptions. Replace 'optional' and 'Range 1-100' with precise bounds: 'page: positive integer, 1-1000 (default: 1)'. For numeric params, state min/max explicitly in description text, not just in schema.
Add pagination tokens to paginated tool responses. If web_search returns 20 results, include a next_cursor field if more results exist, or document the max_results limit and how to request subsequent pages.
Score history
Overall score trend
↓ 26 points across a rubric change (v1 → v2)
13/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
13
2026-07-28+
v2
2026-03-09
F
39
-
v1
No output schemas documented. The source code provided does not explicitly define or document what fields these tools return. LLMs cannot plan downstream actions or extract expected data without knowing the response structure. This affects tool chaining and agent reasoning.
Missing enum constraints for categorical parameters. Parameters like 'sort_order' across 103by_* tools accept free-form strings with no documented valid values. The avby_search tool has string parameters (engine_type, transmission, body_type, drive_type, condition) that should be constrained enums but are not. This invites LLM hallucination of invalid values.
Minimal parameter descriptions for 103by_* tools. Parameters like 'city', 'page', and 'sort_order' have generic descriptions ('City name', 'Results page number', 'Sort order') with no guidance on valid values, ranges, or dependencies. For example, sort_order describes one example value ('reviews') but not others or constraints.
No error handling or recovery guidance. The tool descriptions do not explain what errors might occur, what they mean, or how to recover. For example, web_search does not document rate limits, timeout behavior, or fallback strategies. This leaves LLMs without recovery paths when tools fail.
Ambiguous parameter dependencies in avby_search. The tool accepts optional parameters (model, year_min/max, price_usd_min/max, etc.) but does not document how they interact or whether some are mutually exclusive. LLMs cannot reliably construct valid queries without clear dependency documentation.
Lack of pagination guidance for search tools. web_search specifies 'Maximum 20 results per request' but does not document how to get the next page of results or whether a cursor/offset is returned. web_search_batch does not explain how to retrieve additional results beyond the per-query limit.
Generic city/region parameter documentation. The 103by_* tools document city as 'City name' with examples in a list ('minsk, brest, gomel, grodno, vitebsk, mogilev, baranovichi') but do not formally constrain this to an enum. This violates the 'example values in descriptions lead to LLM misuse' principle, LLMs may try to pass undefined cities.
No tool annotation hints (readOnly, destructive, idempotent). None of the tools declare whether they are read-only, destructive, or idempotent. While all tools appear to be read-only based on descriptions, explicit annotations would help agents understand safe retry strategies.
Document parameter dependencies in avby_search. Clarify: 'If both year_min and year_max are provided, year_min must be <= year_max. If price_usd_min and price_usd_max are provided, price_min must be <= price_max.' Alternatively, split into separate tools (e.g., search_cars_by_year, search_cars_by_price) to avoid ambiguity.
Add tool annotation hints to all tools. Use MCP's readOnlyHint: true for all tools (since they are read-only), and idempotentHint: true (since repeated calls with identical params return identical results).
Improve search_news description to explain pagination and supported sites. Example: 'Returns up to 50 news items. Call with site=onliner.by (default) for all categories, or site=tochka.by;smartpress.by for multiple sources. Use page parameter for pagination (not currently documented).'
Add usage guidance to tool descriptions. Example for web_search: 'Use for broad searches, recent events, and fact-checking. For specialized news, use search_news. For deep page extraction, use fetch_page. For medical specialist lookup in Belarus, use search_specialist_103by.'
Document region/language support clearly. For web_search, list valid region codes (e.g., 'ru-by for Belarusian-Russian, us-en for US English') rather than just 'Region/language code (e.g. ru-by, us-en)'.
Enforce result limits. Cap web_search to 20 results per request (already documented), but also document that max_results >20 will be clamped. This prevents LLMs from requesting 1000 results and bloating context.