Production-ready MCP server for accessing Indian government open data from data.gov.in
This server has moderate definition quality with uneven tool design. Tool names follow verb-first convention (get_, filter_, clear_, paginate_) which is good. However, descriptions are present but generic, parameter descriptions lack specificity about constraints and ranges, and output schemas are not documented. Several tools conflate related concerns (e.g., get_dataset combines filtering, pagination, and field retrieval into one tool). The server returns raw JSON strings rather than structured typed responses, making it harder for LLMs to extract and chain data. Error handling is basic, errors are returned as JSON strings with minimal recovery guidance. No tool annotations (readOnlyHint, destructiveHint) are present despite clear read/write distinction. Parameters lack explicit constraints (enums, regex patterns, min/max bounds) that would guide LLM input validation.
Clear all cached API responses
Filter dataset records by a specific field value
Get statistics about the API response cache
Retrieve data from a specific dataset/resource on data.gov.in
Get field information and schema for a dataset
Get a summary of a dataset including record count and field information
Get information about the MCP server and its configuration
All tools return raw JSON strings (json.dumps()) instead of structured typed responses. LLMs must parse unstructured text rather than working with native types. This violates pattern:response-shaper and increases token waste and hallucination risk.
Output schemas are not documented. The rubric (section D, SCHEMAS & OUTPUT) requires documentation of what fields LLMs should expect so they can plan downstream calls. Currently, LLMs see only a JSON string and must infer structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Retrieve a specific page of data from a dataset
Parameters lack explicit constraints (enums, min/max, regex patterns). For example, 'limit' has a description saying 'max: 100' but no JSON Schema maxItems constraint. 'filters' accepts arbitrary JSON but no schema validates its structure. This invites invalid LLM input that fails at runtime.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are absent. The server clearly distinguishes read-only tools from reversible ones (clear_cache is REVERSIBLE, others are READ_ONLY), but this information is not communicated via MCP annotations. Current spec supports tool annotations; they should be used.
Error handling returns JSON strings with minimal recovery guidance. When 'filters' JSON parsing fails, it returns '{"error": "Invalid filters format..."}' as text. Per pattern:recovery-guide, errors should guide the LLM on next steps (e.g., 'Try passing valid JSON like {"state": "Maharashtra"}').
Descriptions are present but generic. 'Get field information and schema for a dataset' (get_dataset_fields) is ~50 chars, adequate but lacks context about WHEN to call it vs get_dataset. Per pattern:tool-description, descriptions should indicate selection logic: 'Call this first to explore field names before filtering or aggregating.'
Unclear responsibility boundaries between tools. get_dataset, paginate_dataset, and filter_dataset all retrieve subsets of data but in different ways. The LLM must reason about which to use. Per pattern:tool, each should do exactly one thing. Consider: does paginate_dataset add value vs get_dataset with limit/offset? Does filter_dataset add value vs get_dataset with filters param?
Parameter 'filters' in get_dataset accepts a JSON string but does not validate its structure. The description says 'e.g. {"state": "Maharashtra"}' but does not specify which fields are filterable or what types they accept. An LLM might pass invalid field names that fail silently at the API layer.
No batch variants. If an agent needs to fetch summaries for 10 datasets, it must call get_dataset_summary 10 times sequentially. A batch_get_dataset_summaries(resource_ids: list[str]) variant would reduce token waste and latency.