MCP server for searching and retrieving datasets from Peru's National Open Data Platform (PNDA - Plataforma Nacional de Datos Abiertos)
PNDA MCP provides 2 tools with reasonable descriptions and basic parameter documentation, but exhibits significant gaps in schema completeness, error handling, and output structure. Both tools have clear semantic intent (search and detail retrieval) and descriptions exceeding 100 characters, which is good. However, input schemas lack formal type constraints (enums, ranges), output schemas are undocumented, and error handling is minimal or absent. The code shows bare exception handling with no recovery guidance. Parameter descriptions exist but lack detail on constraints, ranges, and format expectations. No input validation is visible. The overall design is functional but falls short of production-grade agent-tool standards.
Get detailed information about a specific dataset from the PNDA (Plataforma Nacional de Datos Abiertos) Peru. Use this when you have an dataset ID from search results and want to retrieve the full content and metadata for that specific dataset from Peru's national open data platform (datosabiertos.gob.pe).
Search for relevant content (datasets) from the PNDA (Plataforma Nacional de Datos Abiertos) Peru. Find information from Peru's national open data platform (datosabiertos.gob.pe) that semantically matches your search terms. Use this when you need to discover content related to a topic, concept, or question. The search understands meaning and context, not just exact word matches.
Output schemas are completely undocumented. dataset_search returns {query, results} and dataset_details returns {id, text, resources, error} but no schema is visible in code or docstring. LLMs cannot reliably extract or chain results without knowing the structure.
Input parameter 'top_k' in dataset_search lacks formal constraints. Description says 'max 25' but no JSON Schema maximum is enforced. LLMs may pass invalid values (negative, >25). Should define 'minimum': 1, 'maximum': 25 in schema and validate in code.
Error handling is bare-minimum. dataset_details has try-except blocks that silently pass without returning actionable error messages. A user/LLM receives only {error: 'Item with ID ... not found'} with no recovery guidance (e.g., 'Try dataset_search() to find valid IDs').
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
dataset_details function has three fallback paths (cache → PNDA API → Pinecone) with inconsistent error handling. Some paths return silently (pass after except) without raising or logging. This breaks auditability and prevents LLMs from knowing what went wrong.
Parameter descriptions lack actionable detail. 'query' is described as 'What you want to search for' but does not specify format constraints, length limits, language expectations (Spanish? English?), or what semantic search means. Similar issue with 'id' parameter in dataset_details.
dataset_search returns relevance 'score' but no documentation of its range, meaning, or interpretation. Is it 0-1? 0-100? Higher is better? LLMs cannot reliably rank results or filter by threshold without this info.
No pagination support in dataset_search despite returning results list. If semantic search returns 10 items by default, what if the user wants more? There is no 'next_cursor' or 'total_count' in the response, and no offset/limit parameters beyond top_k.
dataset_details returns 'text' (title), 'resources' (list of files), and optional 'metadata'/'error' fields inconsistently. Some paths return {id, text, resources}, others return {id, metadata, text}. Inconsistent response structure breaks LLM parsing and chaining.
No validation or sanitization of the 'id' parameter in dataset_details. If an LLM passes a malicious string (SQL injection, path traversal), the code constructs URLs and cache keys without escaping. The requests.get() call is URL-safe but cache key usage is not.
Tool names 'dataset_search' and 'dataset_details' are reasonable but could be more action-focused. 'search_pnda_datasets' and 'get_pnda_dataset' would be clearer (matches get_, search_ verb pattern baseline). Current names are borderline acceptable but not ideal.