Read-only access to DataPulse's Malaysian public dataset catalogue (418 datasets, 10-status health taxonomy, licence/attribution metadata).
DataPulse demonstrates strong definition quality with consistently comprehensive descriptions, well-structured schemas, and proper tool annotations. All 5 tools are explicitly defined with complete input schemas. Descriptions are exceptionally detailed and include usage guidance, prerequisites, and rate-limit warnings. Tool naming follows verb_noun convention (search_, get_, find_). Schemas use proper JSON Schema with type definitions, constraints (minLength, minimum/maximum), and examples. All parameters are described. Tool annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint) are properly applied to all tools. Minor gaps: output schemas are not explicitly documented in the visible source; descriptions are verbose (400-600+ chars vs. production baseline of 194 chars p50), which may waste tokens; no batch variants offered for tools agents might call repeatedly.
Return datasets flagged by the latest published anomaly detection (anomalies), ranked by how far the observed update interval exceeds its threshold. Optionally require a minimum publish-reliability grade; includes pipeline-computed anomaly and reliability evidence so agents do not recompute it. Use it for unusual update intervals; do not use it for worsening freshness, recovery, reliability grades, or structural drift—use find_deteriorating, find_recovering, find_unreliable, or find_schema_drift instead. It reads precomputed anomaly data, so an empty result means no published row survives the selected filters; DataPulse is read-only, requires no API key, and the edge limits clients to roughly one request per second with a small burst, so pace or retry.
Return datasets whose status is aging, stale, or degraded, plus datasets missing from the latest health snapshot. Use when an agent needs to know which data has a freshness or schema-validity risk. Use it to enumerate freshness or schema-risk candidates; do not use it for anomalies, trends, reliability, or drift—use find_anomalies, find_deteriorating, find_unreliable, or find_schema_drift instead. It reads the published health snapshot, so an empty result means no rows met this snapshot-based rule; DataPulse is read-only, requires no API key, and the edge limits clients to roughly one request per second with a small burst, so pace or retry.
Return one bounded, machine-readable Dataset Passport v1 for a canonical dataset ID. It reads the published Passport artifact only; it does not fetch an upstream source or create evidence. The Passport describes observed metadata and evidence availability, not semantic truth, completeness, certification, legal permission, safety, or AI admission. Use it for the bounded Passport artifact; do not use it for current health detail or citation context—use get_dataset or get_provenance instead. It reads precomputed published data, and evidence_available=false identifies an unavailable or unsupported Passport; DataPulse is read-only, requires no API key, and the edge limits clients to roughly one request per second with a small burst, so pace or retry.
Output schemas not explicitly documented in visible source code. Tool descriptions reference returned fields (e.g., 'Returns ranked matches with id, title, source, licence, published status, and score') but formal output schema definitions are not visible in mcp.json or server.py excerpt.
Descriptions are verbose (400 - 600+ characters, exceeding production baseline of 194 chars p50). While detailed and LLM-friendly, they waste tokens on guidance that could be condensed. E.g., rate-limit and retry boilerplate appears identically in all 5 tool descriptions.
No batch variants for tools agents might call in loops. E.g., if an agent needs to verify multiple dataset IDs, it must call get_dataset N times serially instead of one batch_get_datasets call.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 89 | 2026-07-28+ | v2 |
Return full detail for one dataset id, including its latest health status and last-verified timestamp, content_freshness_date, and freshness_signal_source (last_modified, content_parse, or none). Use to fetch the provenance/citation metadata for a dataset found via search_datasets and distinguish unknown-freshness from proven stale data. Use it for one dataset's current published detail; do not use it for citation context or a Passport—use get_provenance or get_data_passport instead. It reads published data, so an absent health row is reported as unknown rather than probed live; DataPulse is read-only, requires no API key, and the edge limits clients to roughly one request per second with a small burst, so pace or retry.
Use for discovery only: find DataPulse's 418 Malaysian public datasets by topic, source, or licence—for example, 'Malaysian public data inflation', licence and attribution, or a government dataset source. Returns ranked matches with id, title, source, licence, published status, and score. This is not trust verification: a status is published pipeline context, not proof that a dataset is current or reliable. For pre-trust use search_datasets → verify_dataset → get_provenance. Use it to discover candidates; do not use it for a trust decision—use verify_dataset instead. It reads published catalogue data, so no match means the published catalogue has no matching entry; DataPulse is read-only, requires no API key, and the edge limits clients to roughly one request per second with a small burst, so pace or retry.