MCP server for DataSinking — full-text Asian financial reports (China, Korea, Japan, Taiwan) as clean Markdown, for Claude, Cursor and other AI agents.
DataSinking presents well-structured tool definitions with consistent naming patterns (all verb_noun format), comprehensive descriptions, and properly typed input schemas. All six tools follow the list_* / get_* / get_* pattern for discoverable operations. Descriptions are detailed and mention important context (e.g., 'Expensive in tokens, prefer get_section', 'keep that attribution when you cite it'). Schemas use proper JSON Schema with type declarations and required fields. However, output schemas are not documented, callers cannot know what fields to expect in responses, forcing LLMs to guess. No tool annotations (readOnlyHint, idempotentHint) despite all being READ_ONLY operations. Error handling guidance is minimal, no recovery hints for 404s or invalid queries. Parameter descriptions are generally good (e.g., 'FMP-style symbol' with examples, enum constraints on doc_type), but some lack depth (e.g., 'Heading keyword' in get_section could clarify substring matching more explicitly in the description text, though it does mention it). Missing: return type documentation, error taxonomy, security notes on API key injection patterns.
Fetch one report's full text (metadata + Markdown body). The `source` field names the official disclosure platform; keep that attribution when you cite it. Expensive in tokens — prefer get_section when you only need one chapter.
Fetch only one section of a report by keyword — much cheaper than get_report, best for RAG.
List the exchanges DataSinking covers. Returns exchange codes (sse / szse / bj / ksc / koe / knx / jpx / twse / tpex) with the number of reports available per exchange. Call this first to discover coverage. Sources: A-shares = cninfo.com.cn, Korea = DART, Japan = EDINET, Taiwan = MOPS.
List a company's reports — metadata only (id, title, period), no body text. Each item carries a `source` field naming the official disclosure platform; keep that attribution when you cite it.
List every section of a report with its size, before you decide what to pull. Returns `sections` (titles, in order) plus `section_details` — same order, one entry per section with `title`, `has_tables`, `chars` and `estimated_tokens`. Use `estimated_tokens` to avoid pulling a chapter that would blow your context, and `has_tables` to know whether the chapter needs special handling (tables are the part RAG pipelines usually get wrong). Then call get_section with a heading keyword — the headings are in the report's own language.
No output/return schemas documented. All six tools lack specification of what fields and structure callers should expect in responses. LLMs cannot plan downstream operations (chaining) without knowing what data is available.
No tool annotations (readOnlyHint, idempotentHint, destructiveHint) despite all tools being READ_ONLY operations. Annotations help clients understand safety guarantees and retry behavior without parsing descriptions.
Error handling and recovery guidance minimal. No documentation of what errors each tool can return (e.g., 404 for invalid symbol, invalid section keyword match) and what LLMs should do next. Particularly important for get_section which explicitly states 'returns 404 with the real headings' but no guidance on how to handle this.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 79 | 2026-07-28+ | v2 |
List the stocks on one exchange, including the report count per company.
Pagination not documented for list_* tools. No mention of whether results are capped, whether there's a next_cursor or offset/limit mechanism, or what the maximum returned count is. list_reports defaults to size=10 and list_stocks to limit=20, but it's unclear if these are actual response sizes or just defaults.
Description for 'section' parameter in get_section is very long and complex. While technically accurate (explains substring matching, language handling, fallback), it could be more concise. Current form: ~220 chars, exceeds guideline of 10-1024 per-parameter (top-level guideline; per-param ~72 chars average in production). Truncation in provided code makes full assessment difficult.