Read-only MCP server exposing scraped HKEx (Hong Kong Stock Exchange) regulatory filings corpus to LLM clients. Serves filings data from configured database sinks over stdio.
The server defines 9 well-named tools with clear action verbs (get_, list_, search_, describe_). All tools have descriptions between 80-250 characters, meeting the 10-1024 character baseline. All visible parameters include type declarations and descriptions. Schema structures are explicit (JSON Schema format with types). Key strengths: consistent naming (verb_noun pattern), paginated results with offset/page_size, clear parameter constraints (max_text_chars, page_size limits stated in descriptions). Key gaps: output schemas are not documented in the source code (no response type definitions visible), error handling is mentioned but not detailed in tool definitions, no enum constraints for categorical parameters (filing_type, document_status, etc. described as comma-separated strings rather than enums), and tool annotations (readOnlyHint) are not declared despite all tools being read-only.
Describe the schema of the configured database sink including tables, columns, data types, and sample counts
Get the current configuration including DATABASE_TARGET, sink IDs, read sink, company table settings, and download worker count
Retrieve a single filing by ID with optional full text, tables, and text windowing. Text can be paged using text_offset and max_text_chars.
Retrieve multiple filings by IDs with optional text and table content, using the same text windowing as get_filing.
Get server metadata including name, version, transport type, configured sinks, and capabilities
Get aggregated statistics on filings by ticker, filing type, document status, or year. Returns counts and coverage information.
Output schemas are not documented. Tool descriptions state what is returned (e.g. 'Returns paginated filing IDs and metadata', 'Returns paginated document chunks') but no structured schema documentation is visible in the source code. LLMs cannot reliably extract response fields for downstream tool composition without explicit schema definitions.
Categorical parameters are free-form strings rather than enums. Parameters like 'filing_type', 'document_status', 'order_by', 'group_by' are described as comma-separated values (e.g. 'e.g., ANNOUNCEMENTS, CIRCULARS') but not declared as enums. This invites LLM hallucination of invalid values like 'ANOUNCEMENTS' or 'STATUS_PENDING' instead of enforcing valid options.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 72 | <=2025-11-25 | v2 |
List all available database sinks with their configuration status, availability, capabilities, and licensing information
Full-text search across extracted document text in the corpus. Returns paginated document chunks with text excerpts and metadata.
Search for filings using filters such as ticker, stock code, filing type, document status, date range, and full-text title/text queries. Returns paginated filing IDs and metadata.
Tool annotations are missing despite all 9 tools being read-only (Risk: READ_ONLY). The MCP spec supports readOnlyHint, destructiveHint, and idempotentHint annotations. Without readOnlyHint, agents and clients cannot distinguish safe read-only tools from write tools without re-parsing descriptions.
Error handling guidance is absent from tool descriptions. The code mentions 'errorReporting: true' but no tool description explains what errors can occur, how to recover, or what to do if a query returns no results. E.g., search_filings should note 'Returns empty list if no filings match filters; try broader filters or different date range.'
Date format constraints are implicit. Parameters 'date_from' and 'date_to' state '(ISO format: YYYY-MM-DD)' in descriptions but do not declare a JSON Schema format: 'date'. This makes validation ambiguous and increases LLM mistakes (e.g. passing '2024/01/15' instead of '2024-01-15').