Dynamic RAG Engine for AI Reliability. We provide mathematically scored context & sanitized data to prevent hallucinations in both static & volatile domains (starting with Korean Finance).
Scoring was not performed
Output schemas not formally documented in MCP tool registration. Tool response types exist as TypeScript interfaces in source code (e.g., chart.ts lines showing `success`, `data` with nested structures) but are not exposed in the MCP schema that LLMs receive. LLMs cannot reliably plan chained calls or extract specific fields without knowing response structure.
Missing when-to-use guidance in tool descriptions. Descriptions like '주식 차트 조회' (stock chart query) and 'Research 보고서 조회' are minimal (5-30 chars) and do not explain when to call this tool vs. similar ones. For example, getNews and getNewsScored both return news but differ only by sentiment scoring, the distinction is not clear from descriptions. LLMs will waste reasoning cycles deciding between them.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 16 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
No error recovery guidance. Tools have no documented error handling or recovery strategies. If a call fails (e.g., invalid ticker code), no guidance tells the LLM what to try next (search_tags to validate? Retry? Try a different market?). Baseline expectation: 'If ticker not found, call search_tags() first to validate the code.'
Parameter descriptions use example values and non-machine-parseable formats. E.g., 'tag_code (search_tags 결과값, 예: STK005930, THM001)' embeds examples in text rather than enforcing via pattern validation. 'from_date (YYYY-MM-DD)' lacks type:string format constraint in schema. LLMs latch onto example values and may pass them literally in edge cases.
Overlapping tool functionality without clear differentiation. getNews (points excluded) and getNewsScored (points included) are nearly identical. searchTags and matchTags both operate on tag data but with different semantics (search by name vs. NER extraction). Tool descriptions do not explain the trade-offs or composition logic.
Missing natural-language identifier support. getNews and getFinancials expect 'tag_code' and 'ticker' as system identifiers (STK005930, THM001). No description or parameter guidance explains how users/agents obtain these codes or if human-readable names (e.g., 'Samsung Electronics', 'Semiconductor') are accepted. Forces callers to invoke searchTags first, breaking the one-intent-one-call pattern.
Parameter validation rules documented only in descriptions, not in schema. Examples: 'limit' has min/max constraints (1-100) defined in Zod, but 'days' (min 1, default 7, no max) has no explicit upper bound documented. DateString formats (YYYY-MM-DD) appear as hints, not JSON Schema format validators. LLMs cannot programmatically enforce these constraints.