Financial news and data crawler with MCP protocol support, providing web scraping, structured data extraction, and financial data access via FastAPI and MCP
FNewsCrawler exposes two web-scraping tools via FastMCP with complete input schemas and descriptions. Both tools have clear, well-structured parameters with proper typing. However, there are notable gaps in error handling guidance, output schema documentation, and composition patterns. The tools follow a verb-noun naming convention but lack the granularity expected for production-grade agent tooling. Parameter descriptions are adequate (72 - 150 chars) but could be more LLM-optimized. Output schemas are not explicitly documented, forcing agents to infer structure. Error handling is minimal, no recovery guidance or error categorization. Security is reasonable (no credentials exposed as parameters), but logging/audit trails are not visible in the tool definitions. The server supports HTTP transport (via FastMCP framework) and includes cancellation support, which are positives for protocol readiness.
从指定URL和CSS选择器中提取结构化信息。支持多种提取类型:text(提取文本内容)、html(提取HTML内容)、attribute(提取指定属性)、mixed(混合提取文本+HTML+属性)
使用pandas实现的成熟表格数据提取函数,支持高级数据清洗和格式化。pandas选项说明:pandas_attrs(JSON格式属性筛选)、pandas_match(根据文本内容匹配表格)、pandas_header(指定表头行号)、pandas_skiprows(跳过的行数)、pandas_na_values(转换为NaN的值)
Output schemas are not documented. Agents cannot plan downstream calls or extract relevant data without knowing what fields to expect (e.g., does extract-structured-data return a list or object? What fields does it contain?)
Error handling lacks recovery guidance. No indication of which errors are retryable vs user-fixable vs fatal. No error messages guide the LLM on next steps (e.g., 'Selector not found. Verify the CSS selector matches the page structure.')
Parameter relationship complexity not documented. extract-table-data has interdependent pandas_* parameters (pandas_header, pandas_skiprows, pandas_na_values) whose valid combinations are unclear. No guidance on which combinations are safe or recommended.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 31 | - | v1 |
Enum constraints missing for string parameters. extract_type should be an enum (text|html|attribute|mixed), not a free-form string. This invites hallucinated values and runtime failures.
Timeout and wait semantics could cause context window exhaustion. wait_timeout (10000ms default) on both tools may cause agent to hang. No guidance on what happens when timeout expires or how to handle partial results.
Tools are not idempotent by definition. Repeated calls with same parameters on dynamic pages may return different data. No guidance on idempotency or whether agents should treat results as stable.
No result limits documented. Both tools could return unbounded data (e.g., extract-structured-data with multiple=true on a page with thousands of matching elements). No cap or pagination guidance risks context window overflow.
Naming ambiguity: 'context_name' is opaque. The description says 'browser context名称,用于会话管理' but does not explain what this controls, what defaults exist, or how to discover valid context names. Should be 'browser_context_id' or 'session_name' with clearer semantics.