MCP Server with RAG Pipeline for SEC EDGAR Financial Filings. Exposes tools for semantic search across ingested SEC 10-K filings, year-over-year filing comparison, risk sentiment analysis, financial metrics extraction, and filing ingestion.
The server exposes 7 tools with reasonable naming and some description, but exhibits significant gaps in schema completeness, parameter documentation, and error handling. Tool names follow verb_noun convention (search_*, get_*, extract_*, compare_*, list_, ingest_) which is good for LLM discoverability. Descriptions are present but many are overly technical without explaining WHEN to use each tool or what happens on failure. Input schemas are partially visible in the docstrings but lack formal JSON Schema definitions in the registration code shown. Output schemas are documented in some cases (e.g., FilingComparison, XBRLMetrics) but not consistently across all tools. Error handling is weak, most tools return generic dicts without typed error responses or recovery guidance.
Compare a 10-K section across two fiscal years for the same company. Computes: Cosine similarity (1.0 = identical, lower = more change), Sentiment delta (Loughran-McDonald scores), Newly added and removed risk topics.
Extract structured XBRL financial metrics for a company. Returns revenue, net income, EPS, total assets, operating cash flow and other key metrics from XBRL-tagged financial statements.
Get Loughran-McDonald sentiment scores for a company's risk factors.
Download, parse, chunk, and embed a 10-K filing into the vector store. This is the main ingestion pipeline: 1. Download filing HTML from SEC EDGAR 2. Extract key sections (Items 1, 1A, 3, 7, 8 and their tables) 3. Chunk with structure-aware rules and metadata 4. Embed and store in ChromaDB
List all ingested filings with metadata: ticker, company name, fiscal year, filing date, chunk count.
Search within a specific 10-K section across all companies.
No formal JSON Schema visible in tool registration. Function signatures show parameter types, but no JSON Schema constraints (enum, minLength, pattern, minimum, maximum) are enforced. This forces LLMs to infer valid values from description text alone, increasing hallucination risk.
Output schemas documented in model names (SearchResult, FilingComparison, SentimentScores, XBRLMetrics) but LLMs cannot see the field definitions. model_dump() returns dicts with no schema hints. LLMs don't know what fields to expect or how to chain results to downstream tools.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Semantic search across all ingested SEC 10-K filings.
Error handling is absent. Tools return generic dict responses or log errors internally without surfacing recovery guidance to the LLM. If a query fails, the agent has no way to know if it should retry, try a different tool, or ask the user for clarification.
ingest_filing is a WRITE operation (modifies vector store) but the description does not explicitly state 'This will modify state' or 'This operation has side effects'. Agents need this labeling to know whether a call is safe to retry.
Parameter validation and constraints are documented in text but not enforced. For example, n_results is documented as 1 - 50 but no minimum/maximum is enforced in the schema. Invalid inputs may be silently accepted or cause runtime errors with no user-friendly feedback.
Enum values for section_filter and section parameters are documented in text ('risk_factors, mda, business, legal, financials') but not declared as formal JSON Schema enums. LLMs are more likely to hallucinate variants like 'risk_factor' (singular), 'management_discussion', etc.
get_risk_sentiment returns an error response {'error': '...'} when ticker not found, but other tools' error handling is not visible. Inconsistent error formats (untyped dict vs structured error) make it harder for LLMs to handle failures uniformly.
No pagination support documented for search tools. If search_filings returns many results, there is no offset/limit or cursor mechanism to retrieve additional results. Large result sets could blow the context window.
Descriptions for search_filings and search_by_section do not explain when to use one vs the other. An LLM may be confused about which tool to select when a query mentions both a query and a section.