Layer B: High-performance Data Normalization Engine (Stateless) - Extracts, cleans, structures, and reconciles unstructured web content for AI agents with strict financial compliance enforcement
Agent-Commerce-Core exhibits mixed definition quality across 9 tools. Tool naming follows action-verb conventions reasonably well (normalize, search_web, extract_via_jina, etc.), but several tools suffer from duplicate names (two 'normalize' tools with different responsibilities), missing or incomplete parameter descriptions, and inadequate output schema documentation. Descriptions vary from detailed (normalize) to minimal (verify_gateway). Parameter schemas are present in most cases but inconsistently documented. Error handling is largely absent from tool definitions, no guidance on retryability, classification, or recovery steps. The codebase shows effort toward structure (Pydantic models, FastAPI validation), but the tool interface itself lacks the rigor expected for production agentic use.
Calculates a deterministic hybrid trust score combining LLM self-evaluation with infrastructure metrics (extraction route stability, data freshness). Returns clamped float between 0.0 and 1.0.
Validates text against strict non-financial compliance policy. Raises HTTPException if forbidden financial/trading terms are detected. Acts as a kill-switch for financial data processing.
Extracts content from a URL using Firecrawl API. Returns markdown-formatted extracted content or None if extraction fails or API key is not configured.
Extracts content from a URL using Jina Reader API. Returns markdown-formatted extracted content or None if extraction fails or API key is not configured.
Executes the core fallback strategy and normalization logic. Fetches and normalizes web content through a cascade of extraction methods (Jina Reader → Firecrawl → Tavily Web Search), enforces compliance rules, and returns normalized data with trust scores.
Duplicate tool name 'normalize', two distinct tools register with the same name (normalize from gemini_normalizer.py endpoint and GeminiNormalizer.normalize() from app/tools/gemini_normalizer.py). LLMs cannot disambiguate and will select randomly or fail.
Two extraction tools (extract_via_jina, extract_via_firecrawl) perform nearly identical operations (fetch content from URL, return markdown). The tool descriptions do not clarify when to use one vs. the other (fallback order, quality expectations, API key availability, latency). This violates single-responsibility and creates LLM confusion.
Output schemas are not formally documented in tool definitions. Descriptions mention return values (e.g., 'Returns markdown-formatted extracted content', 'Returns clamped float between 0.0 and 1.0'), but the exact response structure (field names, types, nested objects) is not specified. LLMs cannot reliably map downstream tool inputs to upstream outputs without explicit schemas.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
GeminiNormalizer.normalize: Executes LLM-based normalization process using Google Gemini. Returns extracted data as string (Markdown or JSON), success status, and metadata including trust scores, inference time, and model information.
Executes Tavily API Deep Search. Returns structured JSON containing raw_content, urls, and a UTC timestamp (fetched_at) for upstream X-Data-Freshness-Seconds calculation.
FastAPI Dependency: Ensures the request is genuinely routed through Layer A (Gateway). Validates internal secret token and returns tenant_id and trace_id for logging/isolation purposes.
Validates if the target URL is reachable before delegating to extraction APIs. Performs HEAD and GET requests to verify HTTP status and accessibility.
Error handling guidance is absent from all tool definitions. No description of when/how tools fail, which errors are retryable, or what the LLM should do if a call fails. Example: verify_url_is_live does not state whether a 404 is retryable, whether it triggers a timeout exception, or what status codes are acceptable. enforce_compliance raises HTTPException but does not document the error format or recovery options.
Parameter descriptions are inconsistent and sometimes vague. Example: 'target_tier' is described as 'Specifies the extraction tier. Options: standard, tier_a1, tier_a2, tier_a3' but the semantic difference between tiers is not explained. 'format_type' supports 'json' and 'markdown' but does not clarify output structure (is JSON a specific schema or arbitrary?). 'webhook' object has no documentation of internal properties beyond url and secret_token.
No tool declares data freshness semantics, rate limits, timeouts, or quota constraints. LLMs have no guidance on whether a tool is safe to call in a loop, what the max QPS is, or how stale data is acceptable. This invites runaway agent behavior and integration failures.
Tool descriptions lack context on prerequisites and dependencies. Example: normalize() mentions a 'cascade of extraction methods' but does not state whether the caller must invoke verify_url_is_live first, or whether it's called automatically. Similarly, enforce_compliance is described as a 'kill-switch' but the tool definitions do not indicate which tools depend on it.