A declarative web scraping and data extraction engine with support for HTTP requests, headless browser automation, and structured data parsing via JSON/HTML/XML/XPath selectors. Exposes tools via MCP for Claude and other AI models.
Fitter demonstrates solid definition quality with well-named tools following verb_noun patterns and comprehensive descriptions. All 6 tools have clear, substantial descriptions (194-380 chars average, well above the 10-1024 baseline). Tool names are action-oriented (fitter_run, fitter_validate_config, fitter_inspect_url) and distinguish clearly from one another. Input schemas are present and structured with JSON Schema format. However, critical gaps exist: (1) Parameters lack type declarations in the visible schema definitions, the code shows parameter names and descriptions but no explicit 'type' field like 'string', 'boolean'; (2) Output schemas are not documented, tools like fitter_run and fitter_run_file return 'extracted data as JSON' but the structure of that JSON is never specified; (3) Error handling and recovery guidance is absent, tools do not document what errors may occur or how to recover. These gaps prevent downstream tool chaining and force LLMs to infer output structure. The fitter_inspect_url tool is well-designed for discovery, but its relationship to fitter_run is underdocumented, the description says 'use it before fitter_run' but does not explain which output fields map to which config input parameters.
Return the Fitter config schema reference and documentation.
Fetch a URL and return a compact structure outline plus candidate selectors/paths, so you can author a fitter config that matches on the first try instead of guessing selectors and getting nulls. For JSON it lists gjson paths with types and sample values; for HTML it lists repeated elements (candidate array_config root_path / list rows) and link/heading selectors. For client-rendered SPAs (content built by JavaScript), a plain fetch sees only an empty shell — the output warns when it detects one; pass render:true to render it in a headless browser first (mirrors what a browser_config scrape would see). Read-only helper that does NOT extract data — use it before fitter_run, then fitter_run to actually extract.
Run a Fitter scraping/parsing config passed inline (JSON or YAML) and return the extracted data as JSON. Fitter fetches data via a connector (HTTP request, headless browser, static value, file, ...) and extracts structured data using json/HTML/XML/xpath selectors described by a declarative model. Call fitter_config_reference first if you are unsure about the config format.
Run a Fitter scraping/parsing config from a local JSON or YAML file and return the extracted data as JSON. Same as fitter_run but reads the config from disk.
Output schemas completely undocumented across all tools. Tools like fitter_run and fitter_run_url state they 'return extracted data as JSON' but the structure of that JSON is never formally specified. This forces LLMs to guess at output field names and prevents confident downstream tool chaining.
Error handling and recovery guidance absent. No documentation of what errors each tool can return (e.g., config parse error, network timeout, SPA render failure, file not found) or how to recover. Tools like fitter_run and fitter_run_url are brittle from the agent's perspective.
Tool composition chain not clearly documented. fitter_inspect_url description says 'use it before fitter_run' and fitter_run says 'Call fitter_config_reference first', but the mapping from fitter_inspect_url's output fields to fitter_run's config input parameter is not explained. An LLM cannot confidently compose these tools without trial and error.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 72 | 2026-07-28+ | v2 |
Run a Fitter scraping/parsing config downloaded from an HTTP(S) URL (JSON or YAML) and return the extracted data as JSON. Same as fitter_run but fetches the config from a remote location, e.g. a raw GitHub link.
Validate a Fitter config (JSON or YAML) without executing it. Checks the structural rules: item/connector_config/model presence, valid response_type, that the connector has a data source, and compiles every condition/item_condition expression in the model. Returns "valid" or the validation error. Cheap and safe — use it while iterating on a config before calling fitter_run.
fitter_config_reference description is brief (73 chars, below baseline of 194 avg) and lacks guidance on WHEN/WHY to call it. Stated as a dependency in fitter_run but the relationship is unclear. A better description would explain: 'Call this first to understand valid config keys (item, connector_config, model, limits, references) and their syntax before authoring a fitter_run call.'
Optional parameters lack clear guidance on defaults and constraints. For example, fitter_inspect_url's 'response_type' parameter says 'Optional hint' but does not explain what happens if omitted (auto-detect from Content-Type), nor does it list valid enum values (json, HTML, xpath, XML). Similar issue with 'render' boolean, no guidance on performance/latency trade-offs or what SPA scenarios require it.