MCP server for test intelligence. Analyzes JUnit XML reports for failures, flakiness, and root causes.
Sift MCP provides 6 well-named tools with consistent schemas and clear intent. All tools follow verb_noun naming (ingest_, analyze_, get_). However, there are significant gaps in parameter descriptions, output schema documentation, and error recovery guidance. Descriptions are adequate (70-120 chars) but lack LLM-optimized guidance on WHEN to use each tool. Parameter descriptions exist but lack format constraints, ranges, and validation rules. No output schemas are documented in the tool definitions. Error handling is basic, errors return text but don't guide recovery or provide alternatives.
Re-analyze a previously stored test report by its ID. Runs the analysis pipeline again and returns updated results.
Get failure history for a specific test name. Shows how many times it failed, when it first and last failed, and related report IDs.
Identify flaky tests that intermittently pass and fail across stored reports. A test is flaky if it has both passes and failures in recent runs.
Get overall statistics: total reports stored, aggregate pass/fail rates, and top failing tests.
Analyze failure severity trends over time. Shows how critical failures are changing in the specified time period with bucketed data.
Ingest a base64-encoded JUnit XML report. Parses the XML, runs the analysis pipeline, stores in database, and returns the analysis summary. Never returns raw XML.
Output schemas completely undocumented. None of the 6 tools specify what fields are returned. LLMs cannot plan downstream tool calls or extract data from responses without visible output schemas.
Parameters use 'e.g.' examples instead of enum constraints. time_range in get_flaky_tests, get_report_stats, and get_severity_trend all say 'e.g., 7d, 30d' instead of declaring an enum ['7d', '30d', '90d'].
Missing bounds and validation rules on numeric parameters. 'limit' in get_failure_history and 'min_runs' in get_flaky_tests have no stated min/max.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 0 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | 2024-11-05+ | v1 |
Parameter descriptions lack format/validity guidance. 'test_name' in get_failure_history says 'Fully qualified test name' without explaining format (e.g., 'com.example.Test#method' vs 'example.Test::method'). 'report_id' in analyze_results says 'UUID of the stored test report' but doesn't explain where to obtain it (output of ingest_report?).
Tool selection guidance is sparse. Descriptions don't explain when to use analyze_results vs get_report_stats, or get_flaky_tests vs get_failure_history.
Error handling provides no recovery guidance. Code shows ErrorResult() returns text like 'invalid base64 encoding: ...' but doesn't tell LLMs what to try next.
No pagination documented. get_failure_history accepts 'limit' but doesn't specify how to iterate beyond the limit (does it return a cursor? total count? next_token?).
Ambiguous parameter semantics. 'bucket' in get_severity_trend has description '(hour, day, week)' but unclear if it's an enum or free-form string. Is 'hour' valid? 'hourly'? '1h'?