UltraRAG is a comprehensive RAG (Retrieval-Augmented Generation) framework with MCP server support, featuring multiple specialized servers for routing, benchmarking, and retrieval operations.
UltraRAG presents severe definition quality issues across multiple dimensions. Tools lack comprehensive parameter descriptions, output schemas are undocumented, and error handling is absent. Most tools show minimal naming clarity and no guidance for LLM selection. The 11 tools are heavily specialized for a RAG system but lack the documentation and structure required for reliable agent integration. STDIO transport and lack of schema visibility further constrain practical utility.
Build and generate server configuration
Check if model should continue or stop based on search token.
Load benchmark data from file with key mapping and optional shuffling.
Check if IRCoT answers are complete based on completion phrase.
Check if r1_searcher answers are complete based on EOS tokens.
Route queries to state1 or state2 based on query value.
Route all queries to state2.
All tools lack documented output schemas. LLMs cannot infer what fields are returned, preventing downstream tool composition and forcing agents to guess at response structure.
Parameter descriptions are minimal or absent. Example: 'query_list' parameters across multiple tools lack actionable guidance on format, content, or expected structure.
Tool names lack clarity and action verbs. Names like 'ircot_check_end', 'search_r1_check', and 'surveycpm_state_router' use domain-specific acronyms without context, making it difficult for LLMs to understand when to invoke them.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Check if Search O1 should stop or continue retrieving.
Check if search-r1 answers are complete based on EOS tokens.
Route SurveyCPM state variables based on current state.
Check if WebNote pages are complete or incomplete.
No error handling documentation. Tools provide no guidance on failure modes, retryability, or what actions an agent should take if a check returns false or data is malformed.
'build' tool has empty input schema ({}). Cannot validate this is intentional vs. a missing definition. If it takes no parameters, description should state this explicitly.
surveycpm_state_router accepts 6 list parameters with deeply nested meanings (state_ls, cursor_ls, survey_ls, step_ls, extend_time_ls, extend_result_ls) but provides no documentation of the relationship between these lists, their cardinality, or expected element types.
get_data 'benchmark' parameter is documented as a dict with required/optional keys, but no schema is provided for the dict structure. LLMs cannot validate input structure or understand available keys.