A lightweight and modular Python library for implementing Retrieval-Augmented Generation (RAG). It enhances the capabilities of Large Language Models (LLMs) by combining document retrieval with natural language inference.
RAGLight exposes 10 tools with HTTP/FastAPI transport. Tool naming follows verb_noun conventions (retriever, class_retriever, ingest, generate, etc.), which is positive. However, there are significant gaps in parameter descriptions, schema completeness, and output documentation. Most tools have basic descriptions (34-194 chars range), but parameter-level documentation is sparse. Several tools lack explicit input validation guidance, and error handling is not documented. The ingest and ingest_upload tools have WRITE semantics but no documented confirmation/dry-run patterns. Schemas are partially visible in Pydantic models (RetrieverInput, ClassRetrieverInput defined clearly), but FastAPI router tools lack visible schema definitions in the source, only tool names and brief descriptions appear in router.py. This makes it impossible to verify if all parameters are properly typed and described for the router-based tools (health, generate, generate_stream, ingest, ingest_upload, collections, get_config, update_config).
Retrieves class definitions and their locations in the codebase.
List all available collections in the vector store
Generate a response from the RAG pipeline based on a question
Generate a streaming response from the RAG pipeline with Server-Sent Events
Retrieve the current LLM configuration
Check the health status of the RAGLight API server
Ingest documents from a local directory, file paths, or GitHub repositories into the vector store
Missing schema definitions for FastAPI router tools (generate, generate_stream, ingest, ingest_upload, collections, get_config, update_config). Source code shows tool names and descriptions but no visible Pydantic input models or explicit schema declarations for these 7 tools. Cannot verify parameter types, required fields, or constraints.
No documented output schemas for any tool. Tools return string responses (e.g. retriever returns 'Retrieved documents:\n...'), but LLMs have no specification of what fields or structure to expect. Output documentation is required for tools returning lists (retriever, class_retriever, collections) to clarify pagination, count fields, and result structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Upload and ingest files directly into the vector store via multipart form data
Uses semantic search to retrieve relevant parts of the code documentation or knowledge base.
Update the LLM configuration at runtime
ingest and ingest_upload perform WRITE operations (modifying vector store state) but lack confirmation/dry-run patterns and explicit error handling guidance. Agents could inadvertently ingest malformed data or wrong documents without rollback options. No recovery guides for partial failures.
update_config modifies runtime LLM configuration but lacks permission gating and audit trail documentation. No verification that the calling agent has authority to change LLM provider/model. Configuration changes should be logged for compliance and debugging.
parameter collection_name in retriever and class_retriever is marked optional but no default behavior is documented. What happens if omitted? Does it search all collections, use the server default, or fail? This ambiguity forces LLM guessing.
Parameter 'question' in generate and generate_stream lacks format/length constraints. No guidance on max question length, whether structured queries are supported, or how the RAG pipeline handles malformed input. Without constraints, agents may pass oversized or incompatible queries.
ingest_upload uses 'files' parameter with format 'binary' but no guidance on file types accepted, size limits, or encodings. Agents don't know if PDFs, markdown, code files, or images are supported. Missing: file type enums, max size constraint, supported encodings.
ingest tool parameters (data_path, file_paths, github_url) are mutually exclusive but not documented. LLMs may pass multiple simultaneously, leading to ambiguous behavior. Must explicitly state 'pass exactly one of: data_path, file_paths, or github_url'.
No error handling guidance for any tool. If generate fails due to LLM timeout, retriever returns empty results, or ingest encounters a corrupt file, what should the agent do next? No recovery patterns, retryable vs fatal error classification, or actionable error messages documented.
generate_stream returns Server-Sent Events but streaming semantics are not documented. How should the agent consume the stream? Is each event a token, a sentence, or a complete response? Does the stream return structured fields or plain text? Streaming contract is undefined.