A Discord archive tool with optional RAG support providing semantic search and SQL query tools via MCP.
discord-archive provides 7 tools with generally good structure, but exhibits inconsistent parameter documentation and lacks explicit output schema definitions. Tool names follow the verb_noun pattern (semantic_search, graph_search, get_context_window, sql_query, get_message_attachments, get_message_details, get_user_profile) which is strong. Descriptions are present and reasonably detailed (150-400 chars each), explaining the semantic or functional purpose. However, parameters lack consistent type-checked constraints, notably, 'sql_query' accepts free-form strings with only a description guardrail, and several tools use default=null inconsistently. Output schemas are inferred from implementation code rather than explicitly documented in tool definitions. Error handling and recovery guidance are minimal, no actionable error messages or retry strategies visible in the tool contracts.
Retrieve context around a chunk: preceding and following messages from the same thread/channel.
Retrieve attachment metadata (URLs, file names, MIME types) for a list of message IDs.
Retrieve full message details (author, timestamp, content, embeds, reactions) for a message ID.
Retrieve a user's profile (username, avatar, status) via Discord API.
Search by social graph proximity (interaction neighbors in guild).
Search Discord archive chunks by semantic similarity. Encodes the query with NV-Embed-v2 and performs ANN search in LanceDB. Returns JSON array of {chunk_id, distance, guild_id, channel_id, author_ids, has_attachments, timestamps}. Lower distance = more similar.
sql_query tool accepts free-form SQL strings with minimal validation. No enum constraints, no input sanitization hints, and no documented protection against SQL injection. Description says 'read-only' but relies on backend enforcement only.
Output schemas are not explicitly documented in tool definitions. Developers must infer structure from implementation code (_serialize_results() function). LLMs cannot plan downstream operations without seeing schema. This violates the documented return type requirement.
Multiple parameters use default=null semantrically (guild_id, channel_id, author_id, before, after, instruction in semantic_search; guild_id, channel_id, instruction in graph_search). Inconsistent handling of optional filtering, some tools omit null filtering, others accept it. No documentation of how null is handled downstream.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 57 | 2026-07-28+ | v2 |
Execute a read-only SQL query against the archive database. The connection is read-only (no INSERT/UPDATE/DELETE allowed). Useful for aggregate queries, joins, or complex filters not covered by semantic_search.
No error handling or recovery guidance in tool descriptions. Tools do not document failure modes (e.g., embedding model load failures, LanceDB search failures, Discord API timeouts). LLM has no guidance on retry strategy or fallback behavior.
Tool composition does not enforce idempotency or transaction safety. semantic_search and graph_search both call embedding_model.load() lazily, concurrent calls could cause race conditions. No confirmation step for destructive operations (though all tools are read-only, confirming this would be clearer).
Parameter descriptions omit format constraints for ID fields. Tools accept author_id, guild_id, channel_id, user_id, message_id as strings or integers with notes about 'precision loss' but no explicit validation rules, regex patterns, or length constraints. LLMs cannot validate before calling.
No pagination support in tools returning lists (semantic_search, graph_search results). Default limit=20 is enforced, but no next_cursor, offset, or total_count returned. If results exceed 20, LLM cannot iterate, violates paginated-result pattern.