MCP server for indexing website documentation, planning crawls, and searching indexed content using full-text search
Definition quality is moderate with significant gaps. All 9 tools have descriptions and input schemas are present and properly typed. However, parameter descriptions are often generic or missing guidance, error handling is inconsistent, and output schemas lack formal documentation. The server follows basic tool definition patterns but falls short of production-grade standards. Naming is generally clear (verb-noun convention), but parameter descriptions frequently lack constraints, ranges, or usage context that LLMs need to invoke tools correctly.
Show the schema and sample rows for a specific database table.
List all tables in the siteindexer SQLite database. Useful for debugging and inspection.
Delete all indexed data for a source (pages, chunks, plans). Use this before re-indexing a source from scratch.
Return stored page content (not live fetch).
List all indexed sources with page counts and first/last fetch timestamps.
Plan an index run without fetching full content. Returns the exact list of URLs the server intends to fetch.
Shortcut that combines plan_index and run_index in a single call, skipping the approval gate. Useful for re-indexing a known source quickly. For new or large sources, prefer plan_index first to review URLs before committing.
Output schemas not formally documented for any tool. LLMs cannot predict the structure of returned data, forcing them to guess at field names and types. This breaks tool chaining and increases error rates.
No confirmation or dry-run gate for destructive operations. delete_source is called with a single parameter and immediately deletes all data for a source. Agents make mistakes, this pattern invites catastrophic errors.
Parameter descriptions lack constraints and guidance. Numeric parameters (max_pages, top_k, limit) have no min/max ranges. String parameters (scope_mode, table) lack enum constraints or lists of valid values. source_name validation rules are stated in prose but not enforced in schema.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Execute an approved indexing plan: fetches all planned URLs concurrently, extracts text content, splits into chunks, and stores everything in the index. Call plan_index first to get a plan_id.
Search indexed documentation using full-text search (BM25).
Error responses provide no recovery guidance. When a tool fails, it returns an error message but does not suggest what the LLM should do next (e.g., 'Unknown plan_id: X, try plan_index() first'). This forces the agent to reason about recovery instead of following guidance.
refresh tool combines two responsibilities (plan + run) in a single call, skipping the approval gate. This violates single-responsibility principle and introduces a security concern: agents cannot review URLs before committing. Tool description warns against use but does not fully explain the risk.
Discovery tools (db_list_tables, db_describe_table, list_sources) lack clear guidance on when to call them. Descriptions are minimal; dependency relationships between tools are not documented (e.g., db_describe_table requires the output of db_list_tables to know valid table names).