WebNexus provides three well-named tools with complete input schemas and reasonable descriptions. All tools follow verb_noun naming convention (crawl_*, search_*, get_*). Schemas are properly typed with JSON Schema format. However, descriptions are generic and lack LLM-optimization guidance; parameters lack detailed format/constraint documentation; output schemas are not documented; error handling is not visible in the code provided; and no security patterns (secret injection, permission gates, audit trails) are evident. The server would benefit from richer parameter descriptions, output schema documentation, and explicit error recovery guidance.
Tools (3)
crawl_websitewritesource verified73/100
Crawl a website and optionally store documents for later search.
get_sourcesread onlysource verified70/100
Get detailed information about documents and their sources used in search results.
search_documentsread onlysource verified77/100
Search crawled documents using vector similarity, keyword search, or hybrid RAG.
Tool descriptions lack actionable guidance for LLM selection. Descriptions are functional but generic (60-65 chars) and do not address WHEN to call each tool vs alternatives, dependencies, or prerequisites.
Parameter descriptions lack format constraints and LLM guidance. E.g., 'strategy' accepts 'single_page', 'batch', 'recursive', 'sitemap' but description does not explain when to use each; 'max_depth' lacks guidance on typical values; 'similarity_threshold' (0.1 default) lacks context for why 0.1 was chosen.
Output schemas are not documented in the code. LLMs cannot plan downstream tool calls or extract needed data without knowing what fields are returned (e.g., does search_documents return 'document_id', 'source_url', 'relevance_score'? Can those IDs be passed to get_sources?).
crawl_websitesearch_documentsget_sources
Recommendations
Enrich tool descriptions to 150-250 chars, explicitly addressing: (1) what the tool does, (2) when to call it instead of alternatives, (3) what it returns. Example for search_documents: 'Search stored crawled documents using vector embeddings (semantic), keyword matching, or hybrid ranking. Use vector search for conceptual queries ("how to reset password"), keyword for exact phrase matching ("error code 403"), hybrid for balanced relevance. Returns document summaries with relevance scores; use get_sources() to retrieve full content.'
Document enum constraints formally in the schema and reference them in descriptions. For 'strategy' in crawl_website, add schema 'enum': ['single_page', 'batch', 'recursive', 'sitemap'] and describe: 'single_page: fetch URL only; batch: all URLs up to max_pages; recursive: follow links to max_depth; sitemap: use robots.txt/sitemap.xml if available.'
Add output schema documentation in tool descriptions or docstrings. Example: 'Returns array of {document_id, snippet (first 200 chars), source_url, relevance_score (0-1), metadata}. Pass document_id to get_sources() for full content.'
Implement explicit error handling with recovery guidance. Examples: 'If crawl_website times out after 30s, retry with max_pages=10; if URL returns 403, check if domain blocks crawlers; if storage fails, ensure database is writable.' Return error codes (e.g., TIMEOUT, PERMISSION_DENIED, STORAGE_ERROR) alongside messages.
Add permission/scope checks before operations. crawl_website modifies state (stores docs), verify calling agent has write:documents permission. Add an audit log entry: {timestamp, agent_id, tool, params, result, status}.
No visible error handling patterns. Code does not show recovery guidance, error categorization (retryable vs user-fixable), or actionable error messages. LLMs will not know what to do on failure.
No security patterns evident. No mention of credential/secret injection, permission gates, scope declarations, or audit trails. The 'store_documents' parameter in crawl_website could unintentionally persist sensitive data without explicit safeguards.
Parameter naming ambiguities. 'index_name' appears in both crawl_website and search_documents but no description clarifies whether indices are created implicitly or must exist first, or how naming conflicts are handled.
'search_type' enum in search_documents is not shown; description mentions 'vector', 'keyword', 'hybrid' but these should be formally declared as enum values in the schema for LLM clarity.
crawl_website description does not clarify idempotency. If called twice with the same URL and index_name, does it re-crawl and overwrite, or skip? Agents need to know for safe retry logic.
search_documents accepts 'max_results' with default 10 but lacks bounds documentation. LLMs may request thousands of results, risking context explosion and token waste. Should state hard limit (e.g., max 100).
search_documents
Clarify index_name lifecycle. Add description: 'Vector index to store/search documents. Auto-created if not exists. Pass same index_name to search_documents() to query those docs. Omit to use default shared index (not recommended for multi-tenant setups).'
Document idempotency explicitly. For crawl_website: 'Idempotent within same index_name: re-crawling same URL overwrites prior vectors. Safe to retry on network failure. store_documents=false does not persist, useful for one-off retrieval without pollution.'
Enforce and document result limits. search_documents should cap max_results at 100 with description: 'Maximum results to return (1-100, default 10). Larger limits degrade LLM reasoning; use pagination for exhaustive retrieval.' Return next_cursor for pagination support.
Add parameter relationship documentation. For search_documents, note: 'use_reranking only applies to hybrid and vector search (ignored for keyword-only). similarity_threshold only enforced for vector/hybrid; keyword search uses BM25 scoring.'
Consider adding a list_indexes() tool to help agents discover available indices and their document counts, enabling better planning and reducing failed searches on non-existent indices.