BM25-based code search engine with CLI and MCP integration for coding agents. Provides full-text search, repository indexing, and file operations for polyglot codebases.
Strong tool design with consistently good naming, comprehensive descriptions, and detailed parameter schemas. Most tools follow verb_noun convention. Descriptions average 180-250 characters with clear intent statements. Input schemas are complete with proper JSON Schema types and constraints. However, output schemas are not explicitly documented in the source code provided, and tool annotations (readOnlyHint/destructiveHint) are missing from the MCP registration layer. The server implements 14 well-designed tools for code search and repository indexing with strong parameter validation (regex patterns, min/max bounds, enums for pattern_type). Error handling guidance is embedded in descriptions but not formalized in a recovery pattern. Session management is clear but lacks idempotency guarantees in documentation.
Delete a session and all associated indexed data permanently. Removes the entire session directory including index, metadata, and file cache. This operation is irreversible. Use with caution.
Find files by glob or regex pattern within a session. Searches file paths against the pattern and returns matching files with metadata. Useful for locating files by name pattern without full-text search.
Find all references to a symbol (function, variable, class name) across the entire indexed codebase. Uses case-sensitive substring matching for language-agnostic symbol resolution. Returns all occurrences with file paths and line numbers.
Get version and build information about the running shebe-mcp server. Returns server version, protocol version and available tools. Use this to check which version of shebe-mcp is running. Fast operation (<1ms).
Get detailed metadata for a specific session. Shows repository path, files indexed, chunks created, index size, last indexed time, and schema version. Useful for validating session state before operations.
Output schemas not documented in source code. LLMs cannot plan downstream tool calls or extract return values without knowing the response structure (e.g., what fields search_code returns, pagination structure, metadata format).
Tool annotations (readOnlyHint for read operations, destructiveHint for delete_session, idempotentHint where applicable) missing from MCP tool registration. Agents cannot determine at registration time which tools are safe to retry or have side effects.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 50 | 2024-11-05+ | v1 |
Index a code repository for BM25 full-text search (REQUIRED before search_code works). Runs SYNCHRONOUSLY (blocks until complete) and returns actual statistics. PERFORMANCE (tested on 6,364 files): Small repos (<100 files): 1-4 seconds, Medium repos (~1,000 files): 2-4 seconds, Large repos (~6,000 files): 10-15 seconds, Very large repos (~10,000 files): 20-30 seconds. Throughput: 1,500-2,000 files/sec (varies with system load). CREATES A SESSION for future search_code queries. Session persists until deleted. Supports polyglot codebases (PHP+SQL+JS+HTML+CSS+Rust+Python+etc). FILE FILTERING: Use glob patterns. Defaults exclude build artifacts (target/, node_modules/, .git/, dist/, __pycache__/). Customize with include_patterns and exclude_patterns. CHUNKING: Default 512 chars/chunk with 64 char overlap. Increase chunk_size (max 2000) for verbose languages (Java, C++), decrease (min 100) for dense code (Python, Ruby).
List files in a session with cursor-based pagination. Returns file paths, types, and sizes. Supports recursive directory traversal. Useful for exploring repository structure without full file reads.
List all available sessions with metadata. Shows session ID, files indexed, chunks created, index size, and last indexed timestamp. Fast operation (<1ms). Useful for discovering which sessions are available before searching.
Show N lines before/after a search result chunk. Takes a chunk reference from search_code results and returns the surrounding context. Useful for understanding code context around search matches.
Read file contents from a session with optional offset-based pagination. Returns file content with optional line range selection. Supports large files by reading only requested portions. Useful for examining source code referenced in search results.
Re-index a session using stored repository path. Convenient for schema migrations or config changes. Automatically retrieves original path and config from metadata. Supports config overrides (chunk_size, overlap). Use force=true to re-index even if config unchanged.
Search indexed code using BM25 full-text search. Returns matching chunks with line numbers and file paths. Supports pagination (cursor-based) and result limiting. Use list_sessions to find available sessions. Fast operation (<100ms for most queries). Returns chunks in context order.
Show the current shebe configuration (from config file or defaults). Displays all indexing options, include/exclude patterns, chunk settings, and other configuration values. Useful for understanding current settings before indexing.
Upgrade session metadata to latest format. Useful for schema migrations when shebe-mcp is updated with breaking changes. Automatically detects current schema version and applies necessary migrations.
Destructive operations (delete_session, index_repository with force=true, reindex_session with force=true) lack confirmation/dry-run mechanics. Agents cannot preview changes before irreversible operations.
Error handling guidance not formalized. Descriptions mention performance ('Fast operation <100ms') but do not specify error scenarios, retryability, or recovery steps (e.g., 'Session not found, call list_sessions() to discover valid sessions').
Session parameter pattern '^[a-zA-Z0-9_-]{1,64}$' inconsistently applied. get_session_info uses this pattern, but search_code does not validate it, LLMs could pass invalid session IDs without schema-level feedback.
Idempotency not documented. index_repository and reindex_session both re-index synchronously, agents need assurance that calling twice with identical parameters produces identical results (essential for retry safety).