Compares log folders (test runs, deployments, nodes -- any set of logs that belong together) to find which one looks wrong, with no labels required. Wraps loglead.delta in a session model so a log root is loaded, masked, and parsed once and then interrogated repeatedly.
LogLead MCP server exhibits significant gaps in definition quality. While tool names generally follow verb_noun patterns and descriptions are present for most tools, schema documentation is severely incomplete. Parameter descriptions in schemas are minimal (typically 1-3 words), and output schemas are not documented anywhere in the provided code. The server has 28 tools, but examination of loglead/mcp/server.py reveals that most tool definitions are inferred from test files rather than explicitly visible in the main server code. Tools like `query_result`, `distance_folder_filename`, and `anomaly_folder_content` have sparse input schemas with inadequate parameter descriptions. LLMs need to know what fields to expect.' This forces LLMs to guess at response structure and plan blind. The instructions in the server docstring are helpful for domain understanding but do not substitute for per-tool documentation. Schema completeness is approximately 30% (names and basic types present; descriptions missing; output schemas absent), placing this server in the 'poor' tier.
Detect anomalies at the file level based on content
Detect anomalies at the folder level based on content
Detect anomalies at the folder level based on file names
Detect anomalies at the line level based on content, identifying point anomalies and patterns
Close and release resources for a log root session
Get detailed information about an open log root session
Calculate distance metrics at the file level based on content
Output schemas completely undocumented. No tool in the server specifies what fields are returned, data types, or structure. LLMs need to know what fields to expect so they can plan downstream tool calls and extract the right data.' An LLM calling `distance_line_content` has no way to know whether it returns buckets, scores, or a formatted report.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
Calculate distance metrics at the folder level based on content
Calculate distance metrics at the folder level based on file names
Calculate distance metrics at the line level based on content, using fuzzy diff
Filter log lines based on specified criteria
List all currently open log root sessions
List all available mask patterns
Find new tokens (words) in a log folder that don't appear in comparison folders
Open a log root for analysis with specified parameters for masking, parsing, and other processing options
Preview a log root without fully loading it to see structure and determine if it's a single file or folder set
Generate statistical plots at the file level based on content
Generate statistical plots at the folder level based on content
Generate statistical plots at the folder level based on file names
Query the results of previous analysis operations
Read all lines from a distance_line_content bucket for detailed inspection
Read log lines from specified log files within a session
Register a custom mask pattern for sensitive information masking
Re-apply masking to a log root with updated mask patterns
Run a complete analysis workflow from a configuration file
Search for log lines matching specified patterns or text
Set or update folder names for an open log root session
Split a single large log file into slices that can be compared
Parameter descriptions are minimal or missing. Tools like `distance_folder_filename`, `anomaly_folder_filename`, and similar analysis tools accept `session_id` parameter with only a 1 - 3 word description. The rubric states: 'Add a non-empty description to every input parameter. LLMs cannot infer parameter meaning from names alone.' No description explains what a session_id references, how to obtain it, or what happens if it's invalid.
No error handling guidance or recovery instructions. Tools provide no indication of what errors can occur, when to retry, or what the LLM should do if a session_id is invalid, a path does not exist, or analysis fails. The rubric requires: 'Error responses must tell the LLM what to do next.' Absence of error documentation forces LLMs to guess.
Tool definitions inferred from test files, not explicit in main server implementation. Most tools in the inventory are documented as defined in `tests/mcp/server.py` rather than `loglead/mcp/server.py`. Per the rubric: 'If you cannot see the actual tool definition in the source (only inferred): cap that tool's overall at 50.' This applies to 15+ tools whose implementations are not visible in the provided code snippet.
Unclear parameter semantics for generic names. Many tools accept parameters named `target_folder` and `target_file` with minimal description (e.g., 'Target folder selector', 'Target file selector'). Per the rubric: 'If parameters are mutually exclusive (e.g. 'user_id' vs 'email'), state this in descriptions.' No guidance on format (exact name vs glob vs list vs int index), the FolderSelector and FileSelector type aliases are defined in code comments but not reflected in parameter descriptions, forcing LLMs to infer from usage context.
No pagination or result-limiting guidance. Tools like `list_log_roots` and `list_mask_patterns` provide no indication of result size, pagination support, or max result limits. Per the rubric: 'Tools returning lists should accept page/offset and limit parameters and return a total count or next_cursor. Without pagination, large results blow the context window.' It is unclear whether these lists are bounded, and if an agent calls them repeatedly, whether results are deduplicated.
Generic, vague tool names for analysis operations. Tools like `query_result`, `run_config` are ambiguous. 'query_result' does not specify what results are queried (anomalies? distances? tokens?), and 'run_config' does not clarify whether it executes analysis, loads configuration, or both. Per the rubric: 'The tool name alone should convey what happens when called. 'update_ticket_status' is clear; 'modify_ticket' is vague.' Renaming to `query_anomaly_results` or `execute_analysis_config` would clarify intent.
No idempotency guarantees or side-effect declarations. Write tools like `set_folder_names`, `split_log_file`, and `register_mask_pattern` do not state whether they are idempotent. Per the rubric: 'If the tool modifies state (creates, updates, deletes, sends), the description must say so. Agents need to know which calls are safe to retry and which have irreversible consequences.' An LLM retrying `split_log_file` does not know whether this creates duplicate slices or overwrites existing ones.