A stdio-based MCP server for searching and reducing logs across 11 cloud and infrastructure providers (GCP, AWS CloudWatch, Kubernetes, Loki, Axiom, Azure Monitor, Datadog, Elasticsearch, New Relic, Splunk, Sumo Logic)
logsift is a STDIO-only MCP server with 3 tools for log searching and reduction. Tool naming follows verb_noun convention (search, list_sources, reduce), which is strong. Descriptions are present and moderately detailed (80-180 chars), above the 10-char floor but below the 50-200 char LLM-optimal range. Input schemas are fully visible with type declarations and descriptions for all parameters. Output schemas are NOT documented in the source code, this is a critical gap. The server lacks error handling guidance, confirmation patterns for destructive operations (search can return thousands of logs), and pagination strategy is implicit. Field naming is consistent (provider, source, text_filter, field_filters, severity_min, start_time, end_time, max_raw_entries), but the design does not match common LLM patterns for log reduction (e.g., no explicit per-tool result limits stated in descriptions, no guidance on when to call reduce vs. incremental filtering).
List available log sources (projects, namespaces, datasets, log groups, etc.) for a given provider
Reduce log entries into summarized clusters using pattern extraction and token budgeting
Search logs across configured backends with text and field filters, severity levels, and time ranges
Output schemas not documented. Tool descriptions state what the tools return informally (e.g., 'search logs across configured backends'), but there is no formal schema specification for search() result structure, list_sources() result structure, or reduce() output. LLMs cannot predict the shape of responses or plan downstream operations without explicit return schemas.
No explicit pagination or result-limit guidance in tool descriptions. search() accepts max_raw_entries (optional, default varies by provider), but descriptions do not state: (1) default limit per provider, (2) recommended max for LLM context, (3) when to use pagination vs. fetch-all. reduce() similarly lacks guidance on token_budget defaults. This invites unbounded result fetches that could exhaust context or timeout.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 49 | - | v1 |
No error handling or recovery guidance. Tool descriptions lack any mention of what errors might occur (API auth failure, malformed time range, provider not configured), whether errors are retryable, or what the LLM should do next. E.g., if LOGSIFT_GCP_PROJECTS is not set and the user calls search(provider='gcp'), the error response should guide the LLM to check configuration rather than blindly retry.
Parameter dependencies and constraints are implicit. field_filters accepts a free-form object with canonical field names (service, host, namespace, pod, container, level, trace_id), but the description does not state: (1) which fields are supported by each provider, (2) what happens if an unsupported field is passed, (3) case sensitivity or value validation rules for level. This forces LLMs to guess and retry on invalid combinations.
Time range parameters lack format clarity. start_time and end_time accept 'RFC3339 format or duration string like 15m', but descriptions do not specify: (1) whether '15m' means the last 15 minutes from now (relative) or a fixed window, (2) whether both must be relative or both absolute, (3) what happens if start_time > end_time. Ambiguity invites invalid input.
No composition or chaining guidance. If search() returns a list of logs and the user wants a summary, reduce() must be called in a second step. Descriptions do not explain: (1) when reduce() is appropriate vs. using severity_min to filter, (2) what the expected input structure for reduce().entries is (does it expect raw search output, or a pre-processed format?), (3) whether reduce() can be called on arbitrary log objects or only logsift search results.