Bagel answers questions about robotics, drone, and IoT data (ROS 1/2 bags, MCAP, PX4/ArduPilot/Betaflight logs, CAN/MF4, live MQTT) by generating DuckDB SQL over the actual messages: never estimate a numeric answer yourself, and show the user the query you ran.
Bagel is a domain-specialized MCP server for robotics/IoT data analysis with 15 well-scoped tools. Strengths: clear naming conventions (verb-noun patterns like describe_data_source, query_messages, run_pipeline), comprehensive parameter descriptions with context about DuckDB schema and time bounds, and proper risk classifications (READ_ONLY, WRITE, DESTRUCTIVE). Schema definitions are visible in tool registration via mcp_compat.tool_annotations(). Weaknesses: (1) output schemas are documented in docstrings but not formally returned in the MCP tool definition; (2) error handling descriptions are minimal, tools lack 'what to do if X fails' guidance; (3) some parameter descriptions could be more prescriptive about constraints (e.g., no explicit mention of SQL injection avoidance, no format examples for Jinja templates); (4) a few tools like subscribe/unsubscribe and upload_artifacts have vague 'args' parameters that accept open-ended dictionaries, reducing schema clarity.
Delete a saved pipeline from PIPELINES_DIRECTORY.
Inspect a robotics, drone, or IoT log first: returns source metadata, available topics, and instructions for summarizing them. Use describe_topic next for field schemas before SQL. Does not return message rows or detect anomalies. Paths are resolved on the Bagel server; runtime and schema support depend on the selected service.
Inspect one known topic before writing SQL or event predicates. Returns its DuckDB schema, original message definition, and query instructions, without message rows. Discover topic names with describe_data_source. Confirm units from the definition or user; do not infer units from field names.
List all available agent capabilities (builtin and user-authored). User capabilities are discovered from USER_CAPABILITIES_DIRECTORY.
List all saved pipelines from PIPELINES_DIRECTORY.
Preview a pipeline's event detection and data reduction before writing to disk. Returns event count, time span covered, and duration of data retained. Use run_pipeline afterward if the preview looks correct.
Output schemas underdocumented in MCP tool definitions. While docstrings describe return structure (e.g., 'Returns rows as dictionaries' for query_messages, 'Returns artifact paths' for run_pipeline), formal JSON Schema definitions for response types are not visible in the tool registration. LLMs need machine-readable output schemas to plan downstream steps and extract fields reliably.
Vague 'args' parameters accept open-ended dictionaries (e.g., describe_data_source, describe_topic, query_messages, read_loggings, preview_pipeline, subscribe). Documented as 'Additional constructor arguments' but no enum or structure definition, forcing LLMs to guess valid keys. This reduces discoverability and invites invalid inputs.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 73 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Answer quantitative questions with read-only DuckDB SQL over one topic: filtering, aggregates, downsampling, and event evidence. Call describe_data_source and describe_topic first; use the returned schema and show the SQL to the user. Returns rows as dictionaries. Use read_loggings for textual diagnostics. Time bounds are inclusive source timestamps in seconds, not offsets from the first message.
Read textual INFO/WARN/ERROR diagnostics from a recorded source, optionally within inclusive source-time bounds in seconds. Returns logging records, or an empty list.
Execute a pipeline to detect events and reduce data to disk. Returns artifact paths. Call preview_pipeline first to review the impact.
Execute a saved pipeline template with variables. Variables are substituted into the Jinja template before execution.
Save a user-authored agent capability (as .poml or .md) to USER_CAPABILITIES_DIRECTORY for discovery by list_agent_capabilities.
Save a Jinja pipeline template to PIPELINES_DIRECTORY for later execution via run_saved_pipeline.
Subscribe to a live MQTT topic or data stream and write messages to a topic sink.
Unsubscribe from a live topic sink.
Upload artifacts to cloud storage (AWS S3 or Google Cloud Storage). Requires EXTELLIGENCE_S3_BUCKET_NAME or equivalent cloud configuration.
Error handling descriptions are minimal. Tool descriptions lack 'what to do if X fails' guidance. For example, query_messages does not say how it fails on invalid SQL or missing topics, or what an LLM should do next (e.g., 'If schema mismatch, call describe_topic first'). This violates the recovery-guide pattern.
Destructive tool delete_pipeline lacks confirmation or dry-run pattern. No warning in description that deletion is irreversible, and no preview_before_delete tool to let agents verify before wiping data. This risks accidental data loss.
Parameter constraints under-documented. save_pipeline/save_agent_capability accept 'filename' and 'content' but do not specify allowed formats (e.g., '.yaml' for pipelines, '.poml' or '.md' for capabilities). LLMs may pass invalid filenames. Descriptions should include format constraints and examples.
Pagination/result limits not documented. query_messages may return large result sets but no indication of row limit, pagination cursor, or how LLMs should handle thousands of matches. This risks context window exhaustion.
subscribe and upload_artifacts have minimal descriptions of required configuration. 'EXTELLIGENCE_S3_BUCKET_NAME or equivalent cloud configuration' is vague, no enum of supported cloud backends (AWS, GCS, Azure), no indication of how to validate configuration before execution.
Tool annotations present (mcp_compat.tool_annotations(read_only=True, idempotent=True) seen in describe_data_source) but inconsistently applied. Not all READ_ONLY tools show annotations in visible code; WRITE and DESTRUCTIVE tools lack explicit destructiveHint annotation in visible registrations.