An MCP server for SRE ticket analysis and incident management, integrating Notion for ticket tracking, Google Gemini for AI-powered diagnosis, and automated rollback execution.
Four tools defined with basic schemas and descriptions, but significant gaps in parameter documentation, output schema clarity, and error recovery guidance. Tool names follow verb_noun convention adequately, but descriptions lack depth and context for LLM selection. Parameters are minimally documented. No output schemas are documented for any tool. Error handling is generic and provides no recovery guidance. The server attempts to implement an SRE workflow (ticket analysis → RAG search → diagnosis submission → rollback execution) but the tool interface doesn't expose enough structure for reliable agent composition.
Menjalankan webhook pemulihan server (rollback) HANYA untuk tiket yang telah disetujui (dicentang) oleh manusia di Notion.
Mengambil daftar tiket error yang baru masuk (berada di kolom 'Analyzing') di Notion.
RAG: Mencari solusi resmi di database SOP/Runbook KuroTech berdasarkan kata kunci error (misal: 'fs' atau 'timeout'). SELALU gunakan tool ini sebelum membuat diagnosis akhir!
Mengirim hasil pemikiran AI kembali ke tiket Notion, lengkap dengan skor keyakinan dan deteksi kebocoran data (DLP).
No output schema documented for any tool. LLMs cannot plan downstream tool calls or extract required fields (e.g., pageId from get_analyzing_tickets for submit_ai_diagnosis).
get_analyzing_tickets accepts no parameters and returns unstructured JSON. Should document the returned object structure (array of {id, title, errorLog}), enable pagination support, and optionally accept filters (e.g., status, limit).
search_kuro_runbook description does not explain what constitutes a match or how results are ranked. Does not document output structure (returned docs, fields). 'SELALU gunakan' (always use) is imperative but doesn't explain WHY, missing context for agent selection.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
submit_ai_diagnosis parameter 'severity' lacks an enum constraint. Description says 'Low, Medium, High, atau Critical' but this is a string hint, not a formal constraint. LLMs may hallucinate variants (lowercase, abbreviated, misspelled).
submit_ai_diagnosis parameter 'confidenceScore' has no range constraint (0 - 100 implied but not enforced). LLMs may pass invalid values like 150 or -5. Should declare min/max in schema and description.
execute_approved_rollbacks description says 'HANYA untuk tiket yang telah disetujui' but does not explain what happens if no approved tickets exist, what the success criteria are, or what fields are required in the Notion page (e.g., 'Approve Rollback' checkbox). Vague behavior invites misuse.
Error handling is generic ('Error KuroSRE: {error.message}'). Does not categorize errors as retryable, user-fixable, or fatal. Does not provide recovery guidance. Example: if a Notion query fails, the LLM has no guidance on whether to retry, call a different tool, or halt.
No parameter descriptions for get_analyzing_tickets and execute_approved_rollbacks (both accept empty objects), but this is poor API design. These tools should accept optional filters (e.g., limit, status filter, approval filter) to reduce unnecessary data transfer and improve agent control.
Tool composition breaks: get_analyzing_tickets returns {id, title, errorLog}, but submit_ai_diagnosis expects 'pageId', field name mismatch. Agent must infer that 'id' = 'pageId'. Should use consistent naming (pageId in both) or document the mapping explicitly.
No confirmation or dry-run pattern for execute_approved_rollbacks (a destructive operation that affects production). High-risk tools should support a confirmation step or at least return a simulation before committing. Current design risks accidental rollbacks.
search_kuro_runbook has no pagination support (accepts keyword, no limit/offset/page). If the runbook has hundreds of matching docs, returning all results wastes tokens and degrades agent reasoning. Should cap results (e.g., top 10) and document the limit.
Tool descriptions are in Indonesian (Notion column names, error messages) while rubric emphasizes English for global LLM compatibility. This limits usability in multi-language agent deployments and increases confusion when LLMs encounter mixed-language schemas.