An AI-powered developer productivity agent that integrates GitHub, Google Calendar, and Google Tasks to manage development workflows, schedule coding sessions, and track productivity.
DevFlow AI is a STDIO-only MCP server with 22 tools spanning GitHub, Google Calendar, Google Tasks, and development task management. Critical gaps: Tool definitions lack input schemas for most tools; many descriptions are present but generic or under-specified; no output schemas documented; error handling provides minimal recovery guidance. The codebase shows the tools are implemented via LangChain decorators and custom Python, but the MCP registration layer (src/mcp/server.py) is not fully visible, forcing inference of tool definitions in several cases. Tool naming is generally good (verb_noun pattern), but parameter descriptions lack detail on formats, ranges, and validation rules. No security annotations (readOnlyHint, destructiveHint) despite several destructive operations (delete_calendar_event, delete_task). Baseline expectation for this rubric is 45-55; this server falls into the lower tier of that range due to schema gaps and missing error guidance.
Natural language interaction with the DevFlow AI agent
Create a new event in Google Calendar.
Create a new development task with priority and time estimate
Create a new development task in Google Tasks.
Delete a calendar event by its summary/title.
Delete a task from Google Tasks.
Find available free time slots in the calendar.
Duplicate tool names: update_task_status appears twice (tools #11 and #17) with different signatures and behaviors. The first accepts task_title (string-based lookup), the second accepts task_id (numeric ID). This violates the single-responsibility principle and creates ambiguity for LLMs on which to invoke.
No output schemas documented for any tool. The rubric requires documented return types for A+ quality. Example: get_my_assigned_issues returns a markdown string with issue metadata, but LLMs cannot predict the structure without an explicit schema. This prevents reliable downstream tool chaining and forces LLMs to parse unstructured text.
No input schema visible for chat_with_agent in the provided code. Only a description is shown. Per hard scoring rule: if a tool has NO input schema at all, its schema score MUST be 0. This tool appears to be inferred from the mcp/server.py file reference, further reducing confidence.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 70 | 2026-07-28+ | v2 |
Get calendar events for a specific date.
View scheduled coding sessions for a specific day
Get all open issues assigned to the authenticated user across all repositories. Use this for questions like 'What are my tasks?' or 'What is on my plate?'.
Get all open pull requests created by the authenticated user. Use this to check the status of your code reviews.
Get productivity statistics and completion metrics
Get statistics about tasks (completion rate, pending count, etc.).
List development tasks, optionally filtered by status
List pull requests for a specific repository.
List issues for a specific repository.
List tasks from Google Tasks.
Analyze and suggest task prioritization.
Perform self-reflection on productivity and task completion
Schedule a focused coding session for a task
Update the status of a task
Mark a task as completed or pending.
Destructive tools lack security annotations (destructiveHint). Tools delete_calendar_event, delete_task, and operations that modify state (create_task, update_task_status) should declare their destructive/write intent. This prevents agents from invoking destructive tools without explicit user confirmation.
Parameter descriptions lack validation details. Example: 'start_time' in create_calendar_event accepts 'ISO format or natural language like 2pm, tomorrow 10am' but does not specify: (a) What timezone is assumed? (b) What is the exact ISO format? (c) Are relative offsets (e.g. '+2 hours') supported? (d) What happens if natural language is ambiguous? This forces LLMs to guess valid inputs.
No error handling guidance visible in tool definitions. Error responses in the code (e.g., 'GitHub Error: ...' or 'Error: ...') lack actionable recovery instructions. Current errors provide only raw exception messages.
Natural identifiers (task_title, event_summary) are used for lookup instead of system IDs in some tools (update_task_status, delete_calendar_event). This forces partial-match lookups and risks collisions if two tasks have identical titles. Best practice: accept both human-friendly names AND system IDs, or use unique system IDs exclusively.
No pagination or result limits documented for list tools. list_repo_issues, list_pull_requests, and list_tasks may return hundreds of items. Without pagination (offset/limit) or a cap on results, the response can blow the context window and degrade LLM reasoning. Code shows a hardcoded limit of 15 for GitHub, but this is not exposed in the tool schema.
chat_with_agent tool purpose is unclear. A tool that forwards user messages to the agent itself creates a recursive loop and is outside standard MCP tool design. This tool should either be removed or its role redefined (e.g., as a debugging/diagnostics tool with explicit caveats).
Missing descriptions for priority levels in several tools. create_task and create_dev_task accept priority enum [low, medium, high, critical] but do not explain what each level means, when to use each, or if critical has special handling (e.g., escalation, notification).