AI agent system for automated PowerPoint slide editing and content generation via natural language commands
This server exhibits severe definition quality issues across nearly all dimensions. Tools lack proper naming conventions (Planner, Parser, Processor are vague nouns rather than action verbs), descriptions are generic and uninformative (e.g., 'Processes parsed slide content' provides no guidance on when to use this vs similar tools), and parameter documentation is minimal. Input schemas exist but are sparse, many parameters lack type specificity or validation constraints. Output schemas are entirely undocumented. The codebase shows Flask HTTP endpoints mixed with legacy pyautogui GUI automation code, suggesting architectural instability. Tools 7-11 are in a 'legacy' folder and appear abandoned. Tool definitions are partially inferred from Flask route signatures rather than explicit MCP schema registration visible in the source. No error handling guidance, no recovery paths, and several tools (e.g., Applier, SharedLogMemory, PowerPointAction) perform destructive operations (writes to PowerPoint, logging) without confirmation patterns or audit trails. Security baseline violated: no mention of credential injection, and the code references raw API keys in environment variables without isolation.
Applies edits to PowerPoint by generating and executing win32com code to modify presentations
Generates action plans for PowerPoint control using natural language understanding
Searches PowerPoint manual JSON file for information relevant to user commands
Parses PowerPoint slides to extract content and structure based on planning tasks
Creates a plan for PowerPoint editing tasks based on user input by requesting plan from LLM
Executes individual GUI actions on PowerPoint including click, type, hotkey, scroll, and drag operations
Processes parsed slide content and generates editing instructions using LLM
Naming does not follow verb_noun pattern. Tools are named with generic nouns (Planner, Parser, Processor, Applier, Reporter) rather than action verbs (create_plan, parse_slides, process_content, apply_edits, generate_report). This violates the foundational pattern:tool convention and forces LLMs to infer intent from descriptions rather than names.
Descriptions are generic, under-specified, and lack context on WHEN to use each tool. Example: 'Processes parsed slide content and generates editing instructions using LLM' does not explain what distinguishes Processor from Planner or LLMModule. Descriptions must answer: What? When vs similar? What does it return? Current descriptions answer only 'what' partially.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 27 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Generates a summary report of the processing results and applied changes
Captures PowerPoint window or full screen and converts images to base64 encoding
Logs and stores execution history including user input, plan, processed data, and results
Analyzes PowerPoint screen images to extract UI elements, text content, and application state
Generates and executes Python code using win32com to edit PowerPoint based on natural language instruction
No output schemas documented. Tools return JSON objects (plan_json, processed_json, result) but the structure of these objects is never defined. LLMs cannot plan downstream tool calls or extract required fields without knowing the response schema.
Destructive operations (Applier, baseline1 modify PowerPoint; SharedLogMemory persists data) lack confirmation, dry-run, or audit trail patterns. No error recovery guidance. Agents could irreversibly modify presentations without safeguards.
Parameters lack validation constraints and error guidance. 'api_key' is exposed as a parameter (security violation: credentials should never appear in tool params). 'model_name' accepts free-form strings with no enum. 'json_data' parameters lack schema documentation, what fields must be present?
Tool definitions appear inferred from Flask route code rather than explicitly registered with MCP schema. No visible MCP SDK initialization, tool registration, or schema declarations in provided source. Tools 7-11 are in a 'legacy' folder and appear abandoned/uintegrated.
Parameter descriptions are minimal or absent. Examples: 'json_data' (what fields?), 'params' in PowerPointAction (what keys are valid?), 'action_type' (enum not documented in description). LLMs cannot validate inputs without explicit parameter docs.
No error handling or recovery guidance. Tools return generic JSON or errors without actionable next steps. Example: LLMModule.get_action_plan returns {"error": "LLM 응답을 JSON으로 파싱할 수 없습니다."}, LLM cannot recover or retry intelligently.
baseline1 tool name violates all naming conventions. Single word, no verb, no semantic clarity. Appears to be a legacy/debug artifact left in production.
No documented input validation, range limits, or format constraints. 'target_size' in ScreenCapture accepts array but no bounds (could be [0, 0] or [999999, 999999]). 'action_type' in PowerPointAction is free-form string.