A collection of skills and tools for AutoGen Studio, including RAG document retrieval, Slack search, Stack Overflow Teams search, web search, web scraping, and various utility tools. Also provides a dynamic MCP client for accessing MCP server capabilities.
This server has 11 tools with significant definition gaps. Most tools have brief descriptions (4-20 words) but lack detailed documentation of expected behavior, error conditions, and output schemas. Parameter descriptions are minimal or absent, making it difficult for LLMs to reason about valid inputs. No input schemas are explicitly visible in the provided source, parameters are inferred from docstrings. Several tools (mcp, story_mode, draw_geometric_structure) have vague names that don't clearly convey action. Output schemas are completely undocumented across all tools. Error handling guidance is absent. Security considerations (especially for tools like fetch_post that POST messages and save_webpage_as_text that writes files) are not addressed.
Split documents into chunks with configurable size and step. Processes files in multiple formats (md, markdown, txt, html, json, jsonl, pdf, docx, pptx, csv) and returns chunked texts with metadata.
Draws and saves a geometric structure diagram with customizable circles, colors, and lines. Creates a PNG file in the diagrams directory.
Processes the given action, either fetching or posting a message. Fetches messages from a Lambda URL endpoint or posts a message with username and content.
Dynamic MCP (Model Context Protocol) client that provides access to various server capabilities. Supports discovery-first approach: list available servers, discover tools for a specific server, and execute specific tools.
Retrieves documents from a FAISS vector store based on semantic similarity to a query. Expands chunks by including adjacent chunks up to a target length.
Scrapes a webpage and saves the extracted text content to a file. Parses HTML using BeautifulSoup and removes HTML tags.
No explicit input schemas visible in source code. Tool parameters are inferred from descriptions only, with no JSON Schema types, constraints, or enums. This violates the requirement that every parameter must have a type definition.
Output schemas completely undocumented. No tool describes what fields are returned, data types, or structure. LLMs cannot plan downstream tool calls or extract specific fields without guessing the response format.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 33 | - | v1 |
Searches DuckDuckGo for the given query and returns results with title, URL, and snippet. Supports regional and safe search options.
Performs web searches using either Google Custom Search API or Bing Search API based on configuration. Returns list of results with title, URL, and snippet.
Searches Slack messages in specified channels and retrieves thread context. Returns simplified JSON with user, text, permalink, and thread messages.
Searches Stack Overflow Teams for questions and retrieves accepted answers. Returns simplified JSON with question titles and answer bodies.
Displays an interactive story interface with AI-generated images and text overlay. Presents a canvas-based GUI with user input capability for narrative-driven interactions.
Tool naming lacks clarity for action verbs. 'mcp' is cryptic and doesn't convey what it does (should be 'execute_mcp_tool' or similar). 'story_mode' is vague about the actual action. Generic names force LLMs to read full descriptions.
Descriptions are too brief (4-50 words). They lack guidance on WHEN to use the tool, WHAT it returns, and WHY an LLM should select it over alternatives. Rubric baseline for A+ tools is 50-200 char descriptions with clear context.
Parameter descriptions missing or minimal. Many parameters (e.g., 'query' in search_slack) have no explanation of expected format, length limits, or special characters. Rubric requires 100% of A+ tool params to have descriptions.
No error handling guidance. Tools provide no documentation of what can fail, how to recover, or what the LLM should do next. fetch_post and save_webpage_as_text are particularly risky (they write data) but have no error classification or recovery hints.
No documentation of pagination or result limits. search_duckduckgo accepts 'max_results' but doesn't specify min/max bounds. retrieve_documents defaults to 5 results but no guidance on context window impact. Rubric requires documented limits and pagination to prevent context window exhaustion.
Credentials and secrets not addressed. search_slack, search_stackoverflow_teams, and search_query likely require API keys, but no mention of server-side injection or secret management. Rubric requires credentials never appear as tool parameters.
Destructive/write operations (fetch_post, draw_geometric_structure, save_webpage_as_text) lack any confirmation or dry-run pattern. No tool declares whether operations are idempotent or have side effects. Agents cannot know if retrying is safe.
Missing tool composition. For example, mcp tool is extremely generic and tries to wrap ALL MCP server capabilities into one tool. This tool violates single-responsibility principle.