Multi-agent system repository containing multiple specialized agents (Access Controller, Portfolio Manager, Loan Underwriter, Campaign Validator, PR Code Reviewer, Research Assistant, Content Moderator, Helpdesk Bot, Social Media Marketing, Portfolio Swarm) deployed via Docker with unified routing and nginx coordination
This multi-agent helpdesk system presents 15 tools across escalation, HR, and IT domains. While tool names follow verb_noun conventions well (normalize_timezone, submit_pto_request, create_support_ticket), the definitions suffer from significant gaps. Descriptions are present but often generic (10-150 chars, below the 194-char production baseline). Most critically, input schemas are visible but lack proper constraint documentation, parameters like 'department', 'urgency', 'leave_type', 'system_name', and 'access_level' declare string types without enum constraints or validation guidance, violating the constrained-input pattern. No output schemas are documented at all. Error handling is absent, there is no guidance for LLMs on recovery paths, retryability, or partial failures. The codebase also reveals tools that combine multiple concerns (e.g., generate_specialist_chat_session both generates AND registers, combining two responsibilities). Security considerations are unexplored, password resets and access requests lack permission gates or audit trail guidance. Parameter relationships (e.g., which fields apply conditionally) are undocumented. Response field naming is not specified, making tool chaining uncertain. Most tools are READ_ONLY (11/15), but WRITE tools lack idempotency or confirmation patterns.
Check the PTO balance for an employee
Check the status of company systems and services
Create an IT support ticket for complex issues
Run diagnostics for hardware-related issues
Generate a secure specialist chat session with unique URL. Registers the session with the FastAPI server for validation
Get information about company benefits
Look up company policies on a specific topic
Constrained input not enforced, 12 tools list enum values in descriptions instead of declaring them as schema enums or constraints. Examples: generate_specialist_chat_session lists departments in description; submit_pto_request lists leave_type; get_benefits_info lists benefit_type; diagnose_hardware_issue lists device_type. LLMs cannot reliably parse free-text descriptions for constraints and will hallucinate invalid values.
No output schemas documented for any tool (0/15). Agents cannot infer what fields to expect, forcing them to guess or waste tokens retrying with malformed parameters. Critical for WRITE tools: create_support_ticket, submit_pto_request, generate_specialist_chat_session, request_software_license, reset_password, request_access should return IDs/confirmation tokens that downstream tools or agents can reference.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 10 | - | v1 |
Get current date/time information, optionally in a specific timezone
Get payroll-related information for an employee
Normalize timezone string - convert abbreviations to IANA timezone names
Parse a date string in various formats and return YYYY-MM-DD format. Supports MM/DD/YYYY, YYYY-MM-DD, Month DD YYYY, and DD Month YYYY formats
Submit a system access request
Submit a request for a software license
Initiate a password reset for an employee
Submit a PTO or leave request for an employee
WRITE operations lack error handling and recovery guidance, 6 tools (generate_specialist_chat_session, submit_pto_request, request_software_license, create_support_ticket, reset_password, request_access) have no documented error cases, recovery paths, or retryability. Agents cannot determine whether to retry, ask the user, or escalate on failure.
Security-critical tools (reset_password, request_access) lack permission gates and audit trail documentation. No guidance on who is authorized to reset passwords or grant access. Agents may bypass approval workflows. Missing audit guidance means agent actions cannot be traced for compliance.
Tool composition issues, generate_specialist_chat_session combines two concerns: generating a chat session URL AND registering it with the server. Should be split into separate tools so agents can compose as needed. Similarly, create_support_ticket may benefit from a confirm_ticket pattern before execution.
Descriptions are generic and lack LLM-optimized guidance, 10 tools have descriptions under 100 chars with minimal context on WHEN to call them vs similar tools. Production baseline is 194 chars. Examples: 'Get payroll-related information' (43 chars) does not explain when check_payroll_info vs get_benefits_info; 'Check the status of company systems' (35 chars) does not distinguish from diagnose_hardware_issue.
Parameter relationships and dependencies undocumented, no guidance on which parameters are mutually exclusive, which are conditional, or what prerequisites exist. Example: submit_pto_request does not hint 'Call check_pto_balance first to ensure sufficient days'; reset_password does not warn that temporary passwords may be emailed separately.
No idempotency or confirmation patterns on destructive/write operations. Agents may retry failed requests, causing duplicate PTO submissions, duplicate tickets, or duplicate access grants. WRITE tools should either support idempotent tokens (idempotency_key) or require explicit confirmation.
Tool chaining broken, create_support_ticket (WRITE) does not document its return schema, so agents cannot extract ticket_id to pass to follow-up tools like add_comment_to_ticket (if it exists). This forces unnecessary lookups or fails silently.
Sensitive data handling not documented, tools like get_payroll_info and check_pto_balance expose employee financial/personal data but do not document permission requirements, data retention policies, or audit logging. Missing guidance on data classification and GDPR/CCPA compliance.