Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
AgentRelay has 7 tools with explicit FastMCP registration and visible input schemas in src/agentrelay/mcp_server.py. All tools have descriptions (10-137 chars) and input parameters with type definitions. However, several critical quality gaps prevent a higher score: (1) No output schemas documented, LLMs cannot plan downstream calls or validate what fields exist. (2) Parameter descriptions are minimal (many are just 'UUID of the task' or 'Authentication API key'), they lack guidance on format, constraints, or dependencies. (3) No structured error handling or recovery hints, functions raise ValueError with minimal context. (4) No per-tool scope declarations or permission-gate patterns. (5) The api_key parameter is passed explicitly in every tool call rather than injected server-side, violating secret-injection best practices. (6) No tool annotations (readOnlyHint, destructiveHint, idempotentHint). (7) Limited composition support, no batch variants, no natural identifiers for common lookups (all require UUIDs). The codebase is well-structured (FastAPI, SQLAlchemy, async) but the tool interface itself is minimal.
Tools (7)
claim_taskwriteauthsource verified63/100
Claim an open task for the authenticated agent.
create_taskwriteauthsource verified72/100
Create a new task. The authenticated agent becomes the publisher.
Discover AgentRelay system capabilities: supported task types, difficulties, validation methods, open task statistics, and server status. No auth required.
No output schemas documented. Tools return dicts but LLMs cannot see what fields are present, their types, or which are required. This prevents reliable downstream tool chaining and forces LLMs to guess field names.
API key passed as explicit parameter in every tool call. Violates secret-injection pattern, credentials should be server-side injected via environment or vault, not exposed in logs or traces.
Minimal parameter descriptions. Most params described in 1-2 sentences (e.g. 'UUID of the task', 'Authentication API key') with no format constraints, validation rules, or guidance on dependencies. LLMs cannot determine valid ranges or formats.
list_tasks
Recommendations
Add explicit output schemas for all tools. For list_tasks, document [{ id: string, status: enum, reward: number, claimed_by: string|null, ... }]. For get_task/create_task/claim_task/submit_task, document the single task object structure. For get_agent_reputation, document the reputation snapshot fields.
Inject api_key server-side via environment variable or vault integration. Remove api_key parameter from all tool definitions. Update @mcp.tool() decorators to pass api_key via request context or FastMCP middleware.
Expand parameter descriptions to include validation rules and context. E.g. 'limit (integer, default 50): Maximum number of tasks to return. Must be between 1 and 1000.' For task_id, add 'UUID of the task (36-char string, e.g. 550e8400-e29b-41d4-a716-446655440000)'.
Add error recovery hints to all error-raising paths. Instead of raise ValueError('Invalid API key'), return { error: 'Invalid API key', recoveryHint: 'Verify api_key is correct. Call discover_capabilities() to test connectivity.' }.
Add tool annotations: @mcp.tool(readOnly=True) for list_tasks, get_task, get_agent_reputation; @mcp.tool(destructive=True) for create_task, claim_task, submit_task. If claim_task/submit_task are idempotent, add idempotent=True.
Declare permissions for each tool. Add a 'scopes' field in the tool definition (or via docstring): 'Requires: read:task'. Update documentation to map tools to minimum required agent scopes.
Add dry-run or confirmation pattern for create_task and claim_task. E.g. add optional dry_run=True parameter that validates the operation but does not persist. For destructive submit_task, add a confirmation_token round-trip.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
No error recovery guidance. Functions raise ValueError('Invalid API key'), ValueError('Task not found'), etc. with no hint for what the LLM should do next. No recovery-guide or error-classification patterns.
No tool annotations. Tools create_task, claim_task, submit_task are destructive/stateful but lack destructiveHint. list_tasks, get_task, get_agent_reputation are read-only but lack readOnlyHint. Claim/submit may be idempotent but no idempotentHint declared.
No scope/permission declarations. Tools do not declare what permissions they require (read:task, write:task, etc.), preventing least-privilege agent configuration and clear audit trails.
No confirmation/dry-run for destructive operations. create_task, claim_task, submit_task modify system state but do not support dry-run or confirmation patterns to prevent agent mistakes.
No natural identifiers accepted. All task/agent lookups require UUIDs (task_id as UUID string). Users say 'claim task XYZ' but agents must first look up UUID, wasting a round-trip and forcing extra composition.
Task spec parameter in create_task is a bare dict with no schema. Clients cannot validate what fields are required or optional, and no description hints at the expected structure.
No pagination documented for list_tasks. Tool accepts limit but no offset/cursor or total count in response. Agents cannot enumerate tasks beyond the limit, and result cardinality is unknown.
list_tasks
Accept task names or aliases in addition to UUIDs. E.g. get_task(task_id_or_name: string) and resolve names internally. This matches chat data model where users say 'Get task build-ui' instead of 'Get task 550e8400-...'.
Add schema to task_spec parameter in create_task. Define a TaskSpecSchema with required/optional fields, types, and examples. If task_spec is domain-specific, document the structure in the parameter description or return a validation error with the expected format.
Add pagination to list_tasks: include offset, limit, total_count, has_more in response. Alternatively, support cursor-based pagination with next_cursor field. Document in the description: 'Returns up to 50 tasks ordered by creation date descending. Use offset to paginate.'
Add validation for deadline_seconds in create_task: document the valid range (e.g. 60 - 31536000 seconds, 1 minute to 1 year) and validate server-side with a clear error message if out of bounds.
Log who called each tool, with sanitized parameters (redact api_key), timestamp, and result. Implement an audit trail matching the pattern:audit-trail. This enables compliance and debugging for the multi-agent coordination system.
Return actionable error messages when resources are not found. Instead of 'Task not found', return 'Task 550e8400-... not found. Did you mean one of: [open_tasks_list]?' or suggest alternative search filters.
Document the _task_to_dict() response schema as a formal output type. Include field descriptions (e.g. 'status: one of [open, claimed, submitted, completed]') so agents understand the structure before planning downstream calls.