Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
HumanAIOS MCP server has 23 tools with inconsistent definition quality. Critical issues: (1) 8 tools (intake_phase1, intake_phase3, assess, human_score, guidance, findings, health, and others from operations/acat/api/app.py) have NO visible input schemas in the provided source, capped at schema=0. (2) Many descriptions are minimal or trivial (e.g., 'Get all agents', 'Get a specific agent by ID', 'Health check endpoint'). (3) Tools like logActivity, sendPasswordResetEmail expose implementation details rather than user intent. (4) No error handling guidance, no output schemas documented, no idempotence declarations. (5) Token service tools (generateAccessToken, generateRefreshToken, etc.) conflate API concerns with agent-facing tools, agents should not call JWT generation directly. (6) Multi-responsibility concerns: logActivity conflates agent ID, activity type, description, and metadata without guidance on which are required or how they compose. (7) ACAT intake/assessment tools lack parameter descriptions entirely in the source excerpt. Baseline: avg tool description = 194 chars; most here are 10-50 chars. A+ tools document output schemas; none visible here.
Tools (23)
assessread only25/100
ACAT assessment endpoint for evaluating session data
cleanupExpiredTokensreversible50/100
Clean up expired refresh tokens and password reset tokens from database
createAgentwriteauth50/100
Create a new AI agent
findingsread only25/100
Empirica findings endpoint for assessment results
generateAccessTokenwrite50/100
Generate a JWT access token for a user
generatePasswordResetTokenwrite50/100
Generate a password reset token for a user
generateRefreshTokenwrite50/100
Generate a JWT refresh token for a user and store it in database
Token service tools (generateAccessToken, generateRefreshToken, verifyAccessToken, verifyRefreshToken, revokeRefreshToken, revokeAllUserTokens, generatePasswordResetToken, verifyPasswordResetToken, markPasswordResetTokenAsUsed, cleanupExpiredTokens) expose authentication internals as agent-facing tools. Agents should not call JWT generation or refresh token management, these are infrastructure concerns. Move to server-side secret injection pattern or remove from agent interface.
Expand all tool descriptions to 100-250 characters. Include WHAT the tool does, WHEN to call it, and key OUTPUTS. Example: 'Get all agents. Use this to discover available agents before creating or updating. Returns an array of agent objects with id, name, type, and description fields.'
Document output schemas for every tool using JSON Schema or structured text. Include return type (object/array), field names, types, and descriptions. Example: 'Returns: { agents: [{ id: string, name: string, type: string, description?: string }], total: number }'
Remove or re-architect token service tools (generateAccessToken, generateRefreshToken, etc.). These are infrastructure concerns, not agent-facing operations. Use server-side secret injection instead: inject JWT handling into the MCP server initialization, never expose token generation to agents.
Add parameter descriptions for all inputs. Expand beyond 'Agent name' to 'Agent name (required, 1-100 characters, alphanumeric and hyphens only)'. Include enum constraints, format hints, and dependencies.
Add error handling guidance to every tool. Examples: 'If agentId not found, try getAgents() to discover available agents. If still not found, create a new agent with createAgent().' Or 'Email send failed: verify recipient address is valid and not rate-limited.'
For logActivity, define valid activity_type enum values and required metadata schema. Example: 'activity_type must be one of: created, updated, deleted, assigned, completed. metadata is optional; if provided, must be a JSON object with string keys and string/number/boolean values.'
Score history
Overall score trend
↑ 40 points across a rubric change (v1 → v2)
40/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
40
2026-07-28+
v2
2026-03-09
F
0
-
v1
auth
50/100
Get a specific agent by ID
getAgentsread onlyauth50/100
Get all agents
guidanceread only25/100
Guidance endpoint for assessment recommendations
healthread onlysource verified45/100
ACAT API health check endpoint
human_scorewrite25/100
ACAT human scoring endpoint for manual assessment review
intake_phase1write22/100
ACAT intake phase 1 endpoint for initial assessment data collection
intake_phase3write22/100
ACAT intake phase 3 endpoint for final assessment data collection
logActivitywriteauth50/100
Log an activity for an agent
markPasswordResetTokenAsUsedwrite50/100
Mark a password reset token as used
revokeAllUserTokensreversible50/100
Revoke all refresh tokens for a user
revokeRefreshTokenreversible50/100
Revoke a refresh token by removing it from database
sendPasswordResetEmailwriteauth50/100
Send a password reset email to a user
verifyAccessTokenread only50/100
Verify a JWT access token
verifyPasswordResetTokenread only50/100
Verify a password reset token
verifyRefreshTokenread only50/100
Verify a JWT refresh token against stored tokens in database
Descriptions are uniformly short (10-50 chars), below the baseline of 194 chars for production tools. Examples: 'Get all agents', 'Get a specific agent by ID', 'Log an activity for an agent' lack actionable context for LLM tool selection. Expand to explain WHAT the tool does, WHEN to use it, and what it returns.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract correct fields without knowing return structure. Add return type definitions with typed fields.
logActivity tool conflates agent ID, activity type, description, and metadata without documenting which are required, relationships, or how they compose. No guidance on valid activity_type enum or metadata schema. High risk of invalid calls.
No error handling guidance for any tool. LLMs cannot recover from failures without knowing whether errors are retryable, user-fixable, or fatal. Add recovery hints like 'User not found. Try getAgent(agentId) or search with partial name.'
getActivities returns paginated results (limit, offset parameters) but no documented return schema. Cannot verify if response includes total_count, has_more, or next_offset. LLMs cannot implement safe pagination without this.
sendPasswordResetEmail, generatePasswordResetToken, and related password tools lack security declarations. No hint that these handle sensitive credentials or require specific permissions. No audit trail guidance.
Parameter descriptions are minimal or missing. Examples: 'Agent name' (for createAgent.name), 'Agent ID' (for getAgent.agentId). Baseline production tools use 50-150 char descriptions that explain format, constraints, and dependencies.
No idempotence declarations. generateAccessToken, generateRefreshToken, logActivity, sendPasswordResetEmail are write operations. Without idempotence markers, LLMs do not know if safe to retry on failure. Could result in duplicate tokens, duplicate activity logs, or duplicate emails.
ACAT Python tools (intake_phase1, intake_phase3, assess, human_score, guidance, findings) are defined in operations/acat/api/app.py with minimal descriptions and no visible input schemas. Not clear if these are REST endpoints or Python callables or MCP tool registrations.
Add idempotence declarations to write operations. Example: 'This tool is idempotent: calling with the same parameters twice produces the same result. Safe to retry on transient failures.' Or document non-idempotent behavior if applicable: 'Repeated calls create duplicate activity log entries. Use a transaction ID or check existing logs before retrying.'
Add security/permission declarations to sensitive tools. Example: 'Requires: auth:write, audit:log. Logs all calls for compliance. Only users with admin role can call this tool.'
Investigate ACAT Python tools (intake_phase1, intake_phase3, assess, human_score, guidance, findings). Provide explicit MCP tool registration with input schemas, output schemas, and descriptions. If these are REST endpoints, wrap them with schema definitions and document the HTTP contract.
Cap result lists at 20-50 items by default. Add limit and offset parameters to discovery tools (getAgents, getActivities). Trim verbose response fields (e.g., audit metadata, internal IDs agent doesn't need) to reduce token cost.
Add confirmation/dry-run pattern for destructive operations: revokeRefreshToken, revokeAllUserTokens, cleanupExpiredTokens. Example: 'Call with dry_run=true to preview changes without committing. Default is false. Use dry_run=true first to validate before actual revocation.'
Unify naming: avoid verb_noun inconsistency (e.g., 'intake_phase1' vs 'assess'). Rename to match agent conventions: 'submitIntakePhase1', 'submitIntakePhase3', 'runAssessment', 'scoreAssessmentManually', 'getGuidance', 'getFindings', 'checkHealth'.