A Next.js demo application showcasing integration with OpenAI's Responses API and Pearl MCP server for expert assistance
Single tool 'askExpert' has a well-structured input schema with clear parameter types and descriptions, but critical gaps prevent a higher score. The tool description is adequate (152 chars, within the 10-1024 baseline), but it does not document the output schema or error handling behavior. The schema defines chatHistory as an array of objects with role/content but lacks clarity on the expected response format, pagination, or error cases. The tool name 'askExpert' is a verb_noun construction but is vague about what 'asking an expert' entails, does it validate input, route to a specific expert, or something else? Parameters are well-documented with descriptions, but the schema lacks specification of the complete output structure that the LLM needs to understand for chaining. No confirmation step documented for a tool that initiates external communication. No mention of rate limits, timeouts, or retry behavior. This is a functional tool definition but falls short of production-grade clarity and error recovery.
Connects user with an expert by gathering clarifying information. Accepts either full conversation history on first call or session_id with follow-up question on subsequent calls. Returns expert response with session_id for session management.
Output schema not documented. The tool description and parameters are clear, but the LLM has no formal specification of what the response contains (e.g., fields in the expert response, session_id format, any metadata). This forces the LLM to infer response structure from examples or trial-and-error.
No error handling or recovery guidance. If the expert connection fails, the session_id is invalid, or rate limits are hit, there is no documented behavior or suggested remediation. Error responses should tell the LLM what to do next (retry, inform user, etc.).
Tool name 'askExpert' is vague about intent and scope. It does not clarify whether the tool validates the question, routes to a specific expert category, performs any pre-processing, or simply passes input through. A name like 'route_question_to_expert' or 'submit_expert_inquiry' would be clearer.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Mutually exclusive parameters (chatHistory vs sessionId) not explicitly constrained. The description states 'Required on first call, omitted on follow-up calls' but does not formally declare them as mutually exclusive. The LLM may pass both, causing ambiguous behavior.
No pagination or result limit documented. If the expert response contains embedded conversation history or multiple expert suggestions, the tool should document limits and pagination to prevent context bloat.