A multi-agent financial analysis system that provides fundamental analysis, news sentiment analysis, research synthesis, and investment decision-making for stocks using MCP tools for data retrieval from Finnhub and Alpha Vantage
Scoring was not performed
No documented output schemas for any tool. LLMs cannot plan downstream calls or extract returned fields without knowing what the response contains. This violates the critical requirement that tools document their return types.
Parameter descriptions lack implementation detail and context. 'workflow_id' is described only as 'Unique identifier for the workflow run', the description does not explain what a workflow is, when to use run_workflow vs stream_workflow, or how to obtain a workflow_id.
Tool descriptions are below baseline quality (average 194 chars across A+ tools; these range 50-150 chars). Description for 'get_workflow_run' is only 'Retrieve details of a specific workflow run by ID' (50 chars), lacks context about what details are returned, when to call it, or dependencies.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 16 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 51 | - | v1 |
No error handling guidance in tool definitions. If a workflow fails, if a PDF upload is rejected, or if a query times out, the definitions provide no recovery hints. LLMs cannot self-correct without actionable error messages.
Parameter 'temp_workflow' in stream_workflow lacks sufficient description, only states 'Whether this is a temporary workflow' without explaining the implications or when to set it true vs false.
No distinction between run_workflow and stream_workflow in their descriptions. Both accept the same parameters and appear to do similar things, but descriptions do not clarify: when should an LLM choose one over the other? The difference (streaming vs synchronous) is implied but not stated.