Serve multiple stdio MCP servers from one container: Claude-Code-style mcpServers config, path-based routing, hub meta-tools, and CIMD-first OAuth 2.1 + API tokens for ChatGPT, Claude and any Streamable-HTTP MCP client.
This is a test fixture server with 30 tools designed to exercise edge cases (crashes, hangs, oversized responses, deprecated patterns). Most tools have minimal or no input schema, and descriptions are uniformly terse (typically under 20 chars). The server is intentionally designed to stress-test MCP clients, not to be production-grade. No tool has a meaningful input schema beyond trivial cases. Parameter descriptions are almost entirely absent (marked 'null' in the schema). Tool descriptions are curt, single-sentence statements. Error handling, recovery guidance, and structured output are not present. This is evaluation under the rubric as-is, not as a production server, but the tooling is objectively low-quality by agentic tool standards.
Writes half a JSON-RPC message and exits.
Registers another tool and announces it.
Emits notifications/prompts/list_changed.
Emits notifications/resources/list_changed.
Emits notifications/tools/list_changed.
One legitimate question and one that must be dropped.
Smuggles a roots/list toward the client.
Parameter descriptions are null or missing for 85% of tools with parameters. Schema declares types but omits actionable usage guidance. LLMs cannot infer intent from 'null' descriptions.
Tool descriptions are uniformly terse (15 - 60 chars), below the recommended 10 - 1024 char range baseline (avg 194 chars for production tools). Descriptions like 'y (17 MiB of y)' and 'Says nothing about itself' do not guide LLM selection.
No output schemas documented. LLMs cannot plan downstream calls or extract result fields. Even test fixtures should declare the shape of returned data.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
Smuggles a sampling/createMessage toward the client.
Asks a question dressed up to mislead.
Always present.
Returns a result of the requested size.
Asks the user to confirm, then reports what they said.
Asks two questions, one after the other.
Exits without answering.
Deletes a thing, and it does not come back.
Writes a non-protocol line between frames.
y (17 MiB of 'y')
Returns after the given delay.
Never returns.
Measures a thing and says so in both channels.
A normal call. Answers normally despite the noise.
Reads a thing. A second line, so the one-line summary has something to cut.
Says nothing about itself.
Proves the server still works after a refusal.
Answers immediately, so a restart can be observed.
Answers immediately.
One of very many.
Bumps the resource and announces it changed.
Milliseconds since this process finished starting.
The pid of this process.
No error handling guidance. Tools that intentionally crash (crash_now, abort_stream, hang_forever) provide no recovery strategy or actionable error message for LLMs.
Several tools explicitly exercise deprecated MCP patterns (Sampling via ask_for_sampling, Roots via ask_for_roots). These should be noted as obsolete in a real implementation.
No input validation. Parameter names lack context (e.g. 'what', 'bytes', 'ms' are ambiguous). Tool descriptions do not specify ranges, formats, or constraints.