A local AI infrastructure setup system for Ubuntu 24.04 with NVIDIA GPU support, integrating vLLM language models, OpenWebUI, Ollama, Jupyter code interpreter, and SearXNG metasearch
Kaepsele is a bash-based benchmark harness with 2 tools. Both tools have acceptable descriptions (158-193 chars, within baseline 34-392 range) and input schemas with type definitions. However, critical gaps emerge: (1) Output schemas are completely undocumented, neither tool declares what fields are returned, preventing LLMs from chaining results or extracting specific data; (2) The codebase shown is Python (vllm_benchmark.py, run_benchmarks.py), not bash, the transport/framework mismatch is unclear, suggesting this may not actually be an MCP server yet; (3) No evidence of error handling, recovery guidance, or input validation; (4) The tools both expose 'api_key' as a parameter, violating the secret-injection pattern; (5) Parameter descriptions lack actionable constraints (e.g., vllm_url format, valid ranges for num_requests/concurrency/timeout); (6) No idempotency, permission gates, or audit trail patterns present. The tools operate on a remote vLLM service (READ_ONLY risk classification is correct, but does not mitigate the other issues). Without visible MCP server registration code, tool definitions are inferred from Python function signatures, capping per-tool scores at 50. Schema documentation is the most critical missing piece.
Executes benchmark tests against vLLM server with various concurrency configurations (10 to 1000 requests, 1 to 500 concurrent connections)
Runs a single benchmark iteration against vLLM with specified number of requests, concurrency level, timeout duration, and output token count
Output schemas completely undocumented. Neither tool declares return types, fields, or structure. LLMs cannot chain results or extract specific data for downstream operations.
API credentials exposed as tool parameters. 'api_key' passed as parameter will be logged in traces, prompt history, and agent logs, violating secret-injection pattern. Must use server-side secret injection.
Parameter descriptions lack actionable constraints. 'vllm_url' has no format guidance (e.g., 'http://host:port'). Numeric parameters (num_requests, concurrency, timeout, output_tokens) lack min/max bounds. LLMs will pass invalid values without guidance.
No error handling or recovery guidance. No documentation of what happens on network failure, invalid vLLM URL, timeout, or API errors. LLMs cannot determine next steps on failure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Tool definitions inferred from Python source, not explicitly registered via MCP tool registration protocol. No visible MCP server scaffolding (transport, stdio handler, tool.list response structure). Unclear if this is actually an MCP server.