A tiny Model Context Protocol (MCP) server that pools free-tier LLMs across multiple providers with automatic failover, routing, and quota management. Enables Claude Desktop, Claude Code, Cursor, and other MCP clients to offload subtasks to free models.
The server exposes 13 tools with consistent naming patterns (verb_noun using 'free_llm_*' prefix), each with descriptions. However, critical gaps undermine quality: (1) Parameters largely lack type annotations, most are inferred from the schema JSON but descriptions don't always clarify constraints, enums, or ranges; (2) Output schemas are not documented, tool descriptions state WHAT they return but not the field structure or typing; (3) Error handling is implicit, with no recovery guidance in descriptions; (4) Many parameters have generic or incomplete descriptions (e.g., 'routing' enum values are defined in schema but descriptions are sparse on when to use each). The baseline for A-grade tools requires 100% param descriptions and documented return schemas, this server achieves ~70% param coverage and 0% documented output structure. Naming is strong (clear intent, consistent prefix), but schema and description quality lag production standards.
Ask a free LLM (pooled across configured free providers, with automatic failover). Offload a self-contained subtask — drafting, summarizing, classifying, brainstorming, a quick lookup — to a free model. The reply tells you which provider/model actually served it.
Compare a prompt across a small bounded panel of free models and render the result as a Markdown comparison table. Bounded to a few models (2-5) — for the broader eligible-target stress test use `tokenmax`. Per-model failures stay visible in the rendered output rather than failing the whole call.
List available provider/model ids. Returns the catalog of all configured providers and their available models.
Ask the SAME prompt to several different free models at once and get every answer back side by side - the agent-facing second-opinion surface. Great for cross-checking a fact, comparing approaches, or reducing single-model bias. Optionally have a strong model synthesize the best combined answer.
Today's per-provider usage + daily-limit headroom. Returns current quota status across all configured providers.
No documented output schemas. Tool descriptions state what results are returned (e.g., 'panel into one best answer', 'comparison table', 'routing analysis') but do NOT document the field structure, types, or keys of the response. LLMs cannot plan downstream tool calls or extract typed data without knowing output structure.
Parameter descriptions lack clarity on constraints and formats. 'routing' enum is defined in schema but description text does not explain when to use 'agent' vs 'quality' vs 'fast', LLMs must infer strategy differences. 'max_tokens' lacks min/max bounds. 'n' (panel size) states '2-8' in description but type is integer with no JSON schema min/maxItems.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 58 | <=2025-11-25 | v2 |
Local quota-mode / headroom advice (no bypass suggestions). Returns quota status and capacity recommendations based on local configuration and usage.
Run a bundled recipe (panel/text) end-to-end. Recipes are curated workflows that apply freellmpool routing to multi-model comparison, PR review, repo summarization, and other agentic tasks.
List available roles and recommended use. Roles are semantic groupings of models by capability and intended deployment context.
Explain where a prompt WOULD route (difficulty + ranked models), $0. Returns routing analysis showing prompt difficulty classification and which models would serve it, without making an actual request.
Run the shared small-panel second-opinion flow. Same behavior as `free_llm_panel` — exposed as its own agent-facing tool so callers can declare their intent (a *second opinion*) without reasoning about the panel primitive. The panel size is bounded (2-5) and the per-model token budget is clamped.
Lifetime tokens served free + estimated cost avoided. Returns aggregate statistics on tokens served and cost savings across the lifetime of the freellmpool instance.
Safe Tailscale Tailnet connection instructions. Returns setup hints and safe base URLs for connecting to a freellmpool instance via Tailscale.
Blast the prompt to a swarm of models; you synthesize them all. A bounded multi-model stress test that sends a prompt to the broadest eligible target set and returns every answer for the agent to synthesize, rendered with a rainbow ANSI banner.
No error handling or recovery guidance in tool descriptions. Tools mention failures are possible (e.g., free_llm_battle states 'per-model failures stay visible') but descriptions do NOT explain what constitutes an error, how to recover, or when to retry. LLMs have no guidance on next steps if a model fails.
Three tools (free_llm_models, free_llm_quota, free_llm_stats) have minimal descriptions (50 chars or fewer: 'List available...', 'Today's per-provider...', 'Lifetime tokens...'). These are below the 10-1024 char production baseline and lack guidance on when or why to call them versus alternatives.
free_llm_recipe has a very generic 'recipe' parameter description ('Recipe identifier (e.g. ...')') without documenting valid recipe values, their differences, or constraints. An LLM cannot determine valid recipes without external knowledge.
free_llm_ask, free_llm_panel, free_llm_second_opinion allow optional 'model' and 'provider' parameters but descriptions do not explain precedence or interaction (e.g., if both are supplied, which takes priority?). Undocumented parameter relationships cause silent misuse.