MCP server that compresses prompts 40-60% using local LLM + embedding validation
The server exposes a single tool 'compress_prompt' with a clear verb-noun naming convention. The tool description is comprehensive (180+ characters) and explains what it does, when to use it, and dependencies. However, the input schema is minimal (single 'text' parameter with basic type and description), and there is NO documented output schema, the response is unstructured text with embedded metadata rather than a typed object. The error handling returns the original text gracefully but does not guide the LLM on recovery paths or categorize errors. The server lacks tool annotations (readOnlyHint, destructiveHint, idempotentHint) and does not follow structured output patterns used by A-grade tools. The implementation is functional but falls short of production-grade quality standards for agent-facing tools.
Compress a prompt to its semantic minimum using a local LLM and embedding validation. Reduces token usage 40-60% while preserving all conditionals, negations, and critical intent. Returns the original text unchanged if validation fails. Requires Ollama running locally with llama3.2:1b and nomic-embed-text.
No output schema documented. Response is unstructured text with embedded metadata ('mode: X | coverage: Y% | tokens: A→B'). LLMs cannot parse this reliably or extract structured fields for downstream operations.
Input schema is minimal. Single 'text' parameter with basic type string and short description. No constraints on length, format, or maximum input size. No guidance on what happens if text is empty or extremely large (performance implications).
No tool annotations present. The tool is read-only and idempotent (compression always produces the same output for same input), but lacks readOnlyHint and idempotentHint to signal this to the client and agent.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | 2024-11-05+ | v1 |
Error handling is graceful but unhelpful to agents. On exception, the server returns the original text with a comment '[token-compressor error: X, returned original]'. This does not categorize the error as retryable, user-fixable, or fatal, nor does it guide recovery (e.g., 'Check that Ollama is running with llama3.2:1b and nomic-embed-text models').
External dependency not enforced. The tool requires Ollama with two specific models (llama3.2:1b, nomic-embed-text) but does not validate their availability at startup or call time. A missing or misconfigured Ollama will fail silently with a vague error message.