Fast local ML inference for agentic systems using Google Coral Edge TPU. Provides TPU-accelerated image classification, visual embeddings, audio classification, text embeddings, importance scoring, and anomaly detection.
The server provides 6 tools with generally clear naming and adequate parameter schemas, but suffers from significant documentation gaps and schema incompleteness. Tool descriptions are present but often lack depth about when to use the tool vs. alternatives, prerequisites, and return structure. Parameter descriptions are sparse, many lack explanation of expected formats, ranges, or constraints. Output schemas are not documented, forcing LLMs to infer response structure. The codebase shows incomplete tool definitions in the visible source (the classify_image tool description is cut off mid-sentence). Error handling is minimal, no recovery guidance, no categorization of error types, and no actionable error messages are evident in the code. Security practices are sound (no exposed credentials), but the tools lack the LLM-optimization details (dependency hints, alternative suggestions) found in A-grade production tools. The server targets a specialized hardware domain (TPU inference), which is niche and well-scoped, but the tool definitions do not compensate for the narrow use case with exceptional clarity.
Classify an image using TPU. Returns top-k class predictions with confidence scores. Fast (~15ms) local inference.
Classify user intent for command routing. Determines if input is a question, command, statement, etc.
Generate semantic embedding for text. Uses fast CPU model (MiniLM-L6). Returns 384-dim vector for similarity search.
Extract visual feature embedding from an image using TPU. Returns a vector that can be used for similarity search or clustering.
Score the importance/salience of content for memory prioritization. Uses semantic analysis to determine if content should be prioritized for storage.
Get TPU status and inference statistics
Tool descriptions lack LLM-critical context: no explanation of WHEN to use each tool, no hints about prerequisites or dependencies, no actionable error recovery guidance. E.g., 'Classify user intent for command routing' does not explain the intent categories, the output format, or how to handle ambiguous inputs.
Output schemas are not documented anywhere in the visible code. LLMs cannot plan downstream tool calls or extract the right data without knowing the response structure. E.g., classify_image returns 'top-k class predictions with confidence scores' but the exact JSON structure is invisible.
Parameter descriptions are generic or missing format/constraint details. E.g., 'model' enum shows two options (mobilenet_v2, efficientnet_s) but the description does not explain when to choose one over the other, which is faster, or which is more accurate. 'top_k' parameter lacks guidance on typical values or performance impact.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No error handling or recovery guidance. If an image is invalid, text embedding fails, or TPU is unavailable, the tools provide no actionable next steps for the LLM. No distinction between retryable errors (transient service failure) and user-fixable errors (invalid image format).
Tool definitions appear incomplete in source. The classify_image description is cut off mid-sentence ('Classify an image using TPU. Returns top-k class predictions with confidence scores. Fast (~15ms) local inference.', no closing details about what happens when the image is invalid or the model fails). This suggests copy-paste or truncation errors.
Parameters that accept file input (image_base64, image_path) lack validation constraints. E.g., 'image_base64' has no description of maximum size, supported formats (JPEG/PNG only?), or what happens if a non-image is provided. 'image_path' lacks file existence or permission error guidance.
score_importance and classify_intent have minimal descriptions (40 chars each) that do not explain the output range, scale, categories, or use in a reasoning pipeline. 'Score the importance/salience of content for memory prioritization' is too vague, does it return 0-1, 0-100, or a category?
No pagination guidance for tools that could return large results (e.g., embed_text with 'texts' batch parameter). No limit specified, no documented max batch size, no guidance on behavior if the batch is too large.