Precision-driven Tool Recommendation system for filtering MCP tools based on conversation context using semantic search, BM25, cross-encoder reranking, and learning-to-rank models
The server provides a single tool (filter_tools) with a basic description and partial schema. Critical issues: (1) The tool name 'filter_tools' is generic and does not follow verb_noun conventions, a better name would be 'recommend_tools', 'rank_tools', or 'search_relevant_tools'. (2) The description is extremely minimal (97 characters) and lacks essential detail about WHEN to use this tool, WHAT it returns, and any prerequisites or context requirements. (3) The input schema is present but severely underspecified: the 'messages' parameter lacks format/structure documentation, 'available_tools' has no clarity on expected object shape or required fields, and 'max_tools' has no min/max bounds or default value. (4) There is NO documented output schema, the response structure is completely undocumented, forcing LLMs to infer what fields to expect. (5) No error handling guidance is provided. (6) The tool appears to be a filtering/ranking utility intended to reduce tool bloat for downstream LLM calls, but this intent is buried in the description and not surfaced clearly. (7) The implementation code visible (toolbench_evaluator.py) shows the tool is part of a complex evaluation pipeline with vector stores and embeddings, but none of this architectural context appears in the tool definition, leaving users/LLMs confused about dependencies or behavior.
Filter tools based on conversation context. Analyzes the conversation messages and returns the most relevant tools from the available tool set.
Tool name 'filter_tools' is generic and violates verb_noun naming. LLMs cannot infer the action or when to select this tool over similar ranking/recommendation tools. Recommend rename to 'recommend_tools', 'rank_tools_by_relevance', or 'filter_tools_by_context'.
Description is only 97 characters and lacks critical guidance. Does not explain WHAT the tool does (filter? rank? recommend?), WHEN to use it (when LLM has too many tools?), WHAT it returns (ordered list? scored items?), or ANY prerequisites. LLM cannot reliably decide whether to call this tool.
Input parameter 'messages' has no description of expected format. Is this an array of strings? Objects with role/content fields? Conversation history? LLM will guess and likely pass wrong structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 35 | <=2025-11-25 | v2 |
| 2026-04-07 | F | 22 | - | v1 |
Input parameter 'available_tools' has no description of expected object shape. Are these tool names (strings), tool objects with {name, description, schema}? Does the tool accept the MCP tool definition format? Completely ambiguous.
Input parameter 'max_tools' has no bounds documentation, default value, or description of behavior when set to 0 or negative. Missing critical constraint information.
NO output schema is documented. The tool description does not specify what fields the response contains, whether it returns a list, the structure of each item, or how to extract tool names/scores/confidence. LLM cannot plan downstream calls or extract filtered tool list.
No error handling guidance. What happens if 'messages' is malformed? If 'available_tools' is empty? If the embedding service fails? LLM receives no recovery instructions.
Tool purpose is ambiguous from the interface. Is this a tool-selection helper (filter N tools down to top K)? A context-aware recommendation engine? A retrieval-augmented ranking system? The name, description, and schema do not align on intent, creating confusion.