Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
Terrorblade defines 3 tools with reasonable naming (verb_object pattern) and documented schemas. All tools have descriptions and parameter types. However, descriptions lack LLM-optimization depth, parameter descriptions are sparse or generic, error handling guidance is missing, and output schemas are not explicitly documented. The tools are read-only and relatively well-scoped, but fall short of production-grade quality due to incomplete parameter documentation and absent error recovery patterns.
Tools (3)
cluster_searchread onlysource verified69/100
Find the most relevant conversation clusters for a query by aggregating top vector hits. Returns cluster summaries including best similarity, number of hits, chat name, and a snippet.
get_clusterread onlysource verified70/100
Retrieve all messages for a specific cluster (group_id) within a chat.
vector_searchread onlysource verified72/100
Perform semantic vector search over Telegram messages for a user.
Output schemas not documented. LLMs cannot plan downstream tool calls or extract the right data without knowing what vector_search, cluster_search, and get_cluster return.
Parameter descriptions are sparse and lack constraints. E.g., 'phone' accepts multiple formats but LLM has no guidance on which to use. 'similarity_threshold' accepts float but no range documented (0.0-1.0? unbounded?). 'top_k' and 'max_clusters' lack upper bounds.
Error handling provides no recovery guidance. Tools do not document what happens on db_path not found, invalid phone, query too short/long, similarity_threshold out of range, chat_id not found, or group_id mismatch. LLM receives no actionable error recovery hints.
vector_searchcluster_searchget_cluster
Recommendations
Document the return type and structure for each tool. E.g., vector_search: 'Returns a list of {message_id, text, timestamp, similarity_score, chat_name} objects, sorted by similarity descending.' cluster_search: 'Returns a list of {cluster_id, summary, best_similarity, hit_count, chat_name} objects.' get_cluster: 'Returns a list of {message_id, text, timestamp, sender_id} objects in chronological order.'
Add constraints and range guidance to all numeric/string parameters. E.g., 'similarity_threshold: float from 0.0 (any match) to 1.0 (exact match), default 0.0' and 'top_k: integer from 1 to 100, default 10'.
Enhance descriptions with WHEN/WHY and composition hints. E.g., vector_search: 'Use this to find individual messages matching a natural language query. Returns raw messages with similarity scores. For conversation clusters, call cluster_search() instead.' cluster_search: 'Aggregates vector search results into conversation clusters. Call this first for high-level topic discovery, then get_cluster() to retrieve full message threads.'
Add error handling documentation to each tool description. E.g., 'If db_path not found, returns error with hint to check configuration. If phone invalid, returns error with expected format (+1234567890 or 1234567890). If no results match similarity_threshold, returns empty list (not an error).'
Document pagination for tools returning lists. E.g., 'vector_search: Results capped at top_k (default 10, max 100). For larger result sets, adjust top_k. If cluster has >1000 messages, results truncated, call cluster_search first to identify relevant clusters, then get_cluster to retrieve full context.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Tool descriptions lack WHEN/WHY context and composition hints. 'Perform semantic vector search' does not tell LLM when to prefer vector_search over cluster_search, whether it returns raw messages or summaries, or if results need post-processing.
No pagination guidance for large result sets. If a cluster has 10,000 messages or vector_search returns 1000+ results at low similarity threshold, tool does not document how to paginate or warn of context window risks.
Parameter 'db_path' with default 'auto' invites silent failures. If default path does not exist or is misconfigured, LLM receives no guidance on how to recover or specify an explicit path.
Tool descriptions are under 200 chars (baseline for good LLM-optimized descriptions). vector_search (78 chars), get_cluster (94 chars), and cluster_search (156 chars) lack depth and fail to fully contextualize when each tool should be called.
vector_searchcluster_searchget_cluster
Add parameter descriptions for 'phone' that clarify both formats. E.g., 'phone: User phone identifier as string, accepted formats: +1234567890 (with +) or 1234567890 (without +, leading digits inferred as country code).'
Add parameter descriptions for 'db_path' explaining the auto behavior. E.g., 'db_path: Path to DuckDB database file. Use "auto" to use default location ($HOME/.terrorblade/messages.duckdb), or "default" as alias for auto. Provide explicit path to use alternate database.'
Document the relationship between include_cluster_messages (vector_search) and whether get_cluster should be called. E.g., 'include_cluster_messages: If true, each result includes a compact snippet of surrounding messages for context. Set to false if you plan to call get_cluster() for full context.'
Add tool annotations if supported by fastmcp. E.g., readOnlyHint=true on all three tools to clarify they do not modify state.
Consider adding a discovery tool like 'list_chats' or 'describe_database_schema' to help LLMs understand available data without trial-and-error calls.