Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
RSSidian presents a mixed picture: tool names follow verb-first conventions and are generally clear, but descriptions lack depth and fail to address key LLM decision criteria. Most parameters have type definitions and ranges, but parameter descriptions are minimal or absent. Output schemas are not documented in the source code provided, forcing inference. Critically, several tools exhibit ambiguous naming (search vs search_semantic are near-duplicates), and two tools (mcp_discovery, mcp_list_subscriptions) appear to be API endpoints rather than MCP tool definitions. Error handling is not visible in the code. The 12-tool portfolio is well-intentioned but falls short of production-grade quality for agent integration.
Duplicate/overlapping tools: 'search' and 'search_semantic' are nearly identical in purpose (semantic search with query, relevance, max_results, refresh params). LLMs cannot reliably distinguish when to use each. Consolidate into one tool or clearly differentiate by documented use case.
Inadequate tool descriptions. Examples: 'Search articles using semantic search' (39 chars) and 'List articles with pagination' (29 chars) lack context on WHEN to use the tool vs alternatives, what it returns, or what data model it assumes. LLMs cannot distinguish similar tools without richer intent documentation. Baseline: 194 chars average for A+ tools.
Parameter descriptions missing or minimal. 'feed_title' parameter in mute_subscription, unmute_subscription, enable_peer_through, disable_peer_through appears in schema but has no explanation of format expectations (e.g., case-sensitive? exact match required?). 'article_id' in get_article lacks range, format, or validation hint.
Recommendations
Consolidate search and search_semantic into a single, well-documented tool. If both are necessary, rename to clarify the distinction (e.g., search_by_keyword vs search_by_embedding) and document when to use each.
Expand tool descriptions to 100 - 200 characters, answering: What does this tool do? When should the LLM use it instead of a similar tool? What data does it require and return? Example for list_articles: 'Retrieve a paginated list of RSS articles sorted by most recent. Use this to discover articles before applying semantic search. Returns article metadata including title, summary, and quality scores. Pagination: limit (1 - 1000, default 100) and offset (0 - max articles). Use offset + limit for multi-page browsing; no total count is provided.'
Add descriptions to all parameters. For feed_title: 'Exact title of the feed (case-insensitive match against list_subscriptions() output).' For article_id: 'Unique integer identifier of an article (obtained from list_articles or search results).' For lookback_days: 'Number of days to look back for articles (1 - 365; longer periods may timeout).'
Document output schemas inline in tool definitions or in a companion schema document. For each tool, state: what fields are returned, field types, and which fields chain to other tools. Example for get_article: 'Returns {id: int, feed_id: int, feed_title: string, title: string, url: string, summary: string|null, published_at: ISO8601 string|null, ...}. Use feed_id or feed_title to reference list_subscriptions().'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Output schemas not documented in source code. While api.py defines Pydantic response models (ArticleResponse, FeedResponse, SearchResponse), these are not visible in the tool definition/registration code itself. LLMs cannot infer field names, types, or which fields chain to subsequent tools. Rule: '100% of A+ tools have documented return types'.
MCP-specific tools (mcp_discovery, mcp_list_subscriptions) appear to be REST API endpoints, not legitimate MCP tool definitions. 'mcp_discovery' returns 'information about available endpoints and capabilities', this sounds like metadata, not an agent-invoked tool. These may belong in the MCP server's capabilities announcement, not as executable tools. Verify whether these should exist at all.
No error handling guidance visible. Code (api.py) raises HTTPException(404) in get_article, mute_subscription, etc., but no recovery hints are provided to the LLM (e.g., 'Feed not found. Try list_subscriptions() to see available feeds.'). Per pattern: 'Error responses must tell the LLM what to do next.'
Opaque parameter values: enable_peer_through and disable_peer_through require 'feed_title' as a string, but no guidance on how to obtain valid feed titles or whether case sensitivity matters. Users operate with visible feed names, but LLMs need to know lookup strategy. Suggest: 'feed_title must match exactly (case-insensitive) a title returned by list_subscriptions().'
ingest_feeds accepts 'lookback_days' (integer) with no range constraint visible. Unbounded integers allow LLMs to pass absurd values (e.g., 999999 days) that could hang the service. Declare min/max (e.g., 1 - 365) in both schema and description.
Pagination not explicitly documented. list_articles and list_subscriptions accept limit/offset but descriptions do not state: (a) what is the maximum safe limit? (b) is a total count returned? (c) what happens if offset exceeds total? LLMs need this to navigate multi-page results correctly.
Tool naming ambiguity: 'enable_peer_through' and 'disable_peer_through' assume LLMs understand 'peer-through' semantics (fetch origin article content via aggregator feed). Without context, the names are opaque. Better names: 'fetch_original_article_via_aggregator' or provide rich descriptions explaining the benefit and when to use.
enable_peer_throughdisable_peer_through
Remove or clarify mcp_discovery and mcp_list_subscriptions. If they are redundant with MCP's native capability announcement, delete them. If they serve a purpose (e.g., dynamic discovery of new feeds added at runtime), rename to describe the action (e.g., get_available_feeds) and document what makes them different from list_subscriptions().
Add error recovery guidance to tool descriptions. Example: 'If feed not found (404), call list_subscriptions() to see available feeds and verify the exact feed_title.'
Document pagination behavior explicitly. For list_articles and list_subscriptions, state: 'Pagination uses limit (items per page) and offset (items to skip). No total count is provided. Iterate by incrementing offset until fewer than limit items are returned, indicating the final page.'
Create an enum or constraint for relevance thresholds if there are common values (e.g., 'low: 0.3', 'medium: 0.6', 'high: 0.8'). Or document in the description: 'Typically 0.3 - 0.7; lower values return more matches but lower confidence.'
Add response examples to tool descriptions. Example for get_article: 'Returns an article object with fields id (int), title (string), summary (string, may be null), url (string), published_at (ISO 8601 timestamp), quality_tier (string: "high"|"medium"|"low"), and labels (comma-separated string, may be null).'
Implement and document error classification. Tools should return: {'error': 'feed_not_found', 'message': 'Feed with title "xyz" not found. Available feeds: ["feed1", "feed2"]', 'recovery': 'Call list_subscriptions() to see all feeds.'} instead of bare HTTP exceptions.
For write operations (mute_subscription, unmute_subscription, enable_peer_through, disable_peer_through, ingest_feeds), document idempotency and side effects. Example: 'Muting a feed is idempotent, calling it multiple times with the same feed_title has no additional effect. Muted feeds are excluded from future ingest_feeds runs.' This helps LLMs decide whether retries are safe.
Add descriptions of what each parameter controls in write operations. Example for enable_peer_through: 'Enables fetching the original article content from the source feed when the current feed is an aggregator. This enriches summaries and quality scores. Idempotent, enabling an already-enabled feed has no effect.'