An MCP server that scrapes Reddit subreddits, creates vector embeddings using Pinecone, searches the vector database, and generates AI responses using Google Gemini
The server has 5 tools with schemas and descriptions present, but multiple critical gaps significantly reduce quality. Tool names are moderately clear but lack consistency in verb-first patterns. Descriptions exist but are often generic and under-optimized for LLM reasoning. Critical issues: (1) No input validation or constraint documentation (enums, ranges, formats) despite accepting user-controlled parameters like topic strings and subreddit/post limits. (2) Output schemas are not documented, tools don't declare what fields agents receive back, forcing LLMs to guess. (3) Error handling is minimal, no recovery guidance or error categorization. (4) Parameters lack specific format/range constraints (e.g., 'subreddit_limit' and 'post_limit' accept arbitrary integers with no bounds documented). (5) Secrets are exposed in source code (yaml config files), not injected server-side. (6) No permission gates or audit trails for data-modifying tools. (7) Dependencies between tools (scrape_reddit_dynamic → create_vector_database → search_vector_db) are not documented, forcing agents to discover correct composition.
Creates a Pinecone vector index with integrated embeddings using llama-text-embed-v2 model and upserts Reddit posts data into the index.
Complete pipeline that scrapes Reddit for a topic, creates a vector database, searches it, and generates a Gemini response. Returns cached results if already processed today.
Uploads vector search results to Google Gemini API and generates a content response based on the query and search results.
Fetches subreddits matching the given topic and returns a list sorted by subscriber count. Scrapes Reddit posts and comments from matching subreddits and saves them in Pinecone-formatted JSON.
Searches a Pinecone vector database index using a text query and saves the top_k results to a file.
Output schemas not documented. Tools return results (e.g., scrape_reddit_dynamic returns a path, create_vector_database returns success/failure) but the response structure, field names, and types are not declared. LLMs cannot plan downstream steps without knowing what they receive back.
No input validation constraints. Parameters like 'topic' (free-form string), 'subreddit_limit', and 'post_limit' lack documented ranges, patterns, or enums. No maximum or minimum bounds shown, inviting LLMs to pass invalid values (negative limits, extremely large numbers). The rubric requires explicit format/range declarations.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 27 | - | v1 |
Secrets exposed in source code. The codebase loads API keys from 'secrets.yaml' hardcoded in tool functions (src/create_vector_database.py, src/gemini_retrieval.py, src/ingestion_dynamic.py), not injected via environment variables or server-level vault. Violates pattern:secret-injection, credentials must never appear in tool parameters or source files.
No error recovery guidance. Tools may fail (network errors, API limits, missing files) but do not return actionable error messages. E.g., scrape_reddit_dynamic catches exceptions with generic 'print(f"Error fetching from r/{subreddit.display_name}: {e}")' without telling the LLM what to try next.
Tool composition not documented. fetch_and_search orchestrates scrape_reddit_dynamic → create_vector_database → search_vector_db → return_gemini_response, but intermediate dependencies (e.g., scrape_reddit_dynamic must run before create_vector_database, which expects 'raw_scrap_results.json' at a specific path) are undocumented. Agents cannot infer the correct call sequence.
Descriptions are generic and under-optimized for LLM reasoning. E.g., 'Creates a Pinecone vector index with integrated embeddings using llama-text-embed-v2 model and upserts Reddit posts data into the index' does not answer WHEN to use it or what PREREQUISITES must hold (e.g., raw_scrap_results.json must exist at current_run_path).
Parameter descriptions lack specificity. 'topic' in scrape_reddit_dynamic is described as 'The topic to search for subreddits', does not clarify format (free text, hashtag, keywords), length limits, or examples. 'subreddit_limit' lacks bounds (minimum 1? maximum 100?).
Data-modifying tools lack confirmation or dry-run support. create_vector_database and fetch_and_search perform irreversible writes (Pinecone upserts, file creation) but do not offer a dry-run or confirmation step. Per pattern:confirmation-request, agents should be able to preview side effects before committing.
Tool naming inconsistency. Names use verb_noun (scrape_reddit_dynamic, search_vector_db, return_gemini_response) but 'return_gemini_response' uses 'return' (passive) instead of 'generate' (active). 'scrape_reddit_dynamic' includes '_dynamic' suffix with no explanation.
No pagination or result limits documented. search_vector_db accepts a 'top_k' parameter (default 100) but the description does not warn that returning 100+ results could exhaust context or degrade LLM reasoning.
No permission gates or audit trails. Tools invoke external APIs (Reddit, Pinecone, Google Gemini) without checking user authorization or logging who called what, when, and with what parameters. Per pattern:permission-gate and pattern:audit-trail, sensitive operations must be gated and traced.