MCP server for searching and analyzing Hacker News stories and discussions using regex patterns and AI summaries
The server has 3 tools with strong descriptions (all 150+ chars), clear naming with action verbs (search_, read_, summarize_), and documented input parameters. However, critical gaps undermine overall quality: (1) Output schemas are NOT documented anywhere, the code shows tool_query projections but no explicit return schema documentation for the LLM; (2) Parameter descriptions exist but lack format/constraint details (e.g., regex pattern syntax hints are vague, no mention of case sensitivity behavior, no guidance on common search patterns); (3) No error handling guidance, tools silently fail on invalid regex or missing stories without recovery hints; (4) Missing tool annotations (readOnlyHint, idempotentHint) despite all tools being read-only; (5) No pagination support documented despite result_limit hints in code; (6) Composition is good (tools chain well: search → read → summarize), but missing per-item result limits and next_cursor guidance. The server reads clearly at intermediate quality, strong fundamentals but incomplete production details.
Fetch a complete Hacker News story with its entire comment thread. Returns one row per comment/story, maintaining hierarchical structure via depth and path columns. The path column uses zero-padded IDs (e.g., '0000012345.0000012346') to maintain thread ordering.
Search Hacker News stories using regex patterns across titles, URLs, story text, and all related comments. Returns stories ranked by relevance (title > URL > text > comments) and sorted by date. Always searches comments to find all relevant discussions. Useful for finding conversations about specific topics, technologies, companies, or keywords.
Get an AI-generated summary of a Hacker News story and its discussion thread, including key points, themes, and sentiment analysis.
NO documented output schemas for any tool. Code shows projections and denormalized tables, but the LLM cannot see what fields are returned. This violates pattern:tool-description and pattern:response-shaper, LLMs must know return structure to chain calls and extract data.
Parameter descriptions lack format constraints. 'pattern' parameter in search_stories has syntax examples in comments but no formal description of regex dialect, case-sensitivity default, or escape-character handling. LLMs cannot infer Python regex behavior.
No error handling guidance. Tools have no documented failure modes. What happens if regex is invalid? Story ID doesn't exist? API timeout? LLMs have no recovery path.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
No tool annotations (readOnlyHint, idempotentHint) despite all tools being read-only. Missing metadata that helps agents reason about safety and retry semantics.
Pagination not explicitly documented. Code hints at result_limit (100, 2000) but tools lack next_cursor or offset/limit parameters. If a story has 5000+ comments, read_story silently truncates without signaling incompleteness to the LLM.
search_stories result_limit of 100 may be high for LLM context. Baseline best practice is 20 - 50 results. Returning 100 items dilutes signal and increases hallucination risk. No mention of stripping irrelevant metadata (e.g., rank_score internals).