AI research paper ingestion and scoring pipeline with MCP server integration for paper discovery, assessment, and sharing workflow
Paper Scout has solid tool naming (all verb-first) and mostly good descriptions (all >20 chars), but suffers from incomplete parameter descriptions, missing output schema documentation, and absent error handling guidance. 8 of 10 tools have complete input schemas with types; parameter descriptions are present but often generic. No tools document their return structure, forcing LLMs to infer outputs. Security practices are adequate (no exposed secrets), but error recovery paths are not described. Composition is reasonable, tools chain logically (get_papers → assess_significance → generate_summary_context → record_share). Overall, definition quality is fair-to-good but lacks the polish required for A-grade production confidence.
Get structured context for assessing a paper's significance. Returns the paper's full details, scores, related papers in the same topic cluster, and relevant interest profile sections. Use this context to make a judgment about the paper's importance.
Get paper details plus past sharing examples for a given platform. Returns everything needed to write a summary in the user's established style.
Get the current interest profile including topics, keywords, tracked labs, and tracked authors.
Get full details for a specific paper including all scores, author info, and citation data.
Get the status of the ingestion pipeline: last run time, stats, next scheduled run, and any recent errors.
Get today's candidate papers sorted by relevance score. Returns papers that scored above the candidate threshold from the most recent pipeline run.
No output schema documentation for any tool. LLMs cannot plan multi-step sequences when they don't know what fields each tool returns. This violates pattern:tool and forces agents to make unsafe assumptions about response structure.
Parameter descriptions are present but often trivial or incomplete. E.g., 'min_score' in get_todays_papers says 'Minimum composite score filter (0.0 to 1.0)' but doesn't explain what 'composite score' means or when to use this filter. 'notes' parameter in update_paper_status lacks guidance on what kind of notes are useful.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Record that a paper was shared on a platform, with the summary used. Feeds back into scoring to improve future relevance.
Full-text search across all stored papers (not just today's candidates).
Add or remove items from the interest profile (topics, keywords, tracked labs, or tracked authors).
Update a candidate paper's review status.
No error handling guidance. Tools do not describe failure modes, recovery steps, or what the LLM should do if a paper_id is invalid, a database query fails, or the profile update violates constraints. Pattern:recovery-guide requires error responses to guide next steps.
get_interest_profile and get_pipeline_status have empty input schemas (only {'type': 'object', 'properties': {}}). While valid, they lack the structured output documentation to tell the agent what fields to expect. This risks the agent failing to extract needed data.
Pagination handling is weak. search_papers and get_todays_papers accept 'limit' but do not describe total counts, cursors, or how to iterate. Large result sets could exhaust context. Pattern:paginated-result requires pagination metadata.
No idempotency guarantees documented. Destructive operations like update_paper_status and record_share do not state whether they are safe to retry. If an LLM retries due to ambiguous failure, it could double-record a share or change status twice.
No confirmation or dry-run pattern for destructive operations. update_paper_status and record_share modify state without a safety step. Pattern:confirmation-request recommends a way for agents to preview actions before committing.
Tool composition is good, but the response from get_todays_papers should include all IDs needed for downstream calls (e.g., paper_id, status, score). Without seeing the actual return structure, we cannot verify that assess_significance can be called immediately after. Pattern:tool-chain requires output to feed input.