A collection of modular skills and tools for Claude Code, including domain-driven design modeling, news extraction, GitHub trending analysis, LangSmith dataset/trace management, and evaluation utilities.
ClawForge is a collection-focused MCP server with 16 tools across GitHub trending, LangSmith dataset/evaluator/trace management, and news extraction. Definitions show moderate quality but with significant gaps: parameter descriptions are present for most tools, but many descriptions are terse (15-40 chars), schema completeness varies, and error handling is not visible. News extraction tools lack output schema documentation. LangSmith tools have reasonable parameter sets but descriptions are often generic. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are declared. Tools are well-scoped (single responsibility), but descriptions don't consistently explain WHEN to use them or what prerequisites exist.
Delete an evaluator by name from LangSmith.
Export LangSmith dataset to a local file.
Extract news content from Netease News (网易新闻) articles.
Extract news content from Sohu News (搜狐新闻) articles.
Extract news content from Tencent News (腾讯新闻) articles.
Extract news content from Toutiao (今日头条) articles.
Extract news content from WeChat public account articles.
Fetch and basic extract GitHub Trending repos. Takes a span parameter (daily, weekly, monthly) and returns GitHub trending repositories.
News extraction tools (extract_news_from_*) lack documented output schemas. Descriptions state 'Extract news content...save to JSON' but do not document the structure of returned data (fields, types, nested objects). LLMs cannot plan downstream processing without knowing what fields to extract.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are declared in any tool. This is a current MCP spec requirement for production tools. Example: 'delete' and 'upload' tools should declare destructiveHint=true and writeHint=true respectively so clients can warn before execution.
Generic and terse tool descriptions. 'list_datasets' → 'List all LangSmith datasets...' (47 chars) lacks WHEN to use it and what it enables downstream. Best practice: 'Retrieve all available datasets from LangSmith with names, IDs, and counts. Call this first to discover dataset names before querying with show() or exporting with export().' (description avg in rubric: 194 chars).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
List all LangSmith evaluators with their names, sampling rates, and target configurations.
List all LangSmith datasets with their names, IDs, descriptions, and example counts.
Show recent traces from LangSmith.
Show examples from a LangSmith dataset by name.
Analyze and show the structure of a dataset file (JSON or CSV), displaying field names and population statistics.
Fetch and display a specific trace by ID from LangSmith.
Upload an evaluator from a Python file to LangSmith.
View examples from a local dataset file (JSON or CSV).
No visible error recovery guidance in parameter descriptions. Tools like 'show' accept 'dataset_name' but lack hints like 'If the exact name is unknown, call list_datasets() first to discover available names.' This forces agents to guess the next step on failure.
'list' tool name is too generic and conflicts with 'list_datasets'. Both list things from LangSmith. Naming should clearly distinguish: 'list_evaluators' vs 'list_datasets'. LLMs confuse similar names when deciding which tool to call.
News extraction tools accept 'save_path' parameter with default 'data/', but no schema information specifies valid paths, absolute vs. relative, or what happens if the directory doesn't exist. Minimal parameter validation guidance.