MCP server for integrating Macrocosmos SN13 social media data into Claude Desktop and Cursor
Macrocosmos MCP has significant gaps in definition quality. While tool descriptions are present and detailed, they lack proper schema enforcement, parameter descriptions are incomplete, and critical validation information is buried in prose rather than structured as enums/constraints. The server exposes 6 tools (with 3 duplicates), indicating possible design or documentation issues. Output schemas are not documented. Error handling is absent from code, no indication of how malformed input or API failures are handled. The source code is incomplete (pyproject.toml and __init__.py shown but only partial macrocosmos_mcp.py), so direct verification of schema registration and error handling is limited.
Create a Gravity task for large-scale data collection from X (Twitter) or Reddit. Use this for collecting large datasets over time (up to 7 days). For quick queries (up to 1000 results), use query_on_demand_data instead. The task registers on the network within 20 minutes and collects data for 7 days. You'll receive an email notification when the dataset is ready for download. Parameters: - tasks (List[dict], REQUIRED): List of task objects, each containing: * platform (str): 'x' or 'reddit' * topic (str): The hashtag/subreddit to monitor - For X: MUST start with '#' or '$' (e.g., '#ai', '$BTC') - plain keywords are rejected! - For Reddit: subreddit name (e.g., 'r/MachineLearning') * keyword (str, optional): Additional keyword filter within the topic - Filters posts to only those containing this keyword - Example: topic='#Bittensor', keyword='dTAO' -> only #Bittensor posts mentioning 'dTAO' - name (str, optional): Name for the task (helps organize multiple tasks) - email (str, optional): Email address for notification when dataset is ready - redirect_url (str, optional): URL to redirect to from the email notification Returns: - gravity_task_id: Unique identifier to track and manage the task Examples: 1. Basic collection: create_gravity_task( tasks=[{"platform": "x", "topic": "#ai"}], name="AI Tweets" ) 2. With keyword filter: create_gravity_task( tasks=[{"platform": "x", "topic": "#Bittensor", "keyword": "dTAO"}], name="Bittensor dTAO mentions" ) 3. Multiple platforms: create_gravity_task( tasks=[ {"platform": "x", "topic": "#ai", "keyword": "LLM"}, {"platform": "reddit", "topic": "r/MachineLearning"} ], name="AI Data Collection", email="user@example.com" )
Create a Gravity task for large-scale data collection from X (Twitter) or Reddit. Use this for collecting large datasets over time (up to 7 days). For quick queries (up to 1000 results), use query_on_demand_data instead. The task registers on the network within 20 minutes and collects data for 7 days. You'll receive an email notification when the dataset is ready for download. Parameters: - tasks (List[dict], REQUIRED): List of task objects, each containing: * platform (str): 'x' or 'reddit' * topic (str): The hashtag/subreddit to monitor - For X: MUST start with '#' or '$' (e.g., '#ai', '$BTC') - plain keywords are rejected! - For Reddit: subreddit name (e.g., 'r/MachineLearning') * keyword (str, optional): Additional keyword filter within the topic - Filters posts to only those containing this keyword - Example: topic='#Bittensor', keyword='dTAO' -> only #Bittensor posts mentioning 'dTAO' - name (str, optional): Name for the task (helps organize multiple tasks) - email (str, optional): Email address for notification when dataset is ready - redirect_url (str, optional): URL to redirect to from the email notification Returns: - gravity_task_id: Unique identifier to track and manage the task Examples: 1. Basic collection: create_gravity_task( tasks=[{"platform": "x", "topic": "#ai"}], name="AI Tweets" ) 2. With keyword filter: create_gravity_task( tasks=[{"platform": "x", "topic": "#Bittensor", "keyword": "dTAO"}], name="Bittensor dTAO mentions" ) 3. Multiple platforms: create_gravity_task( tasks=[ {"platform": "x", "topic": "#ai", "keyword": "LLM"}, {"platform": "reddit", "topic": "r/MachineLearning"} ], name="AI Data Collection", email="user@example.com" )
Duplicate tool definitions: query_on_demand_data appears twice (#1 and #4), create_gravity_task appears twice (#2 and #5), get_gravity_task_status appears twice (#3 and #6). This suggests either poor documentation or unintended registration.
Parameters with platform/source enums are NOT declared as enums in the input schema. 'source' must be 'X' or 'REDDIT' (case-sensitive), but the schema shows type:string with only a description. LLMs will hallucinate invalid values like 'twitter', 'Twitter', 'x', or 'reddit'.
create_gravity_task 'tasks' parameter is defined as type:array with items:object, but no nested schema. The description mentions required 'platform' and 'topic' fields and optional 'keyword', but the JSON Schema does not enforce this structure. LLMs cannot validate before calling.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 24 | - | v1 |
Get the status of a Gravity task and see how much data has been collected. Parameters: - gravity_task_id (str, REQUIRED): The ID of the gravity task to check - include_crawlers (bool, default: True): Whether to include detailed crawler information Set to True to see records_collected and bytes_collected for each crawler Returns: - Task status (Running, Completed, Pending, etc.) - Task name and start time - List of crawler IDs (needed for build_dataset) - When include_crawlers=True: records_collected, bytes_collected per crawler Example: get_gravity_task_status(gravity_task_id="multicrawler-9f518ae4-xxxx-xxxx-xxxx-8b73d7cd4c49")
Get the status of a Gravity task and see how much data has been collected. Parameters: - gravity_task_id (str, REQUIRED): The ID of the gravity task to check - include_crawlers (bool, default: True): Whether to include detailed crawler information Set to True to see records_collected and bytes_collected for each crawler Returns: - Task status (Running, Completed, Pending, etc.) - Task name and start time - List of crawler IDs (needed for build_dataset) - When include_crawlers=True: records_collected, bytes_collected per crawler Example: get_gravity_task_status(gravity_task_id="multicrawler-9f518ae4-xxxx-xxxx-xxxx-8b73d7cd4c49")
Fetch real-time social media data from X (Twitter) and Reddit through the Macrocosmos SN13 network. IMPORTANT: This tool requires 'source' parameter to be either 'X' or 'REDDIT' (case-sensitive). Parameters: - source (str, REQUIRED): Data platform - must be 'X' or 'REDDIT' - usernames (List[str], optional): Up to 5 usernames to monitor. * For X: '@' symbol is optional (e.g., ['elonmusk', '@spacex'] both work) * NOT available for Reddit - keywords (List[str], optional): Up to 5 keywords/hashtags to search * For X: any keywords or hashtags (e.g., ['AI', 'crypto', '#bitcoin']) * For Reddit: subreddit names (e.g., ['r/astronomy', 'space']) or 'r/all' for all subreddits - start_date (str, optional): Start date/datetime in YYYY-MM-DD or ISO format * Examples: '2024-04-01' or '2024-01-01T00:00:00Z' * Defaults to 24 hours ago from current time if not specified - end_date (str, optional): End date/datetime in YYYY-MM-DD or ISO format * Examples: '2024-04-25' or '2024-06-03T23:59:59Z' * Defaults to current time if not specified - limit (int, optional): Maximum number of results to return (range: 1-1000, default: 10) - keyword_mode (str, optional): How to match keywords - 'any' (default) or 'all' * 'any': returns posts matching ANY of the keywords * 'all': returns posts matching ALL of the keywords Default Behavior (when dates not specified): The tool searches the last 24 hours (from current time back to 24 hours ago). Usage Examples: 1. Get recent tweets from specific users: query_on_demand_data(source='X', usernames=['@elonmusk', '@spacex'], limit=20) 2. Search tweets by keywords in last 24 hours: query_on_demand_data(source='X', keywords=['AI', 'machine learning'], limit=30) 3. Monitor specific users AND filter by keywords: query_on_demand_data(source='X', usernames=['@nasa'], keywords=['space', 'mars'], limit=20) 4. Monitor Reddit subreddits: query_on_demand_data(source='REDDIT', keywords=['r/astronomy', 'space'], limit=50) 5. Search across all of Reddit with date range: query_on_demand_data(source='REDDIT', keywords=['r/all', 'space'], start_date='2025-04-01', end_date='2025-04-02', limit=50) 6. Strict keyword matching (requires ALL keywords): query_on_demand_data(source='X', keywords=['AI', 'machine learning'], keyword_mode='all', limit=30) 7. Precise datetime range search: query_on_demand_data(source='X', keywords=['Bitcoin'], start_date='2024-06-01T00:00:00Z', end_date='2024-06-03T23:59:59Z', limit=100) Returns: JSON object containing: - status: "success" or error information - data: Array of posts/tweets with full content, user information, engagement metrics, timestamps, platform-specific metadata, and media attachments - meta: Processing statistics (miners queried, response rates, items returned, etc.) Platform-Specific Notes: - X (Twitter): '@' symbol is optional for usernames - Reddit: Does NOT support username filtering, only subreddit/keyword searches - All timestamps returned in UTC format
Fetch real-time social media data from X (Twitter) and Reddit through the Macrocosmos SN13 network. IMPORTANT: This tool requires 'source' parameter to be either 'X' or 'REDDIT' (case-sensitive). Parameters: - source (str, REQUIRED): Data platform - must be 'X' or 'REDDIT' - usernames (List[str], optional): Up to 5 usernames to monitor. * For X: '@' symbol is optional (e.g., ['elonmusk', '@spacex'] both work) * NOT available for Reddit - keywords (List[str], optional): Up to 5 keywords/hashtags to search. * For X: any keywords or hashtags (e.g., ['AI', 'crypto', '#bitcoin']) * For Reddit: subreddit names (e.g., ['r/astronomy', 'space']) or 'r/all' for all subreddits - start_date (str, optional): Start timestamp in ISO format (e.g., '2024-01-01T00:00:00Z'). Defaults to 24h ago if not specified - end_date (str, optional): End timestamp in ISO format (e.g., '2024-06-03T23:59:59Z'). Defaults to current time if not specified - limit (int, optional): Maximum results to return (1-1000). Default: 10 Usage Examples: 1. Monitor Twitter users: query_on_demand_data(source='X', usernames=['elonmusk', 'spacex'], limit=20) 2. Search Twitter keywords: query_on_demand_data(source='X', keywords=['AI', '#MachineLearning'], limit=50) 3. Monitor Reddit subreddits: query_on_demand_data(source='REDDIT', keywords=['r/MachineLearning', 'technology'], limit=30) 4. Time-bounded search: query_on_demand_data(source='X', keywords=['Bitcoin'], start_date='2024-06-01T00:00:00Z', end_date='2024-06-03T23:59:59Z') Returns: Structured data with content, metadata, user info, timestamps, and platform-specific details.
Output schema not documented for any tool. Users/agents do not know what fields to expect in responses. The code returns JSON-serialized response dicts, but the description lists expected fields (status, data, meta) without formal schema.
No error handling logic visible in source code. If MC_API is not set, a warning is logged but the client is still initialized. If an API call fails, no error classification or recovery guidance is provided.
keyword_mode parameter in query_on_demand_data is described as 'any' or 'all' but NOT declared as an enum. LLMs may pass invalid values like 'AND', 'OR', or 'both'.
limit parameter lacks explicit bounds in schema. Description mentions range 1-1000, but no min/max constraints in JSON Schema. LLMs may pass 0, negative numbers, or millions.
Example values embedded in parameter descriptions (e.g., '@elonmusk', '@spacex', '#bitcoin', 'r/astronomy'). LLMs are known to reuse example values literally in actual calls, causing incorrect requests.
Pagination not supported. query_on_demand_data returns up to 1000 results with no offset/page parameter and no next_cursor. Large result sets blow context windows and degrade agent reasoning.
MC_API credential stored as environment variable at module level. While not a tool parameter, reliance on implicit env var auth is fragile and not documented in tool descriptions. Agents/users have no visibility into authentication state.