Autonomous FTE powered by Claude Code and Obsidian with multiple MCP servers for Gmail automation, LinkedIn posting, and cross-platform watchers
This server has 5 tools with mixed quality. All tools have descriptions and input schemas that are visible in the code. However, several critical issues significantly lower the score: (1) Two tools (send_email, draft_email) have identical schemas and nearly identical descriptions, which violates the single-responsibility principle and causes LLM confusion. (2) Parameter descriptions are minimal, most are 1-3 words without guidance on format, constraints, or expected values. (3) No output schemas are documented, so LLMs cannot predict what these tools return and plan subsequent calls. (4) Error handling is generic (returns status/message dict) with no recovery guidance. (5) send_email requires human approval per description but has no mechanism in the tool definition to enforce or document this. (6) reply_to_message_id parameter exists but lacks explanation of format (Message-ID vs Gmail message ID). (7) LinkedIn tools' visibility parameter should be an enum, not a free-form string. (8) No pagination, limits, or batching, even though the tools are write operations, missing docs on idempotency and retry behavior. Average tool score: 42 (calculated from per-tool assessments below).
Create a LinkedIn post about business updates to generate sales
Create an email draft in Gmail without sending it.
Mark a Gmail message as read by removing the UNREAD label.
Send an email via Gmail. Requires human approval per Company Handbook.
Share an article/link on LinkedIn
send_email and draft_email have identical input schemas and almost identical descriptions. This violates single-responsibility principle and causes LLM to conflate the tools. They should either be merged with a 'mode' enum parameter or the descriptions clarified to explain distinct use cases.
Parameter descriptions are extremely minimal (1-3 words). E.g., 'Recipient email address', 'Gmail message ID', 'The content of the post'. LLMs cannot infer format, constraints, or expected values from these sparse descriptions.
No output schemas documented for any tool. LLMs cannot predict what these tools return (status code, message text, IDs, etc.) and cannot plan downstream calls or extract the right fields.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 44 | <=2025-11-25 | v2 |
LinkedIn visibility parameter in both create_linkedin_post and share_linkedin_article is a free-form string with description 'PUBLIC or CONNECTIONS_ONLY'. Should be declared as an enum type in JSON Schema. Free-form strings invite hallucinated invalid values.
send_email description mentions 'Requires human approval per Company Handbook' but the tool definition has no mechanism to enforce or document this (no confirmation prompt, no approval_status field, no dry_run support visible in schema). If approval is required, either implement a confirm_before_execute pattern or document the approval workflow clearly.
reply_to_message_id parameter in send_email and draft_email lacks format clarity. The code sets 'In-Reply-To' and 'References' headers (RFC 2822 Message-ID format), but the description just says 'Message-ID header of the email being replied to'. Should clarify the exact format expected and how to obtain it (from get_email, Gmail API, etc.).
No error handling guidance in any tool. If send_email fails due to invalid recipient, authentication, or approval denial, the code returns {status: 'error', message: str(e)}, but the tool description doesn't tell LLMs what errors are possible or how to recover. Per pattern, error responses must guide the next action.
create_linkedin_post content parameter has no length limit or format guidance documented. LinkedIn posts have character limits and may have formatting constraints. LLMs need explicit bounds to avoid oversized submissions.
share_linkedin_article url parameter has no validation rule or format documented. Should specify valid URL format (http/https, no query params, max length, etc.) to prevent malformed submissions.
No idempotency semantics documented. send_email and draft_email are write operations that agents may retry on ambiguous failures. Are these operations idempotent? If an LLM retries, will it create duplicates? The description should clarify or the tool should provide a request_id/idempotency_key parameter.