Multi-tier orchestration system with MCP servers for email, LinkedIn, Odoo accounting, and social media integration. Bronze Tier orchestrator manages HITL approvals, Silver Tier provides MCP endpoints for approved actions, Gold Tier extends with Odoo/social media capabilities.
This suite demonstrates critically weak definition quality across all 22 tools. While descriptions are present for most tools, they lack the specificity and LLM-optimization required for reliable agent tool selection. More critically, NO TOOLS expose input schemas in the source code, only parameter lists are inferred from the tool list summary. This means schema validation, type enforcement, and parameter constraints are completely invisible to the scoring and likely absent from the implementations. The server appears to be a multi-tier orchestration system (BronzeTier/SilverTier/GoldTier/PlatinumTier) where MCP tool definitions are either incomplete, inferred, or hidden in server code that was not provided. The descriptions are functional but generic (e.g., 'Health check endpoint' for three different health tools), parameter names lack the verb_noun convention and type suffixes needed for clarity, and there is no evidence of output schema documentation, error recovery guidance, or composition patterns. The tool portfolio shows scattered responsibilities without clear tool chain design, multiple 'health' tools, overlapping social media posting patterns, and no pagination or limit documentation despite working with lists. This server would require substantial refactoring to meet production standards.
Approve and publish a LinkedIn post that was queued for HITL approval
Post to all configured social media platforms (Twitter, Facebook, Instagram) simultaneously
Create a draft invoice in Odoo with partner, amount, and description
Queue or immediately publish a LinkedIn post with optional approval requirement
Health check endpoint for LinkedIn MCP server
Health check endpoint - verifies Odoo connection and returns server status
Health check endpoint for social media MCP server
No visible input schemas in provided source code, schema validation and type enforcement cannot be verified. All 22 tools score 0 on schema dimension because JSON Schema definitions with types and property descriptions are not exposed in the evaluation materials.
Three distinct 'health' tools (odoo, social_media, email, linkedin) with identical naming and descriptions. Tool names do not distinguish between them, LLM will conflate these tools and struggle to select the correct one. Violates naming clarity principle: 'When multiple tools operate on the same resource, their names must make the distinction obvious.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Health check endpoint for email MCP server - verifies Gmail configuration
List invoices from Odoo accounting system with optional filtering by state
List journal entries from Odoo accounting system
List customers or vendors from Odoo with optional filtering by type
List LinkedIn posts pending approval
List published LinkedIn posts made by the AI Employee
List recently sent emails from the audit log
Post to a Facebook Page via Graph API with optional image
Create Instagram media post via Graph API with image and caption (Business/Creator account required)
Post a tweet via Twitter API v2 with optional image
Return a simple P&L summary from Odoo account balances including revenue, expenses, and net profit
Reject a LinkedIn post that was queued for HITL approval
Save email as a draft in Vault/Drafts/ without sending
Send an email via Gmail SMTP (called by Bronze Tier Orchestrator after HITL approval)
Send a test email to verify email server configuration is working
'broadcast' tool combines multiple responsibilities (post to Twitter, Facebook, Instagram simultaneously). Violates single-responsibility principle. Should be split into separate tools that agents can compose independently.
Descriptions lack LLM-optimized structure. Most fall short of baseline 194 chars average and lack actionable detail on WHEN to use a tool vs. similar alternatives. E.g., 'Health check endpoint' tells an LLM nothing about when to call health vs. proceeding with the actual operation.
Parameters documented informally in string descriptions rather than as formal schema properties. No evidence of type definitions, enums, min/max constraints, or format declarations. E.g., 'state' in list_invoices documented as string with examples (draft, posted, all) but no enum constraint visible. Prevents LLM from self-correcting invalid input.
No output schema documentation visible. LLMs cannot plan downstream tool calls (e.g., after listing invoices, what fields can be passed to other tools?). Absence of per-item success/failure on batch operations and no pagination documentation despite working with lists.
Parameter names lack type suffixes and clarity. 'partner_name' assumes the system accepts names, but Odoo often requires partner IDs. No evidence of supporting both IDs and human-readable identifiers. 'image_url' assumes URL string format; no validation or error guidance if format is invalid.
HITL approval pattern (require_approval boolean) is used across social tools but error handling and recovery guidance are not visible. No documentation on what happens if approval is denied, how to retry, or what to do if approval times out.
No error handling guidance visible. LLM has no instruction on what to do if list_invoices returns empty (try different state filter?), if create_invoice fails (validate partner first?), or if email send times out (retry with smaller batch?).
Tool composition broken. No documentation on which tools return IDs that downstream tools accept. E.g., does list_invoices return invoice_id suitable for use in other operations? Can create_invoice output be directly piped to another tool?
Defaults defined but not justified. E.g., limit=20 for list_invoices, limit=50 for list_partners, why different defaults? No documentation on whether these are safe, whether they hit API rate limits, or how agents should adjust for large datasets.