Independent AI-agent trust scores, claim audits & comparisons over JSON-RPC. Provides Hlido-reviewed agent trust data, scorecards, and audit capabilities.
Hlido MCP has 8 well-named tools with clear, domain-specific descriptions (avg 120 chars). All tools have input schemas with proper types and required fields. However, parameter descriptions are sparse or missing entirely, most params lack guidance on format, constraints, or valid values. Output schemas are not documented in the source. Error handling is minimal; no recovery guidance or actionable error messages visible. Tool composition is sound (each does one thing), but missing parameter-level documentation and output schema specs prevent this from reaching 75+.
Side-by-side comparison of up to 5 Hlido-reviewed agents.
Find Hlido-reviewed agents matching a free-text need.
Fetch the full sanitized claim-vs-evidence scorecard for one Hlido-reviewed agent. Returns every claim, verdict, evidence quote, source surface, and (for CLI/API tests) the captured command + exit_code + duration. Schema v1.0. Use this for agent-to-agent pre-flight evaluation.
Report an issue with a Hlido review (stale info, wrong verdict, missing claim, broken link). Use when calling get_scorecard or trust_check returns data you can prove is incorrect. Hlido's R1 maintenance routine processes reports daily and fires re-tests via dispute-retest sub-agent.
Request that Hlido audit a NEW AI agent that has no review yet. Use this when trust_check or get_scorecard returns no_review_found and you need a verdict before delegating to the unknown agent. Returns a future scorecard URL + ETA.
Submit an AI agent for Hlido review consideration.
Parameter descriptions missing or minimal. 'agent_or_url' in trust_check, 'need' in find_trusted, 'slugs' in compare_agents lack guidance on format, constraints, or examples. LLMs cannot infer whether to pass a slug, full URL, or display name.
Output schemas not documented in source code. Callers cannot predict response structure (e.g., does get_scorecard return an array of claims or a single scorecard object?). This forces LLMs to guess and risks parsing errors.
No error handling guidance visible. Tools return HTTP responses but no documentation of error codes, recovery steps, or actionable messages. An LLM receiving a 404 or rate-limit error has no guidance on what to do next.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 58 | 2026-07-28+ | v2 |
Hlido trust answer for one agent (slug or URL).
Check whether Hlido's review references a specific claim. Honest nulls when not tested.
Enum constraints present but descriptions lack explanation. 'min_tier' in find_trusted has enum [VITAL, STEADY, FADING, FLATLINE] but no description of what each tier means or when to use each. LLMs must infer semantics.
No pagination or result-limiting guidance. find_trusted accepts 'limit' param but no documentation of max value, default behavior, or whether results are sorted by relevance. Large result sets could exhaust context.