MCP server for Kalshi & Polymarket prediction markets — EV, Kelly, Bayes, cross-platform arbitrage, live edge signals and an NFL draft assistant, with a publicly graded track record.
Strong, domain-focused tool set with comprehensive descriptions and well-structured schemas. All 6 tools have clear names, detailed descriptions (mean 278 chars, well above the 194-char baseline), and complete input schemas with proper type constraints. Parameter naming is consistent and specific (e.g., marketPrice, yourProbability). Descriptions include outcome-first language ('returns the % edge and a BUY/SELL/SKIP signal') and appropriate use cases. Key strength: fraction-probability guard (isFractionLike check) prevents a common LLM error. Key weakness: output schemas are documented in code (via toolResult helper) but not formally exposed in the tool config objects, LLMs cannot see the response structure from the tool definition alone. No pagination/limits documented for multi-leg combo results. Error handling is reactive (guards against fraction input) but lacks comprehensive recovery guidance for edge cases.
Find the gap between a market price and a historical base rate, with a BUY / SELL / NEUTRAL signal. Includes 12 calibrated base rates (incumbent reelection, Fed decisions, recession timing, S&P returns, etc.) — useful for "is the market mispricing a known tendency", "historical vs. market", "base-rate arbitrage".
Update a prior probability with one or more pieces of evidence using Bayes theorem. Given a prior and a list of evidence items (each with P(evidence | true) and P(evidence | false)), returns the posterior probability and the per-step chain. Use for "update my estimate with new information", "posterior probability", "how does this news change the odds".
Calculate the expected-value edge on a Kalshi or Polymarket prediction-market contract. Given the current market price (in cents, i.e. the implied probability) and your own probability estimate, returns the % edge and a BUY / SELL / SKIP signal with a plain-English read. Use for "is this contract mispriced", "what is my edge", "should I take this position".
Calculate the fair value and EV edge of a multi-leg combo (parlay) on Kalshi or Polymarket. Given the individual leg prices and your joint-win estimate, returns the fair-value band, expected-value edge, and a verdict (SMASH / PLAY / LEAN / PASS / RUN). Optional: provide the offered combo price to see if you should take it. Use for "is this parlay worth it", "combo fair value", "multi-leg edge".
Output schemas are implemented in code (via toolResult helper, returning structured objects with fields like edge_pct, signal, interpretation) but NOT exposed in the tool config definitions that LLMs parse. LLMs cannot see what fields to expect from responses, forcing them to plan blind and parse outputs without formal guidance.
No explicit validation or error recovery guidance for edge cases: bayes_update allows evidence with probabilities that violate Bayesian assumptions (e.g., both likelihoodIfTrue and likelihoodIfFalse = 0); combo_edge lacks maximum bound on legPrices array (preventing runaway array sizes); convert_probability value parameter lacks per-format range hints in description.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 68 | 2026-07-28+ | v2 |
Convert between implied probability, American odds, and decimal odds. Give one value and its format and get all three back (American odds carry no commas, e.g. +441 or -200). Use for "what is +150 as a probability", "convert 62% to American odds", "decimal to implied odds".
Compute the optimal Kelly position size for a prediction-market contract. Given your win probability, the market price (which sets the payout), your bankroll, and a Kelly fraction (full / half / quarter / eighth), returns the dollar stake and a risk rating. Use for "how much should I stake", "what is my position size", "Kelly sizing for this trade".
base_rate_gap accepts an enum of 12 baseRateId values, but the description does not list or hint at the available options. LLMs must reverse-engineer the enum from the schema rather than reading a natural description like 'e.g. incumbent_reelected, fed_hold_unemp_below_4, ...'.
No pagination or result-limit guidance for combo_edge when calculating multi-leg combos. If an agent nests deeply (e.g., 10-leg parlay), the output could grow large without explicit capping.