Model Weights
Every input. Every weight. Nothing hidden.
Hitter Projection
v80 (2026-07-24)Plate Appearances
Away leadoff 4.52 PA down to home 9-hole 3.25 PA
+0.70 PA per implied run, clamp widened to 0.90-1.12 (v74)
Hit Rates
1.5x weight vs season, priored to league averages
w = splitPA/(splitPA + 2200 RHB / 1000 LHB), slope 0.4 (halved v76), clamp 0.90-1.10
Hit-Event Multipliers
Runs effect on all hits; crosswind XBH term on 2B/3B; HR uses total HR effect only
Hand-split savant rows
SP leg = xFIP + K9 (v74: -5.5% hit rate per K9 above 8.4, residual-fit)
14-day mean Cd vs 0.3386 baseline (v89 fit era), slope 10, clamp 0.93-1.08
Peak age 29, +0.6%/-0.3% per year, weighted by sample size so established hitters (whose current-season sample already reflects their age) are unchanged and only thin-sample and rookies move. Birthdates from MLB StatsAPI via mlb_player_bio.
Runs (implied-total distribution)
RBI protection: next-2-batters wOBA, 0.92x to 1.10x
Park, weather, and pitcher NOT re-applied here -- market total already prices them
Projected PA / H / HR / RBI / R / BB / SB assembled into DK and FD fpts
Per-player sigma fit from residuals (v68). Level factor read from mlb_level_recal_config and applied ONCE in-engine at write; raw fpts kept in base_proj_fpts_dk/fd. The old DB trigger compounded its factor and was dropped 2026-07-04. v80 (2026-07-24): component calibration recentered on current-version data (V63 refit on v78-plus rows, n=3075 hitters). Park run environment now nudges single, double and triple rates (mean-preserving; runs and RBI untouched so the Vegas implied total is not double counted). Stolen bases are no longer suppressed by the opposing pitcher contact-suppression factor.
Pitcher Projection
v80 (2026-07-24)Innings Pitched
-4.1% IP per implied run over 4.5, clamp 0.90-1.08, applied post-shrink; fatigue and IL-return caps also apply
Strikeouts
v75/v76: K scales with the opponent-implied and fatigue IP multipliers so proj_k / proj_ip stays a clean per-PA rate for the K-sim
2026-07-24 review: K inputs and weights confirmed sound. K-Elo is already applied here as a multiplier (v72_k_elo) and backtesting showed it is correctly sized, so no change was made. Open item under forward grading: engine K is mildly over-dispersed (actual-vs-proj slope about 0.73); a mean-preserving shrink is being tested in mlb_engine_k_caltest before any deploy. The Strikeout Sim betting model now simulates this engine proj_k directly, its old parallel rate model retired as double counting.
Earned Runs (v74 blend)
K-BB% beat all ERA estimators out of sample (R2 0.224); raw ERA cut from 30%
Park runs, weather, Statcast xwOBA-against, TTO taper, opponent implied (+5.8% ER per run, v74; band scale keyed to pre-adjustment IP, v75)
RETIRED v76: last-30 ERA already in the level blend; re-applying it was a same-signal-twice double count
Fantasy-Point Distribution (v77)
Pitcher DK and FD floor extends 20 percent further below the mean (PITCH_SKEW 1.20) to reflect early-knockout starts; ceiling and mean unchanged. Per-player sigma from residuals (v68).
Win Probability
WIN_IP_SCALE 0.94, recalibrated 2026-07-04 (was 0.67; proj 0.220 vs actual SP win rate 0.310)
Nightly Accuracy Loop
Feeds the adaptive level controller (mlb_level_recal_config, regime-dated window)
Projected IP / K / ER / Win assembled into DK and FD fpts
v74 (2026-07-04) regression-test batch: win scale 0.94, opponent implied into IP/ER, ER blend rebalanced with kwERA, platoon regression, drag recenter, level factor moved in-engine. v75/v76 (2026-07-05) double-count audit: ER band scale keyed to pre-adjustment IP, K consistent with BF, platoon slope halved, recent-form ER retired. Catcher battery nudge applies afterward in its own catcher_base_* columns. v-note 2026-07-24: K input review confirmed the weights sound and K-Elo correctly sized (no change); engine K over-dispersion (slope about 0.73) is under forward grading via mlb_engine_k_caltest before any change; the Strikeout Sim now simulates engine proj_k directly. v80 (2026-07-24): short-outing ER scaler smoothed from a five-step function to a continuous curve (mean-preserving over the IP distribution; removes the 12 percent cliff at 5.0 IP). Win probability recentered about 8 percent toward the actual SP win rate. Strikeout weights unchanged here, delegated to the mlb_engine_k_caltest forward-grading track.
Pure v2 Run Model run_v1.1 (shadow, no Vegas)
expected_runs and win_prob per team in mlb_team_run_model; anchors the Pure v2 (No Vegas) projection set
run_v1.1 (2026-07-20): the first two graduates of the stat evaluation harness are live in this model. Heart share, the share of a starter's pitches over the middle of the plate, raises the opposing team's expected runs: across roughly 780 graded forward starts, the heaviest-heart third of starters ran a 4.49 forward ERA against 3.51 for the stingiest third, and the relationship survives controlling for current ERA. Team defense (Outs Above Average) lowers the opposing team's expected runs and predicts run prevention even after controlling for defense-independent pitching stats. Both factors ship at half their fitted strength as a conservative first pass, cap out at a few percent of a team total, use league-neutral fallbacks when data is missing, and will be refit with the model's other slopes in the scheduled August recalibration. The live projections and optimizer are unchanged; this model powers the Pure v2 shadow stack only.
Data Sources
Statcast Savant · MLB Stats API · Open-Meteo · DraftKings
Refresh Cadence
Lineups 30 min · Odds hourly · Savant daily
Stack
Supabase Postgres + pg_cron · Edge Functions (Deno) · Base44 UI
Audit Trail
Daily snapshots locked at noon CT · top-20 actuals join · 30-day rolling hit rate