The FRED score answers one question for every active motor carrier in the United States: relative to peers of the same size, how much crash harm is this carrier likely to generate over the coming year? The answer is built entirely from federal FMCSA records — fleet size, operating authority, safety rating, roadside inspections, violations, and crashes — and is refreshed weekly across every active US motor carrier (~1.16 million scored; ~300,000 with enough observed record to earn a letter grade, the rest shown Not Rated).
This paper specifies the engine end to end. We define the quantity being predicted: a carrier's forward twelve-month severity-weighted crash burden, in which a fatal crash counts many times a property-damage one. We show how each of six safety signals — crashes and five violation families — is converted into a per-exposure rate, crash burden per power-unit-year and each violation per inspection (vehicle out-of-service per vehicle inspection), rather than through a self-reported mileage model; how a thin carrier history is stabilized against its fleet-size peers with a per-signal Empirical-Bayes (Bühlmann–Straub) credibility estimator, so a sparse record shrinks toward its cohort average while a rich one rides its own experience; how the six credibility-adjusted relativities are combined into a single peer relativity \(R_c\) as a weighted geometric mean, protected by an observed-crash floor that keeps a real crash record from being washed out by clean or absent violation signals; and how \(R_c\) maps through absolute cutpoints — recalibrated to the realized forward-burden gradient, not percentiles — to a peer-relative letter grade (Excellent … Critical), a 0–100 score, and an explicit fatal-crash probability. Every constant and formula below is taken directly from the production engine.
Underwriters do not ultimately care how many violations a truck has — they care about the cost of future crashes. FRED therefore predicts a forward-looking, harm-weighted quantity rather than a backward-looking tally. For carrier \(i\), define the target
$$ B_i \;=\; \text{total severity-weighted crash burden generated by carrier } i \text{ over the next 12 months.} $$A single crash contributes a weight that reflects its severity (§4): a fatality weighs far more than a fender-bender. \(B_i\) is a non-negative quantity that is exactly zero for the large majority of carriers (most have no crash in a given year) and continuously positive for the rest. That two-part structure — a spike of probability mass at zero plus a continuous positive tail — is the defining feature of the data and dictates the model family we use (§7).
Because a 3-truck owner-operator and a 5,000-truck fleet are not comparable on raw counts, the engine works in rates per mile driven and then expresses every carrier relative to its own size peers. The published grade answers “how does this carrier compare to fleets like it,” while the underlying expected burden and crash count answer “how much, in absolute terms.”
All inputs are federal FMCSA datasets, refreshed weekly: the census (fleet size, mileage, domicile, authority, safety rating), roadside inspections, violations, and crashes. The engine scores every carrier with a positive reported power-unit count.
The model is trained on a rolling year-over-year window. Let \(C\) be the most recent crash-mature date (we discard the last 45 days, since crash reports arrive with a lag). We form two consecutive half-open windows of exactly 365 days each:
$$ \underbrace{[\,C-730,\ C-365\,)}_{\textstyle \text{feature window } W_{\text{fit}}} \;\longrightarrow\; \underbrace{[\,C-365,\ C\,)}_{\textstyle \text{outcome window } W_{\text{out}}}. $$Each window includes its left endpoint and excludes its right, so the two windows partition the 730 days before \(C\): no day is double-counted and none is dropped.
The model learns the map from a carrier's safety profile in \(W_{\text{fit}}\) to its realized burden in the following year \(W_{\text{out}}\). To score live carriers we then feed each carrier's most recent twelve months of features (the \(W_{\text{out}}\) window) into the fitted model to predict the forward twelve months. This out-of-time design is what makes the score genuinely predictive rather than a restatement of past events.
| Symbol | Meaning |
|---|---|
| \(E_i\) | Crash exposure of carrier \(i\), in power-unit-years over a twelve-month window (§3) |
| \(E_{i,s}\) | Exposure for violation signal \(s\): the carrier's inspection count in the window — vehicle inspections for the vehicle-OOS signal (§3) |
| \(N_i\) | Crash count in a twelve-month window |
| \(B_i\) | Severity-weighted crash burden in a twelve-month window (§4) |
| \(b(i)\) | Fleet-size cohort of carrier \(i\): one of seven power-unit bands (§3) |
| \(r_{i,s}\) | Credibility-adjusted peer relativity of carrier \(i\) on signal \(s\) (§6); one of six safety signals — crash, vehicle-OOS, severe violation, HOS/fatigue, unsafe driving, speeding |
| \(R_{c,i}\) | Composite peer relativity of carrier \(i\): weighted geometric mean of the six \(r_{i,s}\), floored by observed crashes |
Every relativity in this engine is a rate, and every rate needs an exposure base — the denominator that scales a raw event count into something comparable across a one-truck owner-operator and a thousand-truck carrier. FredScore deliberately does not use self-reported annual mileage for that denominator, and the reason is empirical: the MCS-150 mileage field is too often blank, stale, or implausible to be the load-bearing input to a score that decides insurability. A carrier can under-report to look safe, file once and never update, or fat-finger a figure by an order of magnitude, and there is no roadside record to check it against.
FredScore therefore measures exposure with quantities FMCSA actually observes and timestamps. Crash burden is rated per power-unit-year, and each violation signal is rated per inspection. For a carrier \(i\) scored over a trailing one-year window, the crash exposure is simply its fleet size,
$$ E_i \;=\; u_i \quad\text{(power units over the window)}, $$and the violation exposures are inspection counts: vehicle out-of-service violations are rated per vehicle inspection, while severe, hours-of-service, unsafe-driving, and speeding violations are rated per inspection. Each signal thus carries its own natural denominator (§6), and none of them is mileage.
Self-reported mileage is not discarded — it is demoted. Reported miles no longer set exposure, but they still do two jobs: they disqualify incoherent records (a fleet whose implied miles-per-truck is impossible is left ungraded rather than rated on a fiction), and a stale MCS-150 filing is surfaced as a data-quality flag alongside the grade. Mileage informs trust in the record; it does not compute the score.
Two data-quality guards protect the book from corrupt federal records. A carrier whose reported fleet exceeds 50,000 power units (the real-world maximum is ~20,000), whose census is internally incoherent (far more power units than drivers), or whose reported mileage is implausible for its fleet size is treated as having no trustworthy exposure and is left ungraded (Not Rated) rather than handed a fabricated baseline.
Exposure is only meaningful against peers of like size, so every carrier is assigned to one of nine fleet-size cohorts. All credibility, peer means, and priors below are estimated within a cohort — the cohort is the unit of comparison everywhere downstream. FredScore is built for the ten-plus power-unit underwriting market, so the six in-scope cohorts band that range finely; every one carries hundreds to thousands of carriers with an observed crash, so each cohort's prior is well estimated. The three cohorts under ten power units are advisory: FMCSA data cannot finely separate crash risk in that thin-record population (§9, and the owner-operator study), so they are searchable and shown but held out of the default ranked book.
| Cohort | Power units \(u\) | Scope | Typical carrier |
|---|---|---|---|
pu1 | \(u = 1\) | Advisory | Owner-operator |
pu2_5 | \(2 \le u \le 5\) | Advisory | Micro-fleet |
pu6_9 | \(6 \le u \le 9\) | Advisory | Very small fleet |
pu10_19 | \(10 \le u \le 19\) | In-scope | Small fleet |
pu20_49 | \(20 \le u \le 49\) | In-scope | Small regional |
pu50_99 | \(50 \le u \le 99\) | In-scope | Mid-size fleet |
pu100_199 | \(100 \le u \le 199\) | In-scope | Large regional |
pu200_499 | \(200 \le u \le 499\) | In-scope | Large carrier |
pu500p | \(u \ge 500\) | In-scope | Fleet / mega-carrier |
Banding the ten-plus book this finely — rather than lumping, say, a 20-truck regional fleet with a 500-truck carrier — lets the engine rank each carrier against genuinely comparable peers, which is where the discriminating power lives (§13). The band edges are set where the population supports a stable per-cohort prior.
Not all crashes are equal. The burden of a single crash \(c\) is a weight that rises with its consequences — counting casualties, not merely flagging them:
$$ w_c \;=\; 1 \;+\; 12\cdot\min(\text{fatalities}_c,\,3) \;+\; 4\cdot\min(\text{injuries}_c,\,5) \;+\; 3\cdot\mathbb{1}[\text{hazmat released}]. $$A baseline tow-away crash weighs 1; a one-injury crash weighs 5; a one-fatality crash weighs at least 13 — and a crash that kills two people weighs more than one that kills one. Fatality and injury terms stack: a fatal crash with two additional injuries weighs \(1+12+8=21\). The caps (three fatalities, five injuries) keep a single catastrophic record from dominating a carrier's burden. A carrier's burden over a window is the sum over its crashes,
$$ B_i \;=\; \sum_{c\,\in\,\text{crashes}(i)} w_c, $$and the crash count is simply \(N_i = |\text{crashes}(i)|\). These weights are a deliberately compressed harm index, chosen for statistical stability rather than cost proportionality: DOT's value of a statistical life implies a fatal-to-property-damage cost ratio in the hundreds, and a cost-proportional target would be dominated by rare fatal events and statistically unstable. Tail risk is instead carried by the dedicated fatal-crash model of §11, and Appendix §B shows that carrier rankings are nearly invariant to the choice of fatal weight (8 / 12 / 20). The burden \(B_i\) is what the crash relativity of §5 rates per power-unit-year; the count \(N_i\) drives the frequency projection of §8.
A 12-truck carrier had three crashes last year: a crash with two fatalities and one injury, a hazmat-release crash, and a tow-away. Its burden is
$$ B \;=\; \underbrace{(1+ 2\cdot 12 + 1\cdot 4)}_{\text{2 fatalities + 1 injury}} \;+\; \underbrace{(1+3)}_{\text{hazmat}} \;+\; \underbrace{1}_{\text{tow}} \;=\; 34, \qquad N = 3. $$The fatal and injury contributions of the first crash stack (\(1+24+4=29\)); under indicator weights it would have counted the same as a single-fatality crash. A peer with three tow-aways has the same count but only \(B=3\) — the engine treats the first carrier as by far the larger risk.
A carrier's own history is the strongest signal we have — but for a small fleet it is also the noisiest. One crash for a 3-truck operator could be terrible luck or a genuine pattern; the raw rate alone cannot tell us which. The classical actuarial answer is credibility: blend the carrier's own experience with the experience of its peers, weighting the carrier more as its data accumulates. We implement this with an Empirical-Bayes Gamma–Poisson (Bühlmann–Straub) model, fit separately within each size cohort, and — crucially — applied once per signal. There is no downstream model that re-shrinks the result; the credibility blend below is the estimate.
Consider a single signal \(s\) — say crash burden, rated per power-unit-year, or vehicle out-of-service violations, rated per vehicle inspection. Within a cohort, suppose each carrier's latent rate \(\lambda\) is drawn from a Gamma prior, and events are Poisson given the rate and the carrier's exposure \(E\) on that signal:
$$ \lambda \sim \text{Gamma}(\alpha_g,\beta_g), \qquad V \mid \lambda,E \sim \text{Poisson}(\lambda E), $$where \(V\) is the observed count (or burden) of signal \(s\) and \(g\) indexes the cohort. The prior is estimated from the cohort's pooled data with the standard Bühlmann–Straub estimators for Poisson exposure data. Writing \(r_k = V_k/E_k\) for carrier rates and \(E_\bullet = \sum_k E_k\) for total cohort exposure, the cohort mean is the exposure-weighted rate — total events over total exposure, not a mean of carrier rates:
$$ \mu_g \;=\; \frac{\sum_k V_k}{\sum_k E_k}. $$A subtlety matters here. The raw variance of observed rates, \(\sum_k \pi_k (r_k-\mu_g)^2\), is not the between-carrier variance: conditional on a carrier's true rate \(\lambda\), the observed rate still fluctuates with Poisson sampling noise of variance \(\lambda/E_k\), so the raw quantity estimates \(\operatorname{Var}(\lambda) + \mathbb{E}[\lambda/E_k]\). Using it directly would overstate the prior variance, understate \(\beta_g\), and hand thin-data carriers too much credibility. Bühlmann–Straub removes the sampling noise explicitly: for Poisson data the process variance is \(\hat s^2 = \mu_g\), and the variance of hypothetical means (the true between-carrier variance) is
$$ \hat a_g \;=\; \frac{\sum_k E_k\,(r_k-\mu_g)^2 \;-\; (n_g-1)\,\hat s^2}{E_\bullet \;-\; \sum_k E_k^2 / E_\bullet}, $$where \(n_g\) is the number of carriers in the cohort. If the numerator is non-positive the cohort shows no detectable between-carrier heterogeneity, and \(\hat a_g\) is floored at a small \(\varepsilon\) so that credibility simply falls to zero rather than dividing by zero. The Gamma prior parameters follow:
$$ \beta_g = \frac{\mu_g}{\hat a_g}, \qquad \alpha_g = \mu_g\,\beta_g, $$and this \(\beta_g\) is the Bühlmann credibility constant \(K\) of classical credibility theory. A cohort with fewer than 50 gradeable carriers cannot support a stable prior; there the estimator falls back to near-zero credibility, so every carrier in it is shown essentially at its cohort average rather than on noise.
Gamma–Poisson conjugacy then gives the posterior mean rate for a carrier with \(V\) events over exposure \(E\) in closed form, and we report it as a relativity against the cohort mean \(\mu_g\):
$$ \boxed{\;r_s \;=\; \operatorname{clip}\!\Big(\frac{1}{\mu_g}\cdot\frac{\alpha_g + V}{\beta_g + E},\ 0.01,\ 100\Big)\;} \qquad(r_s = 1 \text{ means “exactly typical for your size.”}) $$The clip to \([0.01, 100]\) is numerical safety only. This is exactly a credibility blend. Rewriting with the per-signal credibility weight \(Z = E/(E+\beta_g)\),
$$ \frac{\alpha_g+V}{\beta_g+E} \;=\; Z\cdot\underbrace{\frac{V}{E}}_{\text{own rate}} \;+\; (1-Z)\cdot\underbrace{\mu_g}_{\text{cohort mean}}, $$so a carrier with little exposure is pulled toward its peers (\(Z\to 0\), \(r_s\to 1\)), while a carrier with a long track record stands on its own experience (\(Z\to 1\)). The prior dispersion \(\beta_g\) sets the “speed” of that transition and is learned from the cohort, not assumed. Because the shrinkage is applied here, once, each \(r_s\) is already the carrier's best credibility-weighted estimate for that signal — the composite of §7 simply combines the six of them.
The engine reads a carrier through six safety signals, each a per-signal Empirical-Bayes relativity from §5, each with its own exposure denominator. These six relativities are the entire read of the carrier's behavior — there is nothing else feeding the grade — and they combine transparently in §7.
| Signal | Counts | Exposure base | Weight \(w_s\) |
|---|---|---|---|
| Crash burden | severity-weighted burden \(B\) (§4) | power-unit-years | 0.200 |
| Vehicle out-of-service | vehicle OOS violations | vehicle inspections | 0.189 |
| Severe violations | high-severity violations | inspections | 0.105 |
| HOS / fatigue | hours-of-service violations | inspections | 0.088 |
| Unsafe driving | unsafe-driving violations | inspections | 0.076 |
| Speeding | speeding violations | inspections | 0.037 |
The weights are log-relativity exponents learned out-of-time by a non-negative Poisson GLM regressing next year's crash burden on \(\sum_s w_s \log r_s\). They are reported raw — not renormalized to sum to one — because their sum is the calibrated spread that the grade cutpoints of §10 are tuned against. Crash burden is the natural lead of a crash-safety score, so its weight is floored to lead the field (\(0.200 \ge 0.189\)) for face validity.
Why only six? Coarse “behavioral” and “equipment” violation buckets were tested as candidate signals. Fit honestly by a non-negative Poisson GLM, both scored zero or negative — they are confounded by inspection intensity (a heavily inspected carrier accrues more violations of every kind without being riskier) and add no predictive signal once the specific, mechanism-level signals above are present. They are therefore excluded: every signal in the engine earns its place empirically.
Note that crashes are rated per power-unit-year while the five violation signals are rated per inspection (vehicle OOS per vehicle inspection). This is deliberate: a violation is a property of an inspection event, so its exposure is the number of inspections the carrier underwent, whereas a crash is a property of operating a truck, so its exposure is the size of the operating fleet. Federal risk scores that themselves depend on crashes (FMCSA BASIC percentiles, prior FRED outputs) are never read as inputs — admitting them would leak the outcome into the signals.
The six relativities of §6 combine into one number: the carrier's experience relativity \(R_c\), where \(R_c=1\) means “exactly typical for your cohort,” below 1 is safer than peers, and above 1 is riskier. This is a transparent, weighted geometric mean of the signal relativities — the log-linear form the GLM weights were fit for — not a black-box prediction:
$$ R_{\text{own}} \;=\; \exp\!\Big(\sum_s w_s \ln r_s\Big) \;=\; \prod_s r_s^{\,w_s}. $$Every input to \(R_{\text{own}}\) is a credibility-shrunk relativity, so a thin record already sits near 1 (cohort average) on every signal and cannot swing \(R_{\text{own}}\) far from typical. That is the correct behavior for the absence of events — but it is exactly the wrong behavior for the presence of a crash. If a carrier's clean violation profile could pull \(R_{\text{own}}\) below 1 while a real, fatal crash sat in its record, the engine would wash out the single most important thing FMCSA observed about it.
So the engine adds an observed-crash floor. Independent of shrinkage, compute the carrier's raw severity-weighted burden relativity against its cohort mean burden rate \(\mu_{B,g}\),
$$ R_{\text{obs}} \;=\; \frac{B}{E_{\text{PU-yr}}\;\mu_{B,g}}, $$and discount it by a variance-aware factor \(\lambda\) that trusts the observation in proportion to how much it moves the estimate. With \(w^2 = \sum_c w_c^2\) the sum of squared crash weights and \(z=1\),
$$ \lambda \;=\; \operatorname{clip}\!\Big(1 - z\,\frac{\sqrt{w^2}}{B},\ 0,\ 1\Big), \qquad R_{\text{floor}} \;=\; \lambda\,R_{\text{obs}}. $$The ratio \(\sqrt{w^2}/B\) is the relative standard error of the burden. For \(N\) equal-severity crashes it reduces to \(1-1/\sqrt{N}\) — the intuitive “one crash proves little, five prove a pattern” rule — but a single high-severity crash has \(\sqrt{w^2}\approx B\), so \(\lambda\approx 0\) and the floor stays inert: one catastrophic record cannot manufacture a false Critical off a tiny sample. The final experience relativity takes the worse (higher) of the shrunk composite and the discounted observed floor:
$$ \boxed{\;R_c \;=\; \max\!\big(R_{\text{own}},\ R_{\text{floor}}\big)\;} $$Because the floor is a max, it can only ever raise \(R_c\) — a real crash record can never be improved away by clean paperwork, but a clean carrier is never penalized by a floor it hasn't earned.
A 20-truck carrier has a genuinely clean violation record: every one of its six signal relativities lands near or below 1, so its shrunk composite is \(R_{\text{own}} = 0.71\) — that alone would grade Satisfactory. But it also had two injury crashes last year (\(B = 10\)) against a cohort mean burden rate that implies a typical fleet its size would carry far less, giving \(R_{\text{obs}} = 1.45\). With \(N=2\) equal-severity crashes the discount is \(\lambda \approx 1 - 1/\sqrt{2} = 0.29\), so \(R_{\text{floor}} = 0.29 \times 1.45 = 0.42\) — the two crashes are real but the sample is thin, so the floor lands below the composite and does not bind. Now suppose instead the same carrier had a single two-fatality crash (\(B = 25\), \(R_{\text{obs}} = 3.6\)): here \(\sqrt{w^2}\approx B\), so \(\lambda \approx 0\) and the severe single record is also held in check — the engine refuses to brand a fleet Critical on one event. The floor bites hardest in the middle: several moderate crashes, where the sample is large enough to trust and \(\lambda\) approaches 1, lifting \(R_c\) firmly above the clean composite.
The grade answers “how does this carrier compare to its peers on severity-weighted risk.” Two further questions are useful to an underwriter: how many crashes to expect next year, and how likely a fatal one is. Severity belongs in the grade; it does not belong in a count projection — a carrier whose burden is high because of one catastrophic crash should not be projected to crash many more times. So the frequency projections use a separate frequency relativity built from the crash count relativity \(r_{\text{count}}\) (the EB relativity of §5 on \(N\), not \(B\)) blended geometrically with the five violation signals at the same weights:
$$ R_{\text{freq}} \;=\; \exp\!\Big(w_{\text{crash}}\ln r_{\text{count}} \;+\!\!\sum_{s\,\ne\,\text{crash}}\!\! w_s \ln r_s\Big). $$The expected forward crash count is then the cohort mean count rate \(\mu_{c,g}\) scaled by fleet size and this frequency relativity, and the expected fatal count follows the same shape with the cohort fatal rate:
$$ \hat N \;=\; \mu_{c,g}\,u\,R_{\text{freq}}, \qquad \hat N_{\text{fatal}} \;=\; \mu_{f,g}\,u\,R_{\text{freq}}. $$Fatal crashes are rare enough that a count is less useful than a probability. Treating fatal crashes as Poisson, the probability of at least one fatal crash in the forward year is
$$ \Pr(\text{fatal crash}) \;=\; 1 - e^{-\hat N_{\text{fatal}}}. $$These are display and pricing quantities, shown alongside the grade; they do not feed back into \(R_c\). The full tail model — how the fatal projection is calibrated on its own — is the subject of §11.
\(R_c\) is dimensionless and centered so that the exposure-weighted cohort average is 1.0. That anchors the scale
but does not, by itself, tell us where to draw grade lines. Those lines are set empirically, against the one thing
that matters: realized forward crash burden. The engine is fit on a trailing window and judged on the
disjoint forward year (§13); a diagnostic (scripts/calibrate_grade_cuts.py) then measures the
average realized forward burden per power-unit-year across fine buckets of \(R_c\), and the grade cutpoints of §10
are placed so that each grade band maps to a materially different realized burden — a calibration to outcomes, not
to percentiles.
The distribution of \(R_c\) is strongly right-skewed. The exposure-weighted average is 1.0 by construction, but the median carrier sits near \(R_c \approx 0.63\): most carriers are at or below their cohort average, and a long tail of high-\(R_c\) carriers pulls the mean up. This is why a naïve percentile split fails — placing the Strong/Satisfactory line at the median would drop it at \(0.63\), splitting the dense central mass into a bloated Strong and Satisfactory double-peak while leaving the genuinely risky tail under-resolved. The absolute cutpoints instead track the burden gradient: Strong/Satisfactory sits below the median so Strong contains only genuinely-below-average carriers, and the Marginal line lands near cohort average (\(R_c = 0.90\), where realized forward burden is roughly 1.0×) so the risky tail underwriting cares about is fully populated.
Where the calibration comes from — and where it doesn't. The gradient, and all the grade discrimination (§13), lives in the in-scope ≥6-PU book, and firms up as fleets grow. For the thin-record 1–9-PU population the same diagnostic found realized forward burden essentially flat across \(R_c\) (roughly 0.7×–1.0× of average everywhere in the dense middle): these carriers are statistically indistinguishable on forward crash burden using FMCSA data. Forcing a fine ranking on them would grade noise. The honest treatment (§10) is to shrink their thin records toward the cohort average, mark them Not Rated where there is no observed record at all, cap single-power-unit carriers below the top grade, and keep the advisory cohorts out of the default ranked book while leaving every carrier fully searchable.
The grade is a direct, deterministic read of \(R_c\) against the absolute cutpoints calibrated in §9. There is no percentiling and no within-cohort ranking — a carrier's grade depends only on its own \(R_c\), so two carriers with the same experience relativity in different cohorts receive the same grade.
| Grade | Experience relativity \(R_c\) | Score | Realized fwd burden / PU |
|---|---|---|---|
| Excellent | \(R_c < 0.52\) | \(\ge 90\) | 0.038 |
| Strong | \(0.52 \le R_c < 0.60\) | 75–90 | 0.054 |
| Satisfactory | \(0.60 \le R_c < 0.90\) | 50–75 | 0.068 |
| Marginal | \(0.90 \le R_c < 1.20\) | 30–50 | 0.099 |
| Poor | \(1.20 \le R_c < 1.60\) | 10–30 | 0.119 |
| Critical | \(R_c \ge 1.60\) | < 10 | 0.176 |
The realized forward burden per power-unit-year rises monotonically across the grades — roughly a 5× spread from Excellent to Critical (0.038 → 0.176) — and that monotonicity is enforced, not merely observed: the out-of-time gate of §13 aborts the score write if any better grade turns out to carry statistically significantly worse realized burden than a worse grade.
The 0–100 score is a monotone interpolation of \(R_c\) with knots aligned exactly to the cutpoints, so the score and the grade never disagree:
$$ R_c: \;[\,0,\ 0.52,\ 0.60,\ 0.90,\ 1.20,\ 1.60,\ 10\,] \;\longmapsto\; \text{score}: \;[\,100,\ 90,\ 75,\ 50,\ 30,\ 10,\ 0\,]. $$Thus Excellent scores ≥90, Strong 75–90, Satisfactory 50–75, Marginal 30–50, Poor 10–30, and Critical below 10.
Confidence is a separate axis from the grade. How much a carrier is graded on its own record versus its cohort average is governed by its exposure credibility \(Z_{\text{exp}} = E_{\text{PU-yr}}/(E_{\text{PU-yr}}+K_{B,g})\), where \(K_{B,g}\) is the compound-Poisson burden Bühlmann constant of the cohort. This drives a displayed confidence tier:
| Tier | Exposure credibility \(Z_{\text{exp}}\) | Reading |
|---|---|---|
| High | \(Z_{\text{exp}} \ge 0.50\) | Graded on the carrier's own record |
| Moderate | \(0.20 \le Z_{\text{exp}} < 0.50\) | Mostly own record, partly cohort |
| Low | \(0.05 \le Z_{\text{exp}} < 0.20\) | Provisional — leans on cohort average |
| Prior-only | \(Z_{\text{exp}} < 0.05\) | Essentially the cohort average |
Confidence then imposes a one-sided upside cap. Only the top grade is gated: a carrier below Moderate confidence (\(Z_{\text{exp}} < 0.20\)) is capped at Strong, so a clean-but-thinly-observed record cannot display Excellent on the mere absence of events. The downside is never capped — a real crash record can still grade a thin carrier Poor or Critical through the observed-crash floor. In the same spirit, single-power-unit carriers are capped at Strong: one power-unit-year is too little crash exposure to credibly prove the top grade, though their downside shows through in full. When a grade is capped down, the 0–100 score is pulled inside the capped band so it never shows out-of-band.
Two further rules sit atop the grade. A carrier carrying a current FMCSA Unsatisfactory or Conditional safety rating is flagged for referral and its displayed grade is capped at Marginal for binding (a superseded or historical rating does not penalize). And a carrier with no observed record — zero inspections and zero crashes in the window — is left Not Rated rather than graded: “no bad events” is not evidence of safety when there is no footprint at all. Carriers with no active operating authority, no reported power units, or untrustworthy exposure (§3) are likewise Not Rated. Referral facts (no authority, adverse rating, chameleon suspicion, stale MCS-150, zero-inspection large fleet, power-unit-versus-driver mismatch) are surfaced separately from the grade, never silently folded into it.
A brand-new 3-truck carrier has two clean inspections and no crashes. Its shrunk composite is \(R_{\text{own}} = 0.44\) — nominally Excellent territory — but with only two inspections its exposure credibility is \(Z_{\text{exp}} = 0.11\) (Low). The upside cap fires: the grade is held at Strong, and the score is pulled to just under 90. The engine is saying, correctly, “this looks good, but there isn't yet enough record to certify it as the best.” Log a year of clean inspections and \(Z_{\text{exp}}\) climbs past 0.20, the cap lifts, and Excellent becomes reachable on the carrier's own earned record.
Fatal crashes are rare, so the severity weighting of §4 (which caps fatalities at three) does not by itself characterize a carrier's tail risk. The engine therefore reports an explicit probability of at least one fatal crash in the coming year, built from the same peer-relative machinery — there is no separate feature model. Let \(\bar f_{g}\) be the exposure-weighted fatal-crash rate of cohort \(g\) (fatal crashes per power-unit-year among rated peers) and \(R^{\text{freq}}_i\) the carrier's frequency relativity from §8. The expected number of fatal crashes over the next year is
$$ \lambda_i \;=\; \bar f_{g(i)}\,\cdot\, u_i \,\cdot\, R^{\text{freq}}_i, $$the cohort's typical fatal rate scaled by fleet size and by how the carrier's overall crash-and-violation profile compares to its peers. Treating fatal events as Poisson, the probability of one or more in the year is
$$ \Pr(\text{≥ 1 fatal crash})_i \;=\; 1 - e^{-\lambda_i}. $$This converts an abstract intensity into a number an underwriter can act on, and is reported alongside the typical-fleet fatal baseline \(\bar f_{g(i)}\,u_i\) for context. Because it rides the frequency relativity rather than the severity-weighted burden, a carrier with a single catastrophic multi-fatality crash is not projected to suffer many more fatal crashes — that event is already reflected in its burden and grade.
A small set of deterministic rules sits on top of the statistical model, encoding facts that should override any learned prediction. They are applied after the §10 percentiling (see “Ordering matters” there):
| Condition | Action |
|---|---|
| Current FMCSA Unsatisfactory safety rating | Grade forced to Critical, score 0 |
| Historical Unsatisfactory/Conditional (stale, or superseded by active authority) | No penalty — surfaced as context only |
| No active operating authority, or no reported power units | Left ungraded (N/A) |
| Implausible fleet size / corrupt exposure (§3) | Left ungraded (N/A) |
| Provisional (thin) exposure | Grade capped at Satisfactory |
These rules are intentionally few. The philosophy is that the model should do the work; overrides exist only for
regulatory facts (an adverse federal rating) and data-integrity protection (un-scoreable records), never as
hand-tuned score nudges. The adverse-rating override fires only for a current rating: an FMCSA
Unsatisfactory cannot coexist with active for-hire operating authority (the authority would be revoked), so an
Unsatisfactory sitting next to an active docket — or one long past its review date — is treated as historical
and does not penalize the score. Currentness is derived from SAFETY_RATING, REVIEW_DATE,
docket status, and revocation history.
The engine's discriminating power is not uniform across fleet size, and honesty requires saying so. A calibration
diagnostic (scripts/calibrate_grade_cuts.py) measured realized forward crash burden across fine
relativity buckets within each size cohort. For the 1–9 power-unit population the result is stark: the
next-year burden per power-unit-year is essentially flat across the whole range of \(R_c\) —
$$
\mathbb{E}\!\left[\,\tfrac{B^{\text{fwd}}}{E}\;\middle|\;R_c=r,\ u\le 9\,\right]
\;\approx\; 0.72\text{--}0.99 \times \bar{B}/\bar{E}
\qquad\text{for all }r\text{ in the dense middle,}
$$
so a 1–9-unit carrier the engine would rank “safe” crashes forward at very nearly the same rate as one it would
rank “risky.” On this population \(R_c\) does not separate future outcomes: the carriers are, to the resolution FMCSA
data affords, statistically indistinguishable. The real grade gradient — the monotone rise in forward burden
across grades, and the per-cohort Gini reported in §13 — lives in the \(u\ge 10\) book, where each carrier accrues
enough power-unit-years for its own experience to carry credibility.
Two consequences follow. First, small fleets are generally omitted from the default ranking: the census overview defaults to a “Fleets only (10+ power units)” view, and portfolio rankings are built from the \(u\ge 10\) population where an ordering is meaningful. Second, small fleets remain fully searchable. Any USDOT carrier can be looked up directly, and its available FMCSA record — authority, inspections, violations, crashes, safety rating — is shown alongside an honest grade-or-Not-Rated status. This matters most for a broker vetting a single owner-operator for a brokered load: that broker needs whatever information exists on the hauler in front of them, even when the population as a whole cannot be finely ranked.
Refusing to finely rank these carriers is a deliberate act of calibration, not an omission. Because FMCSA data cannot separate 1–9-unit carriers on forward burden, forcing a fine ordering would grade noise — it would manufacture the appearance of precision the data does not support. The engine instead does the honest thing: a thin record is shrunk toward its cohort average (§5), \(R_c \to 1\) as credibility \(Z\to 0\); a record with no observed inspections and no crashes at all is marked Not Rated rather than assigned a fabricated grade; and a single-power-unit carrier is upside-capped at Strong (§10), since one power-unit-year of crash exposure cannot credibly earn the top grade on the mere absence of events. The downside is never capped — a real crash record still shows through to Poor or Critical on any fleet, however small. The full empirical treatment of the owner-operator case is set out in the grading owner-operators study.
Surface the data, do not fabricate precision. Where exposure is too thin to distinguish carriers on forward crash burden, the engine shows every carrier's FMCSA record and an honest grade-or-Not-Rated verdict, but it does not invent a fine ranking the underlying data cannot support. Discrimination is claimed only where it has been validated — in the \(10+\) power-unit book.
An actuarial grade is only worth what it predicts. Because the engine grades a carrier on its credibility-weighted own experience, the chief risk is not variance but circularity: a grade that merely re-describes the same window it was built from proves nothing. The validation design therefore forces a strict separation in time. The engine is fitted on the observation window \([C-730,\,C-365)\) — priors, cohort means, and credibility constants come only from those twelve months — and then judged on the disjoint forward window \([C-365,\,C)\), which no fitting step ever saw. Every carrier in the fit population is pushed through the identical production engine, and its realized forward burden per power-unit-year is tabulated by displayed grade. Discrimination that survives this gap is genuinely predictive, not descriptive.
MONO_SE_Z = 2.0
combined standard errors — a statistically significant inversion, not a small-sample tie. When the gate fires it
signals a real ordering defect to diagnose, and the database write is refused before a single carrier's grade
changes. A companion Gini-floor check (forward discrimination must not regress versus the prior run) is logged
by default and made aborting under --strict-gates.
On the July 2026 refresh the ordering is clean. The table below is the payload of that gate: realized forward burden per power-unit-year, measured on the out-of-time window, for the \(\ge 10\)-PU population by displayed grade.
| Grade | Realized forward burden / PU-year |
|---|---|
| Excellent | 0.038 |
| Strong | 0.054 |
| Satisfactory | 0.068 |
| Marginal | 0.099 |
| Poor | 0.119 |
| Critical | 0.176 |
The gradient is strictly monotone and spans roughly five-fold from best grade to worst — a Critical fleet carries about \(5\times\) the forward crash burden of an Excellent one, on a window the engine never fitted. This is the single most important result in the paper: the grades order future risk, out of time, in the direction and by the magnitude an underwriter needs.
Aggregate discrimination is summarized by the exposure-weighted normalized ordered-Lorenz Gini: rank carriers safest-to-riskiest by their final relativity \(R_c\), plot cumulative realized burden against cumulative exposure, and divide the resulting Gini by the oracle that knows the realized order, so 1.0 is perfect ranking and 0 is random. Over the full graded population the engine reaches an overall normalized Gini of about 0.19. That book-level figure is deliberately conservative — it is dragged down by the thin-record small-fleet mass, where FMCSA data simply cannot separate carriers (§12). The honest picture is per cohort: where exposure is rich, the engine ranks sharply; where it is thin, it correctly declines to manufacture precision.
| Cohort | Power units | Normalized Gini |
|---|---|---|
| pu10_19 | 10–19 | 0.20 |
| pu20_49 | 20–49 | 0.25 |
| pu50_99 | 50–99 | 0.32 |
| pu100_199 | 100–199 | 0.41 |
| pu200_499 | 200–499 | 0.58 |
| pu500p | 500+ | 0.78 |
The pattern is exactly the one the design predicts: discrimination climbs monotonically with fleet size, from 0.20 in the ten-to-nineteen-truck cohort to 0.78 in the 500-plus cohort, because a bigger fleet supplies more power-unit-years and more inspections, so its credibility weight \(Z\) is higher and the grade rides its own experience rather than shrinking to the cohort mean. The engine ranks best precisely where the data is richest — and, conversely, does not overstate its power on the small cohorts, where it defers to the peer average by construction. Small-fleet cohorts (1–9 PU) are omitted from this table because their realized forward burden is statistically flat across \(R_c\) (§12): there is no gradient there to score, so reporting a Gini would report noise.
The full validation payload — overall and per-cohort Gini, the per-grade realized forward-burden gradient, the
monotonicity verdict, graded counts, and the graded distribution — is written to
data/processed/fred_v6_run_metrics.json on every refresh, so each run's discrimination is
recorded and can be compared against its predecessor to catch silent drift.
| Grade | Carriers | Share of graded |
|---|---|---|
| Excellent | 2,422 | 0.8% |
| Strong | 79,879 | 26.4% |
| Satisfactory | 177,949 | 58.8% |
| Marginal | 17,935 | 5.9% |
| Poor | 15,074 | 5.0% |
| Critical | 9,369 | 3.1% |
The absolute cut-points make the distribution right-skewed: the cohort average is \(R_c=1.0\) but the median carrier sits near 0.63, so most carriers land at or below their cohort average, with thin tails of Excellent and Critical. Small (1–9 PU) carriers cluster in Satisfactory — the census "Fleets only (10+ power units)" view surfaces the spread among rankable fleets (§12).
Inputs. 15 power units (cohort pu10_19), 18 roadside inspections last year with an
elevated vehicle out-of-service rate, and two crashes — one with a single injury, one tow-away — no fatalities.
Exposure (§3). Crash burden is rated over \(E = 15\) power-unit-years; the violation signals are rated over the carrier's inspection counts. No mileage enters the grade.
Observed burden (§4). \(B = (1+4) + 1 = 6\) over \(N = 2\) crashes.
Per-signal relativities (§5–6). Against its pu10_19 peers the carrier runs above
average on two signals — crash burden \(r_{\text{crash}} \approx 1.5\) and vehicle-OOS
\(r_{\text{veh}} \approx 1.6\) — and near 1.0 on the severe-violation, HOS, unsafe-driving, and speeding
signals. Each is the Empirical-Bayes posterior versus the cohort, so a thin record would have shrunk toward
1.0; this one has enough exposure to move off the peer mean.
Composite and floor (§7). The weighted geometric mean is \(R_{\text{own}} = 1.5^{0.20}\cdot 1.6^{0.189}\cdot(1.0)^{\dots} \approx 1.19\). The observed-crash floor is inert here: the burden is concentrated in the single injury crash, so its variance-aware discount \(\lambda = \max\!\big(0,\,1-\sqrt{\textstyle\sum_c w_c^2}\,/\,B\big) \approx 0.15\) pulls the raw crash relativity below \(R_{\text{own}}\). Hence \(R_c = \max(R_{\text{own}}, R_{\text{floor}}) = 1.19\).
Grade (§10). \(R_c = 1.19\) falls in the Marginal band \([0.90,\,1.20)\) — above the cohort average of 1.0 — for a Marginal grade, score \(\approx 31\). Exposure credibility \(Z_{\text{exp}} \approx 0.45\) puts confidence at Moderate; the Excellent upside cap is irrelevant to a below-average carrier, and the downside is never capped.
Fatal risk (§11). The fatal model returns \(\lambda_f \approx 0.02\), so \(\Pr(\text{≥1 fatal}) = 1 - e^{-0.02} \approx 2\%\) over the year.
Story told. “A 15-truck fleet whose crash burden and vehicle out-of-service rate both run above its same-size peers — about 19% above the cohort average once credibility-weighted — a Marginal read with enough record to trust. Two crashes last year, roughly one expected next; about a 2% chance of a fatal one. The read rests entirely on the carrier's own observed FMCSA record versus its 10–25-truck peers.”
| Constant | Value | Section |
|---|---|---|
| Crash maturity lag | 45 days | §2 |
| Crash exposure | power-unit-years | §3 |
| Violation exposure | inspection count | §3 |
| Fleet-size cohorts | 9 (pu1 … pu500p); 6 in-scope (10+) | §3 |
| Corrupt fleet-size guard | > 50,000 power units | §3 |
| Severity weights (per fatality / per injury / hazmat) | 12 / 4 / 3 | §4 |
| Severity caps (fatalities / injuries) | 3 / 5 | §4 |
| EB prior estimator | Bühlmann–Straub Gamma–Poisson, no trimming | §5 |
| Relativity clip | [0.01 , 100] | §5 |
| Signal weights (crash / veh-OOS / severe / HOS / unsafe / speeding) | 0.20 / 0.189 / 0.105 / 0.088 / 0.076 / 0.037 | §6–7 |
| Observed-crash floor discount | \(\lambda = \max(0,\,1-\sqrt{\sum_c w_c^2}/B)\) | §7 |
| Grade cut-points (absolute \(R_c\)) | 0.52 / 0.60 / 0.90 / 1.20 / 1.60 | §10 |
| Excellent confidence gate | \(Z_{\text{exp}} \ge 0.20\) | §10 |
| Single-power-unit upside cap | Strong | §10, §12 |
| Current FMCSA adverse rating | referral; grade capped at Marginal | §12 |
| Monotonicity gate | ≥10-PU, ratio-estimator SE, \(z = 2.0\) | §13 |
| Validation windows | fit [C−730, C−365); outcome [C−365, C) | §13 |
All per-cohort priors and constants above are re-estimated and logged on every weekly run
(data/processed/fred_v6_run_metrics.json); the values shown are the July 2026 production refresh.
§4 frames the severity weights as a compressed harm index rather than a cost-proportional scale. A fair question
follows: how much do the published grades depend on the chosen fatal weight? Because the severity caps
(≤3 fatalities, ≤5 injuries) mean the fatal weight moves only the burden of the small minority of carriers with
one to three fatal crashes, the answer is "very little." Re-running the whole engine — per-signal relativities,
the geometric composite, the observed-crash floor, and the absolute cut-points — with the per-fatality weight
moved from 12 to 8 and to 20 (the fatal8 / fatal20 severity schemes) leaves grades
highly stable:
| Fatal weight | Spearman vs production | Worst cohort Spearman | Same grade | Within one grade |
|---|---|---|---|---|
| 8 (vs 12) | 0.996 | 0.995 | 96.9% | 99.99% |
| 20 (vs 12) | 0.996 | 0.992 | 95.6% | 99.96% |
Carrier rankings are nearly invariant: rank correlations exceed 0.99 in every cohort, about 96% of carriers keep the identical letter grade, and well over 99.9% move at most one grade. The grade measures which carriers generate harm; the fatal weight mostly rescales how much harm — and the level is re-anchored by calibration regardless. Tail severity itself is priced by the dedicated fatal-crash model of §11.
This document specifies the production FRED scoring engine as of July 2026. Every formula, constant, and threshold is taken directly from the live scoring code; figures use either the stated production run or a faithful simulation of the named formula. The engine is re-fit and re-scored weekly on the full federal dataset.