Every occupation is scored 0–20 on five dimensions; the sum (0–100) maps to a verdict. A verdict you can't audit is just an opinion with a scary font, so the full rubric, anchors, and prompt live in the public repo.
The five dimensions
- Task resistance — how much of the day-to-day task list current AI already does at usable quality (inverted).
- Embodiment — physical work in unpredictable physical environments. Robotics lags language models badly.
- Liability shield — is a licensed human legally required to sign? Who gets sued?
- Trust premium — do buyers specifically pay for a human relationship, presence, or accountability?
- Judgment & accountability — does the role own consequential calls under ambiguity?
Verdicts
- SAFE 67–100 — AI changes the tools, not the job.
- EXPOSED 34–66 — the job survives but shrinks or splits; the routine tier is at risk.
- COOKED 0–33 — the core task list is already automatable; what remains is adoption speed.
What the verdicts are not
- Not a prediction with a date — we score capability exposure, not adoption timelines.
- Not a judgment about worth — "cooked" describes task exposure, not people.
- Not static — verdicts are re-scored as capabilities change; every page shows a review date, and disputes ship with credit.
How a verdict is actually produced
Every number on this site comes out of the same pipeline. It is
written down here so you can find the step you disagree with.
- Population. All 830 detailed
occupations from the BLS Occupational Employment and Wage
Statistics, May 2025 — the 6-digit SOC level. Broader groupings
are excluded because they double-count employment. No occupation
is chosen or dropped for being interesting.
- Blind scoring. Each occupation is scored on its
own, against the fixed rubric above, by claude-opus-5. The model
sees the occupation title, SOC code, category and employment
count — and nothing else. It does not see other occupations'
scores, so it cannot anchor on them or drift as a run progresses.
- Fixed calibration anchors. The prompt carries five
reference points with published score bands (data entry keyers,
paralegals, software developers, registered nurses, electricians).
They hold the scale steady across 830 independent calls. We
re-score those anchors to check the scale has not moved: the last
check landed 4 of 5 in band at medium reasoning effort and 5 of 5
at high.
- The score is computed, not asserted. The model
supplies five dimension scores; the total is calculated from them
here, and the verdict follows from the total by fixed thresholds.
A dimension outside 0–20 is rejected and the occupation re-scored.
The model's own opinion of the total is discarded.
- Justification, separately. A second pass takes
those scores as fixed and explains each one. It is permitted to
refuse — if a score cannot be defended it is recorded as disputed
rather than dressed in a plausible sentence. That makes the
justification an independent check, not a rationalisation.
- Reliability testing. A random sample is re-scored
from scratch in a fresh context and the two runs compared. The
agreement rate, the drift, and the direction of the drift are
published above. We test a random sample precisely because testing
only the wobbly ones would flatter the result.
- Adjudication, by a rule fixed in advance. 143 occupations have been scored twice. Publishing
whichever run happened to come first is a choice made by arrival
order, so we apply one rule to all of them: average the runs,
but only act on the result if it survives the arithmetic. We
recompute the merge five ways — per dimension and on the total,
rounding halves up, down, and to even — and change the verdict only
when every one of those agrees. 10 occupations passed that
test and were re-scored. 18 did not: their label depends
on how a halfway point is rounded, so we publish the original,
mark it contested, and show both runs on the page.
The remaining 115 had two runs that agreed.
We tried the obvious version of this first — average and
republish — and it would have changed 45 verdicts that flip back
the moment you round the other way. A verdict that depends on our
rounding convention is not a finding about anyone's job.
- Evidence, kept separate from scoring. Reported
deployments are collected and linked to occupations, but they are
not an input to the score. The rubric measures capability
exposure; the evidence shows what is being done about it. Keeping
them apart means a wave of press coverage cannot quietly move a
verdict.
- Re-scoring on evidence, not on a clock. Because
re-scoring drifts a couple of points on its own, a scheduled
refresh would mostly republish our own noise. An occupation is
re-examined when new deployments accumulate against it, and a new
score inside the measured noise band updates the review date
without changing the published number.
The scoring prompt, the rubric, every script in the pipeline, and
the raw per-occupation output are in the public repository. If you
think a verdict is wrong, you can read exactly how it was reached.
Known limits of this method
Every number here has a way of being wrong. These are the ones we
know about, stated before anyone has to find them.
- The verdict is a step function over a continuous score.
An occupation at 33 is COOKED and one at 34 is EXPOSED, off a
one-point difference. Re-scoring the same occupation moves results
by a point or two, so near a boundary the label is less stable than
the number behind it. 79 of 830 occupations
sit within one point of a boundary; each says so on its own page.
- EXPOSED is a wide band. It spans 34–66 and holds 524 of 830 occupations. For most of the
register the score and the reasoning carry the information, not the
chip.
- Scores are assigned by a language model, not measured.
Every occupation was scored by claude-opus-5 against the rubric
above, with the same prompt and fixed calibration anchors. That
makes it reproducible and auditable, not authoritative.
- We measured how reproducible it is, and publish the result.
Scored independently a second time, a random sample of 60
occupations produced the same verdict 97% of the time,
with scores moving a median of 2 points and no directional bias
(mean signed drift 0.35 against a standard deviation of 3.6 — the
instrument is imprecise, not optimistic or pessimistic).
The instability is concentrated at the edges. Re-scoring the
occupations that sit within a point of a threshold flipped 28 of them — those pages are roughly nine times likelier
to change label than a typical occupation. Averaging the two runs
resolves some of it, but only 10 cases robustly; the
other 18 are marked contested on their own pages, because
the merged answer changes with the rounding convention. The
underlying problem is not fixable by re-running: a continuous
score has to become one of three words, and near an edge that
conversion is arbitrary.
- The five dimensions are not independent, and we double-count physical work.
Across all 830 occupations, task resistance and
embodiment correlate at 0.86 — physical work
isn't screen work, so it scores high on both for the same
underlying reason. The other three cluster too. A factor analysis
of our own output finds two real dimensions
explaining 83% of the variance, not five: roughly "can a machine
physically do this" and "will law and custom permit it". Because
physicality occupies two of five slots, it carries about half the
score — an accident of how the rubric was divided, not a judgment
anyone made. The practical effect is that manual work scores
somewhat higher, and desk work somewhat lower, than a
non-duplicating instrument would produce. We keep all five because
they explain why a job is protected, which is what you'd
want to argue with; but the weighting is an artefact and you should
know it.
- The verdict boundaries are arbitrary, and they fall where
the data is densest.
Scores are roughly normal — mean 50, standard deviation 17 — so
the cutoffs at 33 and 66 sit almost exactly one standard deviation
either side of the middle, which is where most occupations are.
Nothing in the data clusters at those points; they are thirds of a
scale. That is why 79 occupations sit within a single
point of a different headline.
- The scorer knows where the thresholds are, and it aims at them.
The rubric states that 67 is the bottom of SAFE, and the model has
read it. The result is a pile-up exactly on the line: 35 occupations score precisely 67, while
only 3 scores 66 — a gap that cannot
happen by chance in a distribution this smooth. In practice the
model appears to decide whether a job is SAFE and then assemble
dimensions that clear the bar, which is the reverse of the process
described above. The artefact reproduces in independent re-runs, so
it is a property of the instrument rather than of one bad batch.
What this means when you read a page: for occupations at or near
67, the number is doing the work of a category label and its
apparent precision is not real. Re-scoring that group returned a
mean of 66.5 — the level is about right, so we have not
adjusted it. It is the granularity at the cutoff that is
manufactured, and you should not read "67 vs 66" as a measured
difference.
- There is no cost-to-automate term, and that matters.
Automation is an investment decision, not a capability threshold.
A job can be technically automatable for years and never automated
because the wage is low, the sites are small, or the work varies
too much to be worth the capital. We score what AI can do, not
what anyone will pay to deploy — so treat these as exposure, not
as a schedule.
- We score the modal worker, at the national median.
Not the elite tier, not the worst case, and not your specific
employer. A BLS title often bundles jobs that feel unrelated from
the inside.
- Escape hatches are computed, not curated. Pivot
suggestions come from overlap in O*NET skill and knowledge profiles
plus a training-level constraint. They are a starting point for
research, not careers advice, and some of them are wrong.
Votes are self-reported sentiment, rate-limited but unverified.