Methodology

Every occupation is scored 0–20 on five dimensions; the sum (0–100) maps to a verdict. A verdict you can't audit is just an opinion with a scary font, so the full rubric, anchors, and prompt live in the public repo.

The five dimensions

  • Task resistance — how much of the day-to-day task list current AI already does at usable quality (inverted).
  • Embodiment — physical work in unpredictable physical environments. Robotics lags language models badly.
  • Liability shield — is a licensed human legally required to sign? Who gets sued?
  • Trust premium — do buyers specifically pay for a human relationship, presence, or accountability?
  • Judgment & accountability — does the role own consequential calls under ambiguity?

Verdicts

  • SAFE 67–100 — AI changes the tools, not the job.
  • EXPOSED 34–66 — the job survives but shrinks or splits; the routine tier is at risk.
  • COOKED 0–33 — the core task list is already automatable; what remains is adoption speed.

What the verdicts are not

  • Not a prediction with a date — we score capability exposure, not adoption timelines.
  • Not a judgment about worth — "cooked" describes task exposure, not people.
  • Not static — verdicts are re-scored as capabilities change; every page shows a review date, and disputes ship with credit.

How a verdict is actually produced

Every number on this site comes out of the same pipeline. It is written down here so you can find the step you disagree with.

  1. Population. All 830 detailed occupations from the BLS Occupational Employment and Wage Statistics, May 2025 — the 6-digit SOC level. Broader groupings are excluded because they double-count employment. No occupation is chosen or dropped for being interesting.
  2. Blind scoring. Each occupation is scored on its own, against the fixed rubric above, by claude-opus-5. The model sees the occupation title, SOC code, category and employment count — and nothing else. It does not see other occupations' scores, so it cannot anchor on them or drift as a run progresses.
  3. Fixed calibration anchors. The prompt carries five reference points with published score bands (data entry keyers, paralegals, software developers, registered nurses, electricians). They hold the scale steady across 830 independent calls. We re-score those anchors to check the scale has not moved: the last check landed 4 of 5 in band at medium reasoning effort and 5 of 5 at high.
  4. The score is computed, not asserted. The model supplies five dimension scores; the total is calculated from them here, and the verdict follows from the total by fixed thresholds. A dimension outside 0–20 is rejected and the occupation re-scored. The model's own opinion of the total is discarded.
  5. Justification, separately. A second pass takes those scores as fixed and explains each one. It is permitted to refuse — if a score cannot be defended it is recorded as disputed rather than dressed in a plausible sentence. That makes the justification an independent check, not a rationalisation.
  6. Reliability testing. A random sample is re-scored from scratch in a fresh context and the two runs compared. The agreement rate, the drift, and the direction of the drift are published above. We test a random sample precisely because testing only the wobbly ones would flatter the result.
  7. Adjudication, by a rule fixed in advance. 143 occupations have been scored twice. Publishing whichever run happened to come first is a choice made by arrival order, so we apply one rule to all of them: average the runs, but only act on the result if it survives the arithmetic. We recompute the merge five ways — per dimension and on the total, rounding halves up, down, and to even — and change the verdict only when every one of those agrees. 10 occupations passed that test and were re-scored. 18 did not: their label depends on how a halfway point is rounded, so we publish the original, mark it contested, and show both runs on the page. The remaining 115 had two runs that agreed.

    We tried the obvious version of this first — average and republish — and it would have changed 45 verdicts that flip back the moment you round the other way. A verdict that depends on our rounding convention is not a finding about anyone's job.
  8. Evidence, kept separate from scoring. Reported deployments are collected and linked to occupations, but they are not an input to the score. The rubric measures capability exposure; the evidence shows what is being done about it. Keeping them apart means a wave of press coverage cannot quietly move a verdict.
  9. Re-scoring on evidence, not on a clock. Because re-scoring drifts a couple of points on its own, a scheduled refresh would mostly republish our own noise. An occupation is re-examined when new deployments accumulate against it, and a new score inside the measured noise band updates the review date without changing the published number.

The scoring prompt, the rubric, every script in the pipeline, and the raw per-occupation output are in the public repository. If you think a verdict is wrong, you can read exactly how it was reached.

Known limits of this method

Every number here has a way of being wrong. These are the ones we know about, stated before anyone has to find them.

  • The verdict is a step function over a continuous score. An occupation at 33 is COOKED and one at 34 is EXPOSED, off a one-point difference. Re-scoring the same occupation moves results by a point or two, so near a boundary the label is less stable than the number behind it. 79 of 830 occupations sit within one point of a boundary; each says so on its own page.
  • EXPOSED is a wide band. It spans 34–66 and holds 524 of 830 occupations. For most of the register the score and the reasoning carry the information, not the chip.
  • Scores are assigned by a language model, not measured. Every occupation was scored by claude-opus-5 against the rubric above, with the same prompt and fixed calibration anchors. That makes it reproducible and auditable, not authoritative.
  • We measured how reproducible it is, and publish the result. Scored independently a second time, a random sample of 60 occupations produced the same verdict 97% of the time, with scores moving a median of 2 points and no directional bias (mean signed drift 0.35 against a standard deviation of 3.6 — the instrument is imprecise, not optimistic or pessimistic).

    The instability is concentrated at the edges. Re-scoring the occupations that sit within a point of a threshold flipped 28 of them — those pages are roughly nine times likelier to change label than a typical occupation. Averaging the two runs resolves some of it, but only 10 cases robustly; the other 18 are marked contested on their own pages, because the merged answer changes with the rounding convention. The underlying problem is not fixable by re-running: a continuous score has to become one of three words, and near an edge that conversion is arbitrary.
  • The five dimensions are not independent, and we double-count physical work. Across all 830 occupations, task resistance and embodiment correlate at 0.86 — physical work isn't screen work, so it scores high on both for the same underlying reason. The other three cluster too. A factor analysis of our own output finds two real dimensions explaining 83% of the variance, not five: roughly "can a machine physically do this" and "will law and custom permit it". Because physicality occupies two of five slots, it carries about half the score — an accident of how the rubric was divided, not a judgment anyone made. The practical effect is that manual work scores somewhat higher, and desk work somewhat lower, than a non-duplicating instrument would produce. We keep all five because they explain why a job is protected, which is what you'd want to argue with; but the weighting is an artefact and you should know it.
  • The verdict boundaries are arbitrary, and they fall where the data is densest. Scores are roughly normal — mean 50, standard deviation 17 — so the cutoffs at 33 and 66 sit almost exactly one standard deviation either side of the middle, which is where most occupations are. Nothing in the data clusters at those points; they are thirds of a scale. That is why 79 occupations sit within a single point of a different headline.
  • The scorer knows where the thresholds are, and it aims at them. The rubric states that 67 is the bottom of SAFE, and the model has read it. The result is a pile-up exactly on the line: 35 occupations score precisely 67, while only 3 scores 66 — a gap that cannot happen by chance in a distribution this smooth. In practice the model appears to decide whether a job is SAFE and then assemble dimensions that clear the bar, which is the reverse of the process described above. The artefact reproduces in independent re-runs, so it is a property of the instrument rather than of one bad batch.

    What this means when you read a page: for occupations at or near 67, the number is doing the work of a category label and its apparent precision is not real. Re-scoring that group returned a mean of 66.5 — the level is about right, so we have not adjusted it. It is the granularity at the cutoff that is manufactured, and you should not read "67 vs 66" as a measured difference.
  • There is no cost-to-automate term, and that matters. Automation is an investment decision, not a capability threshold. A job can be technically automatable for years and never automated because the wage is low, the sites are small, or the work varies too much to be worth the capital. We score what AI can do, not what anyone will pay to deploy — so treat these as exposure, not as a schedule.
  • We score the modal worker, at the national median. Not the elite tier, not the worst case, and not your specific employer. A BLS title often bundles jobs that feel unrelated from the inside.
  • Escape hatches are computed, not curated. Pivot suggestions come from overlap in O*NET skill and knowledge profiles plus a training-level constraint. They are a starting point for research, not careers advice, and some of them are wrong.

Votes are self-reported sentiment, rate-limited but unverified.