← Risk register SOC 15-2051 · reviewed 2026-08-11

Data Scientists

262,440 US workers · median $120,230/yr · Tech

EXPOSED

The modal data scientist spends most of the week on work AI already does credibly: writing pandas/SQL, cleaning and joining tables, building baseline models, tuning hyperparameters, generating charts, and writing up findings in decks and notebooks. What resists is upstream and downstream — turning a vague business question into a measurable target, deciding whether an observed lift is causal, catching leakage and biased sampling before a model ships, and standing behind a recommendation that moves pricing, credit, or headcount. No license protects the role, and it is fully screen-based, so the shrinkage lands on the analysis-execution tier while the framing-and-accountability tier holds.

10-year outlook: By 2035 fewer people write the analysis and more people are paid to define the question and sign off on the decision; teams get smaller and skew senior.

US employment, 2021–2025+147.6%
105,980262,440 workers

Headcount grew steadily across the period.

The job count is not the verdict

This line is counted by the Bureau of Labor Statistics — the one figure on this page that isn't a judgement of ours. Headcount moves on demand, offshoring, demographics and the business cycle, and automation is one term among several, often not the loudest.

So a falling line is not evidence that AI did it, and a rising one is not evidence that it won't. Both happen in this register: some occupations resist automation and shrink anyway, others are highly automatable and keep growing.

BLS projection, 2024–2034

+33.5% 245,900 → 328,300 on the projections basis

Exposed, but growing

AI can already do a lot of these tasks, and the BLS still expects +33.5% more of these jobs by 2034. Demand for the output is growing faster than the work is being automated away — the mechanism BLS gives for software developers, and the combination people most often misread as an error.

Different clocks. The score is what current AI could do to this work today. The projection is how many of these jobs will exist in 2034. Everything between the two — how fast employers actually adopt, whether demand grows in the meantime — is why they can point opposite ways without either being wrong.

~23,400 openings a year on average, including replacing people who leave.

One email if this score changes. Watch as many occupations as you like from the same address — no account, and nothing is sent on a schedule, only when a verdict actually moves.

Also known as — 25 job titles this covers

Titles reported by people doing this work, from the US Department of Labor's O*NET survey. If your job title is here, this page is about your work even though the name doesn't match.

AnalystAbstractorData AnalystData ModelerData EngineerData ArchitectData EconomistData ScientistData ConsultantData SpecialistReports AnalystBusiness AnalystData CoordinatorResearch AnalystApplied ScientistReporting AnalystTableau DeveloperResearch ScientistApplication AnalystBusiness ConsultantData Mining AnalystData Science InternInformation AnalystStatistical Analyst

Added by hand, not from the survey. O*NET last sampled titles before some of these were in common use, so these are our judgement that the title belongs here — treat them as weaker than the list above. How we decide.

Analytics Engineer

Score — 37/100 resistance

Holding it up: judgment & accountability (14/20). Weakest point: liability shield (2/20).

Five dimensions, 0–20 each, summed. Higher means more protected. The arithmetic is shown so you can check it: 10 + 2 + 2 + 9 + 14 = 37. · Scored 2026-08-11, and re-examined when evidence accumulates rather than on a schedule.

Task resistance 10/20

Mixed — a routine tier and a judgment tier At 10, roughly half your week — feature engineering scripts, gradient-boosting baselines, cross-validation loops, matplotlib panels, notebook write-ups — is now competently drafted from a prompt, while defining the label for "churn" in a business with three contract types, spotting that your training set leaked post-outcome fields, or choosing a difference-in-differences design over a naive A/B split still needs you at the whiteboard; it sits below 14 because those framing hours are a minority of logged time, and above 6 because a model that ships without them fails in production.

Embodiment 2/20

Fully desk- and screen-based A 2 reflects that everything from data pull to stakeholder deck happens on a laptop against cloud compute, with the only physical residue being whiteboard sessions and in-person readouts that videoconference replaces without loss.

Liability shield 2/20

No licence, no signature requirement At 2 there is no board, no exam, no continuing-education requirement — a self-taught bootcamp graduate and a PhD hold the same title, and when a scoring model produces disparate impact under ECOA or the FCRA, the bank's compliance officers and counsel answer for it while your name appears nowhere on a filing.

Trust premium 9/20

Some relationship component A 9 recognizes that the product manager who has watched you flag three bad metrics will accept your "that lift is seasonality" without a deck, and that institutional knowledge of which internal tables are trustworthy is genuinely personal — but the artifacts you ship are dashboards and models that keep working after you leave, so the relationship accelerates the work rather than being the work.

Judgment & accountability 14/20

Exists to be accountable for ambiguous calls 14 is earned by the calls with no correct procedure: whether a 0.7% AUC gain justifies a model that is unexplainable to a regulator, whether to tell leadership their favored experiment was underpowered from the start, whether known selection bias in the training population is tolerable for a pricing or credit-limit decision — you own these before anyone else can audit them, though it stops short of 17+ because a director or model-risk committee typically signs the deployment.

Confidence: high · reviewed 2026-08-11 · how scoring works

What this job involves — and which parts are yours

The verdict above describes this occupation as a whole. Almost nobody does the typical version of a job — tick what's actually in your week and see how your own mix sits.

AI already does these at usable quality

These still need a person

Active moats on the surviving side: judgment, trust

How to future-proof this job

Where to go deeper on what this job runs on: Khan Academy — reading and vocabulary, all levels, free free · Coursera — critical thinking and logic, audit free free to audit · Coursera — active listening and communication skills free to audit · Toastmasters — public speaking practice at local clubs worldwide low · Coursera — critical thinking and logic, audit free free to audit · MIT OpenCourseWare — full course materials across every department, free free

All 35 skills ranked by how many jobs they open →

Where this experience transfers — nothing clears the bar

No occupation passed every test: close enough to data scientists on skills and subject matter, at least 10 points more resistant, no big jump in training, no new licence, no pay cut, and not shrinking on its own. That happens for 223 of the 654 occupations here that aren't SAFE, and it is worth stating plainly rather than leaving the section off.

The usual reason is that exposure travels with the skill profile. The jobs most similar to yours tend to be exposed for the same reasons yours is, so the near neighbours don't clear the gap — and the ones that do are a different kind of work, not a transfer of what you already know. Read that as a limit of this method, not a verdict that you're stuck: it only compares whole occupations, and it cannot see specialisation, industry, or anything you'd bring that isn't in a federal skill survey.

Here is that claim on your own job rather than in the abstract. These are the three occupations closest to this one by skill and subject matter — the places the work would most naturally transfer — with what the register scores them:

Statisticians EXPOSED 37/100 (+0) · 83% overlap
Actuaries EXPOSED 45/100 (+8) · 78% overlap
Statistical Assistants COOKED 14/100 (-23) · 73% overlap

That is the whole problem in three lines. The nearest work is not meaningfully safer, so there is no move here that trades a similar skill set for a better verdict. This is not us running out of ideas — it is what the neighbourhood looks like.

What would move this occupation up is the other direction, and on this page it's the more useful one.

What would move this back up — beyond any one person

The moves above are yours to make. This is the other half: what would have to change in the world for the occupation itself to score higher. None of it is in any one person's gift, but it is where the floor actually comes from. Scores here are not a one-way ratchet. Only two of the five dimensions — task resistance and embodiment — track what machines can do. The other three track law, what buyers will pay for, and who is answerable, and those move in both directions, often in response to the same pressure AI creates. If every lever below landed, this occupation would score around 58/100, still EXPOSED.

5 specific changes that would raise this score
  • already happening liability shield +7

    Model-risk governance rules extending SR 11-7-style validation duties beyond banks: EU AI Act high-risk obligations (Annex III: credit scoring, employment screening, insurance pricing) require a named human to document data governance, bias testing and sign the conformity assessment; Colorado SB 24-205 and NYC Local Law 144 bias-audit regimes push the same. If firms designate data scientists as the accountable signer of record for high-risk model documentation, this rises from 2 to 8-10.

  • already happening task resistance +4

    Task-mix shift, no law needed: as codegen absorbs the pandas/SQL/baseline-model/charting tier, the remaining week concentrates on causal identification, experiment design under interference, leakage and sampling-bias detection, and metric definition — areas where AI output cannot be verified without the same expertise. If routine execution falls below ~30% of hours, task_resistance moves toward 13-14, though the count of workers needed to do the residual falls.

  • plausible judgment accountability +4

    Formalized model-risk sign-off with personal attribution: an internal model inventory naming an owner per model, plus regulator-facing attestation (as bank MRM already does) making the data scientist the person who answers for a bad decision, not an anonymous team.

  • plausible liability shield +3

    Insurer requirement: tech E&O / AI liability policies conditioning coverage on a documented human validation step by a named quantitative reviewer before model deployment — mirrors how cyber insurers made MFA a condition.

  • plausible trust premium +3

    Narrow route only: expert-witness, regulatory-submission, and litigation/audit contexts (FDA statistical review, antitrust damages models, algorithmic discrimination cases) where a courtroom or agency requires a deposable human author of the analysis. Buyers pay for the attributable human, not the analysis.

The limit. Screen-based work keeps embodiment near 2 permanently, and no licensure body (no CPA/PE equivalent) exists or is being seriously proposed for data science, so liability_shield gains would come via employer/insurer designation rather than a true personal license — a weaker, more revocable shield. Even with all levers, headcount can shrink sharply while the residual role scores higher.

These are conditions, not forecasts — what would have to happen, not what will. Specific rules, cases and bills are named so you can go and check whether they exist and where they stand; verify before relying on any of them. Nothing here is legal or financial advice.

Where this work is, and what it pays there

BLS metro figures for 288 areas. The verdict above does not change by city — the rubric judges what the work involves, not where it happens — but pay and headcount do, and the national median hides a very wide range.

Most of these jobs

New York-Newark-Jersey City, NY-NJ 23,160 $135,980 +13%
San Francisco-Oakland-Fremont, CA 10,460 $170,110 +41%
Dallas-Fort Worth-Arlington, TX 10,120 $127,750 +6%
Los Angeles-Long Beach-Anaheim, CA 9,850 $129,740 +8%
Washington-Arlington-Alexandria, DC-VA-MD-WV 9,260 $132,200 +10%
Seattle-Tacoma-Bellevue, WA 8,370 $164,740 +37%
Chicago-Naperville-Elgin, IL-IN 7,940 $107,640 -10%
Boston-Cambridge-Newton, MA-NH 7,930 $132,040 +10%

Best paid

San Jose-Sunnyvale-Santa Clara, CA 6,060 $185,080 +54%
San Francisco-Oakland-Fremont, CA 10,460 $170,110 +41%
Idaho Falls, ID 230 $167,840 +40%

Percentages are against this occupation's national median of $120,230. Counts are jobs in that metro, not vacancies. Metros where the BLS suppressed the cell are absent rather than shown as zero.

Who is actually doing this

The score above is about what the work exposes. This is reporting about real deployments in this occupation — the difference between "could be automated" and "somebody automated it."

1 of 1 reported case, with sources

Quick take — do you do this job?

Has AI actually changed your work? One tap, anonymous, and the running tally is public. Nothing else is asked of you.

Self-reported and unverified — a sentiment signal, not a survey. One response per person per occupation; you can change your answer.

Field reports — what people say has changed

No field reports yet. A written account takes a paragraph rather than a tap, goes to an editor before it appears, and is the one thing on this page the rubric cannot produce on its own.

File a field report

Concrete beats general: a tool that arrived, a task that moved, a headcount decision you watched happen. Don't include anything that identifies you or your employer if that would put you at risk.

Watch this verdict
Kept current

Rather than check back: get the digest and we'll tell you what changed — or watch a single occupation from its own page.