{
  "source": "Cooked Index — occupational AI risk register",
  "page": "https://cookedindex.com/jobs/data-scientists/",
  "methodology": "https://cookedindex.com/methodology",
  "notice": "Verdicts are re-examined as evidence accumulates. Re-fetch before relying on this; the page above always carries the current score.",
  "scored_at": "2026-08-11",
  "model": "claude-opus-5",
  "occupation": {
    "title": "Data Scientists",
    "soc_code": "15-2051",
    "category": "Tech",
    "us_employment": 262440,
    "median_annual_wage": 120230
  },
  "verdict": "EXPOSED",
  "risk_resistance": 37,
  "contested": false,
  "near_boundary": false,
  "dimensions": {
    "task_resistance": 10,
    "embodiment": 2,
    "liability_shield": 2,
    "trust_premium": 9,
    "judgment_accountability": 14
  },
  "reasoning": {
    "task_resistance": "At 10, roughly half your week — feature engineering scripts, gradient-boosting baselines, cross-validation loops, matplotlib panels, notebook write-ups — is now competently drafted from a prompt, while defining the label for \"churn\" in a business with three contract types, spotting that your training set leaked post-outcome fields, or choosing a difference-in-differences design over a naive A/B split still needs you at the whiteboard; it sits below 14 because those framing hours are a minority of logged time, and above 6 because a model that ships without them fails in production.",
    "embodiment": "A 2 reflects that everything from data pull to stakeholder deck happens on a laptop against cloud compute, with the only physical residue being whiteboard sessions and in-person readouts that videoconference replaces without loss.",
    "liability_shield": "At 2 there is no board, no exam, no continuing-education requirement — a self-taught bootcamp graduate and a PhD hold the same title, and when a scoring model produces disparate impact under ECOA or the FCRA, the bank's compliance officers and counsel answer for it while your name appears nowhere on a filing.",
    "trust_premium": "A 9 recognizes that the product manager who has watched you flag three bad metrics will accept your \"that lift is seasonality\" without a deck, and that institutional knowledge of which internal tables are trustworthy is genuinely personal — but the artifacts you ship are dashboards and models that keep working after you leave, so the relationship accelerates the work rather than being the work.",
    "judgment_accountability": "14 is earned by the calls with no correct procedure: whether a 0.7% AUC gain justifies a model that is unexplainable to a regulator, whether to tell leadership their favored experiment was underpowered from the start, whether known selection bias in the training population is tolerable for a pricing or credit-limit decision — you own these before anyone else can audit them, though it stops short of 17+ because a director or model-risk committee typically signs the deployment."
  },
  "rationale": "The modal data scientist spends most of the week on work AI already does credibly: writing pandas/SQL, cleaning and joining tables, building baseline models, tuning hyperparameters, generating charts, and writing up findings in decks and notebooks. What resists is upstream and downstream — turning a vague business question into a measurable target, deciding whether an observed lift is causal, catching leakage and biased sampling before a model ships, and standing behind a recommendation that moves pricing, credit, or headcount. No license protects the role, and it is fully screen-based, so the shrinkage lands on the analysis-execution tier while the framing-and-accountability tier holds.",
  "outlook": "By 2035 fewer people write the analysis and more people are paid to define the question and sign off on the decision; teams get smaller and skew senior.",
  "what_would_raise_it": {
    "levers": [
      {
        "dimension": "liability_shield",
        "change": "Model-risk governance rules extending SR 11-7-style validation duties beyond banks: EU AI Act high-risk obligations (Annex III: credit scoring, employment screening, insurance pricing) require a named human to document data governance, bias testing and sign the conformity assessment; Colorado SB 24-205 and NYC Local Law 144 bias-audit regimes push the same. If firms designate data scientists as the accountable signer of record for high-risk model documentation, this rises from 2 to 8-10.",
        "plausibility": "already happening",
        "would_add": 7
      },
      {
        "dimension": "liability_shield",
        "change": "Insurer requirement: tech E&O / AI liability policies conditioning coverage on a documented human validation step by a named quantitative reviewer before model deployment — mirrors how cyber insurers made MFA a condition.",
        "plausibility": "plausible",
        "would_add": 3
      },
      {
        "dimension": "task_resistance",
        "change": "Task-mix shift, no law needed: as codegen absorbs the pandas/SQL/baseline-model/charting tier, the remaining week concentrates on causal identification, experiment design under interference, leakage and sampling-bias detection, and metric definition — areas where AI output cannot be verified without the same expertise. If routine execution falls below ~30% of hours, task_resistance moves toward 13-14, though the count of workers needed to do the residual falls.",
        "plausibility": "already happening",
        "would_add": 4
      },
      {
        "dimension": "judgment_accountability",
        "change": "Formalized model-risk sign-off with personal attribution: an internal model inventory naming an owner per model, plus regulator-facing attestation (as bank MRM already does) making the data scientist the person who answers for a bad decision, not an anonymous team.",
        "plausibility": "plausible",
        "would_add": 4
      },
      {
        "dimension": "trust_premium",
        "change": "Narrow route only: expert-witness, regulatory-submission, and litigation/audit contexts (FDA statistical review, antitrust damages models, algorithmic discrimination cases) where a courtroom or agency requires a deposable human author of the analysis. Buyers pay for the attributable human, not the analysis.",
        "plausibility": "plausible",
        "would_add": 3
      }
    ],
    "ceiling_note": "Screen-based work keeps embodiment near 2 permanently, and no licensure body (no CPA/PE equivalent) exists or is being seriously proposed for data science, so liability_shield gains would come via employer/insurer designation rather than a true personal license — a weaker, more revocable shield. Even with all levers, headcount can shrink sharply while the residual role scores higher."
  },
  "adjudication": null,
  "employment_history": {
    "points": [
      {
        "y": 2021,
        "emp": 105980,
        "wage": 100910
      },
      {
        "y": 2022,
        "emp": 159630,
        "wage": 103500
      },
      {
        "y": 2023,
        "emp": 192710,
        "wage": 108020
      },
      {
        "y": 2024,
        "emp": 233440,
        "wage": 112590
      },
      {
        "y": 2025,
        "emp": 262440,
        "wage": 120230
      }
    ],
    "from": 2021,
    "to": 2025,
    "change_pct": 147.6,
    "comparable_from": 2021,
    "spans_soc_revision": false
  },
  "pivots": [],
  "license": "https://cookedindex.com/terms"
}