Data & AI · Updated 16 September 2026

Data Scientist salary and career roadmap

Use statistics, experiments and models to answer questions the company cannot answer by querying the database.

Salary data checked · outlook from BLS projections released · by Bilal Tahir

$120,230BLS median (May 2025) source
$110,000Median entry salary source
$180,000Senior median source
$199,130BLS 90th percentile source
+35%10-yr growth source
24,800Openings per year source
18-36 monthsTime to first job source
$200-$400Typical cert cost source
What it is
A Data Scientist is a specialist who estimates what would happen rather than reporting what already did, by designing and reading out A/B tests, building forecasting and propensity models, and using causal inference when a clean experiment is impossible.
Salary
A Data Scientist in the United States earns a median of $110,000 entering the field, $140,000 at mid-career and $180,000 at senior level; the top end is $199,130. (BLS Occupational Employment and Wage Statistics, May 2025 (15-2051), widened with Levels.fyi, )
Outlook
Employment is projected to grow 35% over the ten years to 2036, with about 24,800 US openings a year. (BLS Occupational Outlook Handbook, 15-2051 Data Scientists, )
Time to first job
A career changer starting from zero typically needs 18 to 36 months at 10-15 hours a week to reach a first offer.
Cost to get in
The cheapest verified route in (PhD to industry) costs about $0; the most expensive costs about $12,000.
Roadmap
Our Data Scientist roadmap is 8 steps and about 890 study hours; the fastest route in is Domain expert converting at about 12 months.
Degree
No degree is legally required, but a quantitative master's or PhD is a hard filter in pharma, biostatistics, clinical research and quantitative finance; product analytics teams at technology companies and most mid-size employers are where demonstrated skill substitutes for the credential.
Automation exposure
Medium. Large language models already do exploratory analysis, boilerplate feature engineering and first-draft modelling code well, which erodes the junior end of Data Scientist work; the US Bureau of Labor Statistics cites the integration of AI into business workflows as a reason the occupation is projected to grow 35 percent to 2035. Experiment design, causal inference, spotting an artefact and persuading an organisation to act are not close to automated.

What does a Data Scientist do?

A Data Scientist is a specialist who estimates what would happen rather than reporting what already did, by designing and reading out A/B tests, building forecasting and propensity models, and using causal inference when a clean experiment is impossible. A data scientist sits between the analyst and the engineer. The same work is also posted as Decision Scientist, Product Data Scientist, Applied Scientist, Research Scientist, Machine Learning Scientist.

A data scientist sits between the analyst and the engineer. Where an analyst reports what happened, a data scientist estimates what would happen: designs and reads out A/B tests, builds forecasting and propensity models, does causal inference when a clean experiment is impossible, and builds the metric frameworks the rest of the company argues over. In practice the job splits into two archetypes. Product / decision science - heavy on SQL, experimentation, causal methods and stakeholder influence, common at consumer tech, marketplaces and fintech. Modelling / applied science - heavy on Python, scikit-learn, feature engineering and model evaluation, common in insurance, credit risk, healthcare, pricing and demand forecasting. Read the job description carefully, because the interview loops for the two are almost entirely different.

People who thrive are comfortable with uncertainty and willing to say "the data cannot answer that." The best ones are unusually good writers: a data science deliverable is usually a document, and its influence depends on the argument rather than the model. Strong statistical intuition matters more than algorithm breadth - knowing when your sample is too small, when your holdout is leaking, when a lift is seasonality.

The tradeoffs are real. BLS puts the May 2025 median at $120,230 with 35% projected growth, so the field is not shrinking, but the entry-level rung is competitive and many postings still expect a quantitative master's or PhD - particularly in pharma, finance and research-heavy teams. A large share of the job is unglamorous: cleaning data, reconciling numbers, rebuilding a pipeline someone left behind, and discovering that the question was not worth answering. And since 2023 the title has been partly hollowed out at the top end by ML and AI engineering, which now absorb the production modelling work.

Why Data Scientist pay is high

Data scientists are paid for decisions made under uncertainty at scale. A pricing model, a churn intervention, a credit risk score or a correctly-read experiment moves revenue in percentage points on a large base, and being wrong is expensive in a way the company can measure. The supply side is genuinely constrained: the combination of statistical rigour, production-grade Python, business judgment and the ability to write a persuasive memo is rare, and the two halves of that combination are usually trained in different places. BLS records a $120,230 median against a $67,240 tenth percentile and a $199,130 ninetieth, and Levels.fyi US total compensation - skewed toward tech employers who attach equity - runs $190,000 at the median and $363,000 at the ninetieth. The spread within a single title is the tell: this is a role where the top quartile is doing quantitatively different work from the bottom quartile, and is paid accordingly.

What's good

  • BLS median of $120,230 with 35% projected growth to 2035 - one of the strongest wage-and-growth combinations in the US labour market.
  • Genuine intellectual variety: the same week can involve causal inference, forecasting and a metric argument.
  • Portable across industries; the statistics travels even when the domain does not.
  • Visible influence - a well-argued memo can redirect a roadmap or a pricing strategy.
  • Strong optionality: ML engineering, product management, analytics leadership and quantitative finance are all reachable.
  • Remote and hybrid work is widely available.

What's hard

  • Many postings still filter on a quantitative master's or PhD, particularly in pharma, finance and research teams.
  • The title is inconsistent - some 'data scientist' jobs are dashboard work, some are ML engineering, and the description often does not tell you which.
  • A large fraction of the job is data cleaning, reconciliation and pipeline archaeology.
  • Political friction: results that contradict a decision already made are unwelcome, and you deliver them anyway.
  • Entry-level competition is intense and junior hiring softened through 2025-26.
  • The best-paid production modelling work has migrated to ML and AI engineering titles, which pay meaningfully more.

What a Data Scientist does all day

  • 09:00 - Read the overnight experiment dashboard; one test tripped a guardrail metric and needs a decision before the 10am product sync.
  • 09:30 - Product sync: argue that the metric drop is a logging change, not a real regression, and commit to confirming by noon.
  • 10:30 - Deep work: write SQL to reconstruct the event stream before and after the release; confirm the instrumentation theory.
  • 12:30 - Retrain the churn model on last quarter's data; the AUC improved and you spend 45 minutes finding out why before trusting it (it was leakage from a feature computed after the label date).
  • 14:00 - Review a junior analyst's experiment design; push back on a sample size that gives 35% power.
  • 15:00 - Write the memo for the pricing test: methodology, result, caveats, recommendation. Two hours, and it is the most valuable thing you do this week.
  • 16:30 - Pair with a data engineer on a dbt model so your feature logic stops living in a notebook.
  • Weekly - Experiment review board where every team's readouts get challenged.
  • Quarterly - Planning: defend which analyses and models are worth a quarter of someone's time.

Data Scientist salary in 2026: by level

US, annual, USD. Base plus typical bonus where the source reports it.

A US Data Scientist earns a median of $110,000 entering the field, $140,000 at two to four years and $180,000 at senior level, with the top end at $199,130. These are base-salary figures in US dollars as of September 2026, synthesised from BLS Occupational Employment and Wage Statistics, May 2025 (15-2051), widened with Levels.fyi. Pay varies about 20-35% by metro.

BLS OES May 2025 for Data Scientists (SOC 15-2051): median $120,230, 10th percentile $67,240, 90th percentile $199,130, 275,600 employed. Top-paying industries are publishing/broadcasting/content providers ($142,240 median), computer systems design ($132,380), credit intermediation ($129,490), management of companies ($128,050) and insurance carriers ($108,650) - a useful reminder that the same title pays very differently by sector. Levels.fyi US data scientist total compensation is higher because it over-samples tech: $103,875 (10th) / $140,000 (25th) / $190,000 (median) / $260,000 (75th) / $363,000 (90th), with base at $170,000 median. Levels.fyi entry level is $120,570 median total comp (10th $76,383, 90th $200,000) and senior is $235,000 median (10th $144,000, 90th $450,000). Big-tech entry medians on Levels.fyi: Amazon L4 about $194,892, Google L3 about $179,938, Atlassian P30 about $173,550, Microsoft level 59 about $157,788, Meta IC3 about $156,345. Structure: base plus 10-20% target bonus everywhere, plus RSUs at public tech companies which can be 25-50% of the package at senior level and are the main reason the tech distribution has such a long right tail. Non-tech employers (insurance, healthcare systems, government, most of the Fortune 500 outside tech) pay base-heavy packages 20-40% below the Levels.fyi figures. Geography: SF Bay Area, Seattle and NYC run 20-35% above the national median.

35%projected 10-year growth
24,800openings per year
mediumautomation exposure

BLS projects data scientist employment to grow 35% from 2025 to 2035 - much faster than the 3% all-occupation average - adding 95,400 jobs from a 2025 base of 275,600, with about 24,800 openings a year once replacement demand is counted. BLS attributes the growth to demand for data-driven decisions and the integration of AI into business workflows. Two caveats. First, that code is broad: it absorbs senior analysts, applied scientists and some ML practitioners, so headline growth overstates how easy the classic 'data scientist' job is to get. Second, the entry rung is the tight part of the market - Indeed's Hiring Lab reported entry-level postings down 7.5% year over year as of May 2026 with hiring tilting toward seniority. Automation risk is medium: LLMs already do exploratory analysis, boilerplate feature engineering and first-draft modelling code well, which erodes the junior end. Experiment design, causal inference, knowing which result is an artefact, and persuading an organisation to act are not close to automated.

“Employment of data scientists is projected to grow 35 percent from 2025 to 2035, much faster than the average for all occupations.”

US Bureau of Labor Statistics, Occupational Outlook Handbook, Data Scientists (15-2051), source,

Data Scientist salary by city

National bands scaled by metro wage differentials from the BLS May 2025 OEWS release.

Data Scientist pay is highest in San Jose / Silicon Valley (mid-career median about $198,800, ×1.42 the national figure) and lowest among large metros in Salt Lake City (about $133,000). The multiplier moves the offer, not what you keep after rent and state tax.

Data Scientist median pay by US metro, 2026, USD per year, derived from the national bands above.
MetroEntryMidSeniorvs national
San Jose / Silicon Valley$156,200$198,800$255,600×1.42
San Francisco Bay Area$148,500$189,000$243,000×1.35
New York City$140,800$179,200$230,400×1.28
Seattle$132,000$168,000$216,000×1.2
Boston$126,500$161,000$207,000×1.15
Washington DC metro$123,200$156,800$201,600×1.12
Los Angeles$118,800$151,200$194,400×1.08
Chicago$115,500$147,000$189,000×1.05
Austin$115,500$147,000$189,000×1.05
San Diego$114,400$145,600$187,200×1.04
Denver$112,200$142,800$183,600×1.02
Philadelphia$112,200$142,800$183,600×1.02
Dallas$110,000$140,000$180,000×1.0
Minneapolis$110,000$140,000$180,000×1.0
Raleigh-Durham$110,000$140,000$180,000×1.0
Houston$108,900$138,600$178,200×0.99
Atlanta$107,800$137,200$176,400×0.98
Phoenix$104,500$133,000$171,000×0.95
Miami$104,500$133,000$171,000×0.95
Salt Lake City$104,500$133,000$171,000×0.95

A multiplier raises the number on the offer letter, not what you keep. San Francisco pays about 35% more than the national median for these roles, but median Bay Area rent and California state income tax eat most of that for anyone below the senior rung. Texas, Florida, Washington, Tennessee and Nevada levy no state income tax, which is worth roughly 4-10% of take-home versus California or New York City, where city tax stacks on top of state tax. Run the comparison on after-tax income minus housing before you move. Fully remote roles are the edge case worth chasing: a national pay band spent in a 0.88 cost market beats a 1.35 salary spent in a 1.6 cost market for most people.

How the multipliers are derived

Anchor: the BLS May 2025 OEWS national mean wage across all occupations is $33.54/hr ($69,770/yr). Metro all-occupation means from the same release: San Jose-Sunnyvale-Santa Clara $57.32 (1.71x national), San Francisco-Oakland-Fremont $48.19 (1.44x), Washington-Arlington-Alexandria $44.20 (1.32x), Seattle-Tacoma-Bellevue $44.13 (1.32x), Boston-Cambridge-Newton $43.09 (1.28x), New York-Newark-Jersey City $41.50 (1.24x), Denver-Aurora-Centennial $39.28 (1.17x), Atlanta-Sandy Springs-Roswell $34.57 (1.03x), Chicago-Naperville-Elgin $34.42 (1.03x, May 2024), Dallas-Fort Worth-Arlington $33.96 (1.01x), Phoenix-Mesa-Chandler $33.48 (1.00x). We damp the top end of those raw ratios. All-occupation means exaggerate the gap for the careers on this site, because national pay bands, remote hiring and company-wide equity grids compress geographic spread for high-skill professional roles more than they do for service and hourly work. Levels.fyi shows the same damping: its Bay Area software engineer average total compensation of about $291k sits roughly 1.3x-1.4x the US median, not 1.7x. Non-US multipliers convert local market rates to USD and are directional, not survey-grade.

How to become a Data Scientist

Every route we could verify, with honest time, cost and difficulty.

There are 5 routes we could verify into Data Scientist work. The fastest is Domain expert converting at about 12 months; the cheapest is PhD to industry at about $0. The route that produces the most career-changer hires is analyst first and Data Scientist second: 18 to 30 months of production experience substitutes for a graduate degree in a way no certificate does, and you are paid throughout. Going direct from zero means competing with master's graduates for the same junior slot.

Promotion from data analyst

The most reliable route for someone without a quantitative graduate degree. Get an analyst job, spend 18-30 months owning experiments and forecasting inside one domain, then move to a data scientist title - usually by changing companies, because internal reclassification is slow. You arrive with production data experience and stakeholder credibility that no bootcamp graduate has.

30 months$500 cost

Quantitative master's degree

MS in statistics, data science, analytics, economics or operations research. Still the default filter for pharma, finance, insurance and research-heavy teams, and the cleanest path for someone with a non-quantitative bachelor's. Online options such as Georgia Tech's OMSA (about $10,000 total) and UT Austin's MSDS (about $10,000) are the cost-effective versions; full-time campus programmes run $50,000-$120,000.

24 months$12k cost

Domain expert converting

Actuaries, epidemiologists, economists, bench scientists, market researchers and financial analysts already hold the statistics half of the job. Adding Python, SQL and version control takes 6-9 months and the transition often happens inside the same industry at a similar or higher salary. The strongest career-change path if it applies to you.

12 months$800 cost

Self-taught with a serious portfolio

Possible without a degree, but the bar is much higher than for analyst roles: you need a Kaggle competition placement or an equivalent public project, an end-to-end deployed model, and clear evidence of statistical judgment rather than model-zoo tinkering. Works best at startups and mid-size non-tech companies. Budget 18-24 months and expect to apply broadly.

24 months$1k cost

PhD to industry

A quantitative PhD (statistics, physics, economics, computational biology, OR) still converts directly into senior or applied-scientist-track roles, often skipping the first rung entirely. Only relevant if you were already going to do one - it is a 4-6 year detour, not a career-change strategy.

60 months$0 cost

Data Scientist roadmap: 8 steps, 890 hours

Turn it into dates ·

Becoming a Data Scientist from zero takes about 890 study hours across 8 steps, roughly 18 to 36 months at 10-15 hours a week plus a job search. Step one is Python and SQL to working fluency. Python and SQL come first because they are the daily tools and the first technical screen. Statistics comes second because it is the round self-taught candidates most often fail. Experimentation and causal inference come later because they are the highest-value differentiator and only make sense once you can compute the estimates they depend on.

  1. 1

    Python: data structures, functions, list/dict comprehensions, pandas (merge, groupby, pivot, reshape, time-series indexing), NumPy vectorisation, virtual environments, and Git. SQL: everything in the analyst track plus window functions, query plans and how to avoid scanning a billion rows.

    Why now: These are the two tools you will use every day for the rest of your career, and the SQL screen is still the first technical filter in most data science loops. Weak pandas is the most common reason take-homes look amateurish.

    140 h this step140 h cumulative
  2. 2

    Random variables and distributions, expectation and variance, the central limit theorem, estimation, hypothesis testing and its failure modes, confidence intervals, power and sample size, multiple comparisons, Bayes' rule and conditional probability, bootstrapping and permutation tests. Do the exercises by hand and in Python.

    Why now: This is what separates a data scientist from an analyst, and it is the round most self-taught candidates fail. Interviewers probe statistical judgment - 'is this result real?' - far more than algorithm trivia. It is also the knowledge that stops you shipping an expensive wrong answer.

    150 h this step290 h cumulative
  3. 3

    Linear and logistic regression, regularisation, decision trees, random forests, gradient boosting (XGBoost/LightGBM), k-means and PCA, plus the parts that actually matter: train/validation/test discipline, cross-validation, class imbalance, calibration, precision-recall versus ROC, and how data leakage happens in practice.

    Why now: Every modelling loop tests this. Note the emphasis: interviewers care far less about whether you can name ten algorithms than about whether you can explain why your AUC of 0.97 is probably leakage. Andrew Ng's specialization is the standard grounding because it teaches the intuition rather than the library.

    130 h this step420 h cumulative
  4. 4

    Design an A/B test end to end: hypothesis, primary metric and guardrails, randomisation unit, power calculation and runtime, peeking and sequential testing, novelty and primacy effects, CUPED variance reduction, and what to do when the result is flat. Then the observational toolkit: difference-in-differences, propensity score matching, instrumental variables, regression discontinuity, synthetic control.

    Why now: In product and decision science this is the job, and it is the highest-leverage differentiator for a career changer because almost nobody self-studies it. It is also the part LLMs are worst at, because the hard question is what to measure and whether the comparison is valid, not how to compute it.

    100 h this step520 h cumulative
  5. 5

    Take one model from notebook to something that runs on a schedule: a clean repo, a requirements file, a training script, tracked experiments in MLflow, a FastAPI endpoint or a batch job in Airflow, tests, and a Docker image. Add dbt if your target industry uses a warehouse-centric stack.

    Why now: The most common criticism of self-taught candidates is 'notebook scientist'. Having one project that runs unattended and can be handed to someone else changes how your portfolio reads, and it is the skill that lets you move toward ML engineering later if the money there appeals.

    90 h this step610 h cumulative
  6. 6

    Project A: a causal or experimental analysis on public data - estimate the effect of something and defend the identification strategy. Project B: an end-to-end predictive model with a business framing, a deployed inference path, a monitoring plan, and an honest section on what you would do with more data. Each gets a README written for a hiring manager and a 600-word write-up.

    Why now: Hiring managers skim for evidence of judgment, not for volume. Two projects with real methodological decisions written up clearly beat a repo of twenty Kaggle notebooks. The write-up doubles as your answer to 'tell me about a project' in every interview.

    140 h this step750 h cumulative
  7. 7

    Decide between product/decision science (experimentation, causal inference, metric design, product sense) and applied modelling (forecasting, risk, pricing, recommender systems, deep learning). Then pick an industry - tech, insurance, healthcare, finance, retail - and learn its standard problems and regulatory constraints.

    Why now: The two archetypes interview completely differently, and generic 'data scientist' applications are what get filtered. Specialising also determines whether the master's degree question matters: pharma, biostatistics and quantitative finance effectively require one; product analytics and most tech roles do not.

    60 h this step810 h cumulative
  8. 8

    Round 1: timed SQL, window functions under pressure. Round 2: statistics and probability verbally - power, p-values, selection bias, simulation. Round 3: a take-home or case, typically 4-8 hours, judged on framing and write-up more than accuracy. Round 4: product sense or ML design, plus behavioural. Do at least six mock interviews with another person.

    Why now: Candidates lose offers on presentation, not knowledge - rambling case answers, take-homes with no conclusion section, models presented without a baseline. Practising out loud is the highest-return preparation hour there is. Expect 3-6 weeks from first recruiter call to offer.

    80 h this step890 h cumulative

Best certifications for a Data Scientist

Which ones matter, what they cost, and how often people pass.

No certification is required or expected for a Data Scientist job, and none of them substitutes for a degree or production experience. The most useful item on the list is the Machine Learning Specialization from Stanford Online and DeepLearning.AI at $49 a month for about two months and 95 hours, which is a curriculum rather than a credential. The Databricks Certified Machine Learning Associate at $200 is the only proctored exam here and it only matters if your target employers run Databricks; the DataCamp certifications carry little weight with hiring managers.

medium value

IBM Data Science Professional Certificate

IBM (via Coursera)

Cost
Free to audit; certificate included in Coursera Plus ($59/month or $399/year), or $49/month standalone
Study
About 160 (12 courses, 4 months at 10 hrs/week)
Pass rate
No exam; completion-based. ACE/FIBAA recommended for up to 12 US college credits.
high value

Machine Learning Specialization (Andrew Ng)

Stanford Online and DeepLearning.AI (via Coursera)

Cost
$49/month, about 2 months at 10 hrs/week; or included in Coursera Plus; financial aid available
Study
95 across 3 courses (33 + 34 + 28)
Pass rate
No exam; graded programming assignments. 4.9/5 from 39,344 reviews, 838,610 enrolled.
medium value

Databricks Certified Machine Learning Associate

Databricks

Cost
$200 registration fee (plus local tax); every Databricks exam is flat $200 as of 2026
Study
40-60 hours of prep with prior Spark/Python experience
Pass rate
Not published by Databricks
medium value

Google Advanced Data Analytics Professional Certificate

Google (via Coursera)

Cost
$49/month, under 6 months at 10 hrs/week (most finish for under $300); or included in Coursera Plus
Study
About 153 across 7 courses
Pass rate
No exam; completion-based
low value

DataCamp Data Scientist Certification (Associate and Professional)

DataCamp

Cost
Included with DataCamp Premium (about $28/month billed annually; about $35-$39/month monthly)
Study
Timed exams plus a practical exam, 30 days to complete from registration
Pass rate
Not published; certification valid two years

Skills employers screen for

Advanced SQL including window functions and query optimisationPython: pandas, NumPy, scikit-learn, statsmodelsInferential statistics: hypothesis testing, confidence intervals, power analysisExperiment design and A/B testing, including variance reduction (CUPED) and sequential testingCausal inference: difference-in-differences, propensity scores, instrumental variables, synthetic controlRegression, classification, tree ensembles (XGBoost/LightGBM) and honest model evaluationTime-series forecastingFeature engineering and leakage detectionData visualisation and written communication of resultsGit and reproducible notebooks/pipelinesPythonSQL (Snowflake, BigQuery, Databricks, Redshift)Jupyter / VS Codescikit-learnXGBoost / LightGBMstatsmodelsPyTorch (for the modelling archetype)dbtAirflow or DagsterMLflowGit / GitHubTableau or LookerR (still standard in pharma, biostatistics and academia)

Soft skills that decide offers: Framing a business problem as an estimable quantity, Writing memos that survive being forwarded without you in the room, Saying 'the data cannot answer this' and keeping the relationship, Prioritising: most requested analyses are not worth doing, Working with engineers without pretending to be one.

Best courses for a Data Scientist

Checked on the provider's page on 16 September 2026. Some links are affiliate links.

We list 6 courses for Data Scientist work, checked on the provider's page. Take one structured course per gap - Python and pandas, then statistics, then machine learning - and stop there, because the two deep projects in step 6 are what a hiring manager actually reads. Course spend can stay under $400 using Coursera Plus at $399 a year.

Coursera · Google Google Advanced Data Analytics Professional Certificate 153 h · $49/mo after 7-day free trial; included in Coursera Plus ($59/mo or $399/yr) · ★ 4.8 The natural second step after the beginner Google cert: it adds the Python, statistics and regression work that separates a $70k analyst from a $120k data scientist. DataCamp · DataCamp Data Analyst in Python 36 h · DataCamp Premium from $14/mo billed annually (Basic tier free) Browser-based coding with instant feedback means you actually write code from hour one, unlike video-first certificates where beginners stall. Coursera · University of California, Davis Learn SQL Basics for Data Science Specialization 61 h · Included in Coursera Plus ($59/mo or $399/yr); standalone subscription from $49/mo · ★ 4.6 SQL is the single most-tested skill in analyst interviews, and this specialization ends with a 35-hour capstone on real distributed data rather than toy tables. Coursera · IBM IBM Data Science Professional Certificate 158 h · Free to enroll; certificate included in Coursera Plus ($59/mo or $399/yr) · ★ 4.6 The broadest zero-to-portfolio data science program with nearly a million enrollments, and it ships with ACE and FIBAA credit recommendations. Coursera · Johns Hopkins University Data Science Specialization 287 h · Included in Coursera Plus ($59/mo or $399/yr); free to audit individual courses · ★ 4.5 The most statistically rigorous of the big data science programs; take it if you want to be able to defend your model choices, not just call .fit(). Coursera · University of Michigan Applied Data Science with Python Specialization 137 h · Included in Coursera Plus ($59/mo or $399/yr); free to enroll, paid certificate · ★ 4.5 Assumes you can already code and moves fast, so it is the efficient choice for career-changers coming from engineering or quantitative degrees.

Data Scientist interview questions and format

A Data Scientist loop is typically four to six rounds over three to six weeks: a recruiter screen, a timed SQL and Python technical screen, a take-home or live case of four to eight hours, a statistics and experimentation round, a product-sense or machine-learning design round, and a behavioural round with the hiring manager. The statistics and experimentation round is the stage that filters most self-taught candidates.

Typically 4-6 rounds over 3-6 weeks. (1) Recruiter screen. (2) Technical screen: timed SQL (45-60 minutes, window functions guaranteed) and often a short Python exercise. (3) Take-home or case: a dataset with an open question and a 4-8 hour expected effort, or a live case. Some companies have replaced take-homes with a 90-minute live analysis to blunt AI assistance; others allow AI and grill you on the output. (4) Statistics and experimentation round: A/B test design, power, p-values, selection bias, sometimes a simulation. (5) Product sense or ML design, depending on archetype - metric definition and diagnosis for product roles, end-to-end model design for modelling roles. (6) Behavioural with the hiring manager and a cross-functional partner. A/B testing, statistics, ML and product sense together account for over half of the questions asked.

Questions that come up

  • Design an A/B test for a change to the checkout flow. What is the primary metric, the randomisation unit, and how long does it run?
  • Your test shows +1.4% with p = 0.11 after two weeks. The PM wants to ship it. What do you say?
  • Explain p-value to a product manager in two sentences without using the word 'probability of the hypothesis'.
  • Your model's AUC jumped from 0.78 to 0.96 after you added a feature. What do you check first?
  • Daily active users fell 8% last Tuesday. Diagnose it.
  • When would you use logistic regression instead of gradient boosting even though boosting scores better?
  • How do you handle severe class imbalance, and why is accuracy the wrong metric?
  • You cannot run an experiment because the change rolled out to everyone at once. How do you estimate the effect?
  • Write SQL to compute 7-day rolling retention by signup cohort.
  • Walk me through a project where the honest answer was that the data could not support a conclusion.

Prep

Data Scientist FAQ

Do I need a master's degree to become a data scientist in 2026?

Not universally, but it matters more here than in any other data role. Technology product analytics, startups and most mid-size companies will hire on demonstrated skill; pharma, biostatistics, clinical research, quantitative finance and many Fortune 500 research teams treat a master's or PhD as a hard filter. If you lack one, the highest-probability path is analyst first and Data Scientist second, because two years of production experience substitutes for the degree in a way that no certificate does. If you do want the credential cheaply, online programmes such as Georgia Tech OMSA and UT Austin MSDS run around $10,000 in total.

Should I target data scientist or data analyst first if I am switching careers?

Analyst first, almost always, if you are switching without a quantitative degree. The skill overlap is roughly 60 percent, the hiring bar is much lower, and internal movement from analyst to Data Scientist is a well-worn path that takes 18 to 30 months. Targeting Data Scientist directly from zero means competing with master's graduates for the same junior slot in a market where Indeed's Hiring Lab reported entry-level postings down 7.5 percent year over year in May 2026. The exception is a quantitative background already in hand - actuarial, economics, epidemiology, physics - in which case go straight at it.

How much does data scientist pay differ between technology companies and non-technology employers?

Substantially, and the gap is mostly equity. The US Bureau of Labor Statistics puts the national median at $120,230 for May 2025, with publishing and broadcasting at $142,240 and insurance carriers at $108,650. Levels.fyi, which over-samples technology employers, shows a $190,000 median total compensation and $363,000 at the 90th percentile. A senior Data Scientist at a public technology company and one at a regional insurer can do similar work for a $100,000 difference in total compensation. The tradeoff is that the technology roles are far more competitive and restricted stock is volatile.

Is Kaggle worth the time for a self-taught data scientist with no work history?

In moderation, and only early. Kaggle teaches model tuning and gives you a public, verifiable signal, which is genuinely useful when you have no employment history in the field. It teaches almost nothing about problem framing, data collection, experiment design or stakeholder work, which is most of the job and most of what senior interviewers probe. Use it for one or two competitions to build credibility, then move on to the two deep projects in step 6 of this roadmap, where you had to decide what to measure and defend the decision in writing.

Will AI take the data scientist job, or just change what it involves?

It is changing the mix rather than removing the job. Large language models write exploratory analysis code, first-draft models and documentation faster than any human, which compresses the junior end and raises the bar for what a Data Scientist should spend time on. The durable parts are experiment design, causal reasoning, deciding which question is worth answering, and taking responsibility for a number a company acts on. The US Bureau of Labor Statistics explicitly cites AI integration as a growth driver for the occupation, projecting 35 percent growth to 2035, but the entry-level rung is where the pressure shows first.

Should I learn R or Python first to become a data scientist in 2026?

Python for almost everyone. R remains standard in pharma, biostatistics, academic research and some marketing-mix modelling, and if you are targeting those industries you should know it. Everywhere else Python is the default and candidates who only know R get filtered at the technical screen, which usually pairs SQL with a short pandas exercise. Learn Python to the level of pandas, NumPy, scikit-learn and statsmodels first, add R only if your target industry demands it, and do not spend time on the language argument itself.

How long does it take to get a first data scientist job while working full time?

Eighteen to 36 months of deliberate work at 10 to 15 hours a week if you are starting from a non-quantitative background and going direct, which is about 890 study hours plus a search. Via the analyst route it is 12 to 18 months to the analyst job and another 18 to 30 months to the Data Scientist title. That is longer on paper, but you are paid throughout and the success rate is much higher. Estimates shorter than a year assume you already have the statistics or the programming half of the job.

What does a Data Scientist actually do all day, hour by hour?

Reading experiments, writing SQL, and arguing about whether a number is real. A typical day opens with the overnight experiment dashboard and a product sync where you have to say that a metric drop is a logging change rather than a regression, then two hours of SQL reconstructing the event stream to prove it. Afternoons bring a model retrain where the improved AUC turns out to be leakage, a review of a junior colleague's underpowered test design, and two hours writing the memo for a pricing test. The memo is usually the most valuable thing produced that week.

Is data scientist a good career for someone switching at 40 from an unrelated job?

It depends on what your unrelated job was. If it was quantitative - actuarial work, economics, epidemiology, engineering, market research - you already hold half the skill set and the switch takes 6 to 12 months of Python, SQL and version control, often inside the same industry. If it was not, the honest timeline is 18 to 36 months, and the degree filter in pharma, finance and research teams will apply to you. At 40 the higher-probability plan is a Data Analyst job first, which pays a median of about $68,000 entering, then move across from inside.

Data scientist versus machine learning engineer: what is the actual difference?

A Data Scientist is accountable for whether a conclusion is true; a machine learning engineer is accountable for whether a model runs. The scientist designs experiments, estimates effects and writes the memo that changes a decision. The engineer builds training and serving infrastructure, handles latency, monitoring and retraining, and needs real software engineering rather than deeper statistics. Machine learning engineering pays more at the top and has absorbed most production modelling work since 2023. If you enjoy the argument about what to measure, stay on the science side; if you enjoy systems, move to engineering.

How much does a data scientist make in the San Francisco Bay Area versus the national median?

Roughly 35 percent more on the offer letter. Our location index puts the San Francisco Bay Area at 1.35 times the national band and San Jose at 1.42, against the US Bureau of Labor Statistics national median of $120,230 for May 2025 - so about $162,000 in the Bay Area and $171,000 in San Jose for the same seniority, before equity. Seattle sits at 1.2 and New York City at 1.28, while Atlanta at 0.98 and Phoenix at 0.95 sit just below national. California income tax and Bay Area rent absorb most of that premium below the senior rung.

Sources

Every number on this page traces to one of these. Page checked 16 September 2026.

  1. bls.gov/ooh/math/data-scientists.htm
  2. bls.gov/news.release/ocwage.t01.htm
  3. bls.gov/ooh/math/operations-research-analysts.htm
  4. levels.fyi/t/data-scientist
  5. levels.fyi/t/data-scientist/levels/entry-level/locations/united-states
  6. levels.fyi/t/data-scientist/levels/senior/locations/united-states
  7. levels.fyi/companies/meta/salaries/data-scientist
  8. levels.fyi/companies/google/salaries/data-scientist
  9. levels.fyi/companies/amazon/salaries/data-scientist
  10. hiringlab.indeed.com/2026/07/23/the-labor-market-is-tilting-toward-seniority/
  11. coursera.org/professional-certificates/ibm-data-science
  12. coursera.org/specializations/machine-learning-introduction
  13. coursera.org/professional-certificates/google-advanced-data-analytics
  14. coursera.org/courseraplus
  15. databricks.com/learn/certification
  16. datacamp.com/certification/data-scientist
  17. databricks.com/learn/certification/machine-learning-associate