Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

20 Data Science Interview Questions to Probe Skills—Not “Fake” Candidates

Andrew Fogg’s 2016 KDnuggets article lists 20 data science interview questions, not 25. Here is what they cover and how to use them thoughtfully.
From TheFinanceBase Team5 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The source article behind this topic contains 20 interview questions, not 25. Andrew Fogg published the list on KDnuggets on January 1, 2016; the available source does not establish a 25-question version. Its prompts can help interviewers explore a candidate’s breadth, but they are not a validated test for detecting deception or predicting job performance. Use them to prompt reasoning, ask about assumptions and trade-offs, and adapt follow-ups to the role.

What the questions can—and cannot—tell you

Data science draws on more than one discipline. In the source article, Kirk Borne describes it as applying mathematical, computational, visual, analytical, statistical, experimental, problem-definition, model-building, and validation techniques to data. A candidate may be strongest in one area; the useful interview question is whether they can explain how their skills fit the work, recognize limits, and collaborate across disciplines.

“Fake” is a provocative label in the source title, not an evidence-based category. The questions have no published pass threshold or demonstrated predictive validity. A polished definition alone is weak evidence: ask candidates to explain their reasoning, state assumptions, identify failure modes, and connect concepts to relevant work. Evaluate answers against the responsibilities of the job rather than treating the list as a universal checklist.

The 20 questions in the source article

These are the topics and prompts listed by Fogg. Their wording is summarized here; the questions are grouped by the kind of work they probe. The first item retains the source’s unusually specific formulation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Modeling, validation, and predictive performance

  1. How would you validate a model created to predict a quantitative outcome using multiple regression?
  2. What is regularization, and how can it help?
  3. What is the difference between precision and recall?
  4. What are false positives and false negatives, and when does each matter?
  5. How can outliers affect a model, and how would you handle them?

Statistical reasoning and evidence

  1. What is statistical power?
  2. What is selection bias, why is it important, and how can you avoid it?
  3. What is the difference between correlation and causation?
  4. How would you interpret a statistical claim reported in a published study?
  5. How can you tell whether an apparent finding may be due to chance?

Experiments, data, and applications

  1. Give an example of using experimental design to answer a question about user behavior.
  2. What is the difference between “long” and “wide” format data?
  3. How would you approach a recommendation-system problem?
  4. How would you visualize data to communicate a result?
  5. How would you handle a rare event or highly imbalanced outcome?

Additional breadth prompts

  1. How would you decide which variables or features to use?
  2. What is resampling, and when might you use it?
  3. How would you explain a model result to a nontechnical audience?
  4. How do you decide whether a model is useful for the problem at hand?
  5. What would you do if a model performed well during development but poorly on new data?

The source’s broader subject areas include regularization and validation, precision and recall, power and resampling, false positives and negatives, selection bias, experimental design, data shape, interpreting published statistics, outliers and rare events, recommendation systems, and visualization. The questions above organize those areas for interview use; they should not be mistaken for a separately verified 25-item list.

How to probe answers rather than score definitions

Choose questions that match the role, then follow the candidate’s explanation. A research-oriented role may call for closer examination of study design and statistical assumptions; a product experimentation role may require clearer thinking about treatment, outcomes, and user behavior; an applied modeling role may emphasize validation, generalization, and decisions made from predictions. These are interview-design suggestions, not a comparative ranking established by the source articles.

  • Ask for assumptions: What must be true for the proposed method or conclusion to be credible?
  • Ask about failure modes: What could bias the data, distort the metric, or make a result fail to reproduce?
  • Ask about trade-offs: What would change if false positives were more costly than false negatives, or if the data were sparse?
  • Request an example: Invite the candidate to describe a relevant project, their contribution, the checks they used, and what they learned when results changed.
  • Match depth to the job: Distinguish conceptual understanding from hands-on implementation, and assess each only where it matters for the role.

Listen for a chain of reasoning, not just terminology: how the candidate defines the problem, chooses a method, tests assumptions, evaluates results, and communicates uncertainty. The source list does not supply a scoring rubric, so hiring teams should define job-specific evidence before interviews rather than infer a universal pass mark.

What strong answers may reveal about overfitting

The companion article by Gregory Piatetsky describes overfitting as finding results due to chance that cannot be reproduced by later studies. It warns that repeatedly testing hypotheses without appropriate statistical control can create findings that shrink or disappear on repetition. A candidate discussing validation should be able to explain how a method fits the data and question, what leakage or repeated testing could do, and how performance on new data will be assessed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The companion lists several ways to reduce overfitting risk: keep hypotheses simple, use regularization, conduct randomization testing, use nested cross-validation, adjust for false discovery rate, and preserve a reusable holdout set. These are not interchangeable fixes; the candidate should explain why a choice suits the data and evaluation plan. For technical background on feature reduction and regularization, the companion points to Statistical Learning with Sparsity: The Lasso and Generalizations.

Make experimental-design questions concrete

The companion’s illustrative example asks how page-load time affects user satisfaction. A useful follow-up is to ask what would be changed, what outcome would be measured, and how behavior would be recorded. The article discusses comparing page variants and measures such as latency, frequency, duration, or intensity. Ask the candidate to explain what comparison would answer the question and what could make the result misleading. This is an example from a 2016 article, not a universal experimental protocol.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Use long- and wide-format data as a reasoning prompt

The companion describes “tall” data as having many more records than features and “wide” data as having relatively few records and many features. That distinction can lead to a useful follow-up: how might the relationship between observations and features affect model choice or the risk of overfitting? The article cautions that methods suitable for tall data may overfit in wide settings and mentions feature reduction approaches such as Lasso. Ask candidates to connect the data shape to the specific problem, rather than treating “long” or “wide” as a test of vocabulary alone.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase09 OCT 267 minMortgage Escrow FAQs: Taxes, Insurance, Shortages, and Refunds
  2. The Money DeskBlogTheFinanceBase09 OCT 265 minHow Mortgage Escrow Accounts Work and What Homeowners Pay For
  3. The Money DeskBlogTheFinanceBase09 OCT 265 minHow to Read a Stock Chart, Volume and Market-Cap Data
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.