Free tools Windows power users keep installed
One-click scans. No signup required.
Data-science interviews vary by role. Product and analytics candidates usually face more SQL, metrics, experimentation and stakeholder cases; machine-learning candidates can expect deeper modeling, coding, deployment and system-design questions. Use the questions below as a role-aware practice bank: explain the idea, state assumptions, work a small example and anticipate the follow-up.
Interviewers commonly assess data manipulation, experiment design, metric choice, hypothesis testing, machine-learning implementation and behavioral communication, as outlined in Microsoft’s interview-preparation guide. Broader preparation guides also group interviews into coding, statistics, machine learning, case studies and behavioral communication (Coursera; Yale career guidance).
How to prioritize the questions
| Target role | Prioritize |
|---|---|
| Product or analytics data scientist | SQL, funnels, retention, metrics, A/B testing, causal reasoning and communication |
| ML-focused data scientist | Validation, leakage, feature engineering, model evaluation, coding, deployment and monitoring |
| Research or deep-learning role | Probability, optimization, derivations, architectures and experimental rigor |
| Junior candidate | Fundamentals, SQL joins and aggregation, Python, statistics and one end-to-end project |
| Senior candidate | Trade-offs, prioritization, governance, monitoring, technical debt, mentoring and failure recovery |
Statistics and probability interview questions
1. Mean, median and mode
The mean is the arithmetic average and is sensitive to outliers. The median is the middle observation and is more robust for skewed data. The mode is the most frequent value and is useful for categorical data. Choose the summary that reflects the distribution and decision; household income, for example, is often better represented by a median.
Follow-up: Explain when a trimmed mean or a transformation would be preferable.
#1 Best Overall
2. Variance, standard deviation and standard error
Variance is average squared deviation from the mean. Standard deviation is its square root in the original units. Standard error describes uncertainty in an estimated statistic, such as a sample mean; standard deviation describes variation among observations.
3. Conditional probability and Bayes’ theorem
P(A | B) = P(B | A)P(A) / P(B). Explain prior, likelihood, evidence and posterior, then emphasize base rates. A rare-fraud detector can have high sensitivity yet generate many false positives when the underlying event is uncommon.
4. Correlation versus causation
Correlation measures association, not a causal effect. Confounding, reverse causality, selection bias and coincidence can create an association. A causal claim needs a treatment, outcome, counterfactual and identification strategy; randomization generally gives stronger evidence than observation.
5. What is a p-value?
It is the probability of results at least as extreme as those observed, assuming the null hypothesis and model assumptions are true. It is not the probability that the null is true or that the result occurred “by chance.” Report effect size, confidence intervals, sample size and design as well.
6. What is a confidence interval?
A 95% interval is produced by a procedure that would contain the true parameter in 95% of repeated samples under its assumptions. It is not, strictly, a 95% probability statement about a fixed parameter in this particular interval. Smaller samples and higher variance make intervals wider.
7. Type I, Type II error and power
Type I error is rejecting a true null (false positive); Type II error is failing to reject a false null (false negative). Power is 1 − β, the chance of detecting a specified effect. Sample size, variance, effect size and significance threshold determine power.
8. Central Limit Theorem
Under suitable conditions, the normalized sampling distribution of a mean approaches normality as sample size grows. This supports approximate intervals and tests; it does not make the raw data normal and does not cure dependence, biased sampling or extreme heavy tails.
9. Detecting and handling outliers
First determine whether a value is an error, a legitimate rare event or another population. Use domain rules, plots, IQR or robust statistics, transformations, winsorization or robust models as appropriate. Do not delete observations automatically; outliers can strongly influence linear coefficients and target definitions.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute10. Law of large numbers
With increasing independent observations, a sample average converges toward its expected value under the law’s assumptions. Unlike the CLT, it concerns convergence rather than the approximate distribution of the estimate. Dependence or biased sampling can invalidate a naive conclusion.
Experimentation and causal-inference questions
11. Design an A/B test
- Define the decision and hypothesis.
- Choose the randomization unit, primary metric and guardrails.
- Estimate sample size and duration.
- Predefine exclusions, stopping rules and analysis.
- Randomize, verify balance and monitor instrumentation.
- Estimate effects with uncertainty, inspect contamination and segments, then weigh practical impact.
Check novelty effects, interference, noncompliance and selective exposure before treating an estimate as causal.
12. Statistical versus practical significance
Statistical significance concerns compatibility with a null model; practical significance asks whether the effect changes a decision enough to matter. Large samples can make trivial effects significant. Include effect size, interval, cost, risk and expected business impact.
13. Sample-ratio mismatch
This is a material difference between intended and observed treatment allocation. Investigate randomization, eligibility, logging, bots and repeated users before interpreting results; it may invalidate the test.
Recommended Free Tools
14. Simpson’s paradox
An aggregate relationship can reverse after splitting by a relevant variable because of confounding or unequal group composition. Inspect segment-level results and the data-generating process.
15. Causal effect without an A/B test
Possible designs include difference-in-differences, regression discontinuity, instrumental variables, matching, synthetic controls and interrupted time series. State assumptions: parallel trends, exclusion restrictions or continuity at a threshold, for example.
SQL interview questions
16. INNER, LEFT, RIGHT and FULL OUTER JOIN
INNER keeps matches; LEFT preserves every left row; RIGHT preserves every right row; FULL preserves both sides. A predicate on the right table in WHERE can turn a left join into an inner join; use ON when unmatched left rows must remain.
17. WHERE versus HAVING
WHERE filters rows before aggregation. HAVING filters groups after aggregation, such as groups with COUNT(*) > 10.
Rank #3
18. Second-highest salary
Clarify whether “second” means distinct salary or second row. For the second distinct value:
SELECT salary
FROM (
SELECT salary, DENSE_RANK() OVER (ORDER BY salary DESC) AS r
FROM employees
) x
WHERE r = 2;
19. Window functions
They calculate rankings, running totals, moving averages and lag/lead comparisons while retaining row detail. Explain PARTITION BY, ORDER BY, frames, and the difference among ROW_NUMBER, RANK and DENSE_RANK.
20. Seven-day rolling average
Clarify calendar days versus observed rows, missing dates and whether today is included. ROWS BETWEEN 6 PRECEDING AND CURRENT ROW is seven rows, not necessarily seven calendar days; build a date spine when calendar days matter.
21. Duplicate records
Define the business key, group by it and find COUNT(*) > 1. Determine whether repeats are retries, legitimate events or copies, then deduplicate with a deterministic rule such as latest ingestion time and measure metric impact.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
22. Retention
Define cohort date, activation event, interval and qualifying activity. Assign each user to a cohort, calculate elapsed periods, count active users and divide by the original cohort. Clarify time zones, late events, reactivation and whether retention is user-, activity- or revenue-based.
23. Optimize a slow query
Inspect the execution plan; filter and project early where safe; verify keys, data types and cardinality; pre-aggregate before joins; use indexes, partitions, clustering or materialized views when the engine supports them. Identify whether scanning, sorting, shuffling, skew or network movement is the bottleneck.
Python, pandas and coding questions
24. List, tuple, set and dictionary
Lists are ordered and mutable; tuples are ordered and generally immutable; sets hold unique hashable elements; dictionaries map hashable keys to values. Choose by required ordering, mutation and membership behavior.
25. Shallow versus deep copy
A shallow copy creates a new outer object but shares nested references. A deep copy recursively copies nested objects, which can be expensive or unsuitable for objects tied to external resources.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRank #4
26. Missing values in pandas
Measure missingness by column, segment and time; determine whether it means unknown, not applicable, not collected or zero; test whether missingness is informative; then choose deletion, imputation, indicators or a domain category. Fit imputers on training data only.
27. Vectorization
Vectorized NumPy/pandas operations run optimized array code and are often faster and clearer than Python row loops. They are not universal: complex branching and memory limits may favor another approach, so benchmark.
28. Debug an unexpected pipeline result
- Reproduce on a small sample.
- Check row counts after each transformation.
- Inspect keys, join cardinality, nulls, duplicates, types, dates and time zones.
- Compare aggregates before and after steps and test known cases.
- Validate independently, add assertions and document prevention.
29. Mutable versus immutable objects
Mutable objects change in place; immutable objects do not. This affects aliasing, function arguments, caching and dictionary keys. Lists and dictionaries are mutable; strings, integers and tuples are immutable, although a tuple may contain mutable members.
30. Data larger than memory
Read chunks, select columns, filter early, use efficient dtypes and columnar formats, push work into SQL, aggregate incrementally or use Spark, Dask, Polars or an analytical database. Diagnose whether memory, CPU, disk or network is limiting.
Machine-learning questions
31. Bias-variance trade-off
High bias underfits; high variance overfits. Complexity can lower bias while raising variance. Regularization, more representative data, feature selection, ensembles and cross-validation target generalization, not merely training error.
32. Overfitting prevention
Separate train, validation and test data; use appropriate cross-validation, regularization, early stopping, simpler models, feature selection and realistic temporal or operational holdouts. Cross-validation estimates generalization but does not guarantee it (scikit-learn documentation).
33. Data leakage
Leakage occurs when prediction-time-unavailable information enters fitting or evaluation: post-outcome fields, future aggregates, preprocessing before splitting or incompatible entity overlap. It inflates offline scores and harms production performance (scikit-learn’s common-pitfalls guide).
34. Logistic regression versus trees
Logistic regression suits interpretable, sparse, high-dimensional problems with approximately linear log-odds, stable coefficients and low latency. Trees handle nonlinear interactions and mixed types with less preprocessing but may need calibration and are harder to explain.
Best Value
35. L1 and L2 regularization
Regularization penalizes the objective. L1 encourages sparse coefficients and feature selection; L2 shrinks coefficients and stabilizes correlated features. Select strength without using the final test set, and scale features for many regularized linear models.
36. Precision, recall, F1, ROC-AUC and PR-AUC
Precision is the share of predicted positives that are correct; recall is the share of actual positives found; F1 combines them. ROC-AUC evaluates ranking across thresholds, while PR-AUC is often more informative for rare positives. Choose a threshold and metric from error costs (Google’s metrics glossary).
37. Imbalanced classification
Use cost-appropriate metrics, class weights, threshold tuning, calibrated probabilities and carefully separated over/undersampling. Evaluate on a realistic validation set and quantify false-positive and false-negative costs.
38. Regression evaluation
MAE treats errors linearly; MSE and RMSE penalize large errors; R² measures explained variance relative to a baseline; MAPE is unstable near zero; quantile loss supports asymmetric costs and intervals. Inspect residuals by segment and time.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →39. Explain predictions
Use coefficients where appropriate, permutation importance, partial dependence, individual conditional expectation, SHAP-style explanations and counterfactuals. Distinguish predictive association from causality and test explanation stability.
40. Monitor a production model
Track schema, missingness, feature and prediction distributions, latency, throughput, calibration, business outcomes, delayed-label performance, drift, segment fairness and retraining triggers. Include ownership, alerts, versioning, rollback and incident response.
Advanced and role-specific questions
- Design a recommendation, retrieval or ranking system and choose offline and online metrics.
- Detect training-serving skew and design a feature-store contract.
- Split time-series data and distinguish covariate shift from concept drift.
- Compare batch with online inference and design capacity safeguards.
- Evaluate embeddings or a generative-AI application, including quality, safety, latency and cost.
- Design an experiment when users can receive multiple treatments.
- Explain a technical result to an executive, prioritize conflicting requests, or describe a project failure and what changed afterward.
Behavioral questions to prepare
- Tell me about yourself and your most relevant project.
- Describe an analysis that failed or changed direction.
- Tell me about a stakeholder disagreement.
- How did your work change a decision?
- How do you handle ambiguous requirements?
- How do you prioritize requests with competing deadlines?
Use STAR—Situation, Task, Action, Result—and add what you learned and would change.
A framework for strong answers
Conceptual questions
- Define the idea.
- Explain why it matters.
- Give a concrete example.
- State assumptions and limitations.
- Name a trade-off or follow-up.
SQL questions
Clarify grain and keys, define null and duplicate behavior, write and explain the query, test edge cases and discuss performance.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Modeling questions
Define target and prediction-time boundary, establish a baseline, choose a split and metric, validate, analyze errors and subgroups, then specify monitoring.
Case studies
Frame the decision, user, north-star metric and guardrails; identify data and assumptions; propose analysis or experiment; discuss confounders, risks and the action you would take.
Quick Recap
Final preparation checklist
- Map the job description to product, analytics, ML or research priorities.
- Practice joins, windows, retention and messy-data SQL under time limits.
- Explain statistics aloud without textbook shorthand.
- Prepare two end-to-end project stories with measurable outcomes.
- Rehearse one experiment case and one modeling case.
- Review leakage, validation, drift and monitoring.
- Prepare thoughtful interviewer questions and verify permitted tools.
Practice resources by need
| Need | Potential fit | Limitation |
|---|---|---|
| Structured beginner curriculum | Coursera or DataCamp | Less efficient when only targeted drills are needed; current pricing was not established here. |
| SQL and coding repetition | LeetCode, DataLemur or StrataScratch | Does not replace experimentation, case studies or behavioral practice. |
| Broad data-science practice | Interview Query | A full subscription may be unnecessary for a small free-practice need. |
| Mocks and feedback | Exponent or a human coach | Costs more than self-study and quality depends on feedback. |
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




