Recommended Free Tools
These 21 cheat sheets are a revision system, not a promise of the exact questions any employer will ask. Start by matching them to the role: a product data scientist may spend more time on SQL, experimentation, metrics, and communication, while an ML-focused role may emphasize modeling, evaluation, and system design. Interview loops commonly include a recruiter or hiring-manager screen, project discussion, coding, statistics, modeling, a business case, and behavioral interviews, but employers vary. Coursera’s 2026 preparation guide describes a similar progression.
Use each sheet as a compact reference: definitions and formulas, decision rules, common traps, and one worked practice prompt. Then close it and solve problems from memory. Current guidance from DataCamp also stresses statistics, machine learning, Python or R, SQL, experimentation, product sense, and communication.
Choose your interview track first
| Track | Prioritize | Usually secondary |
|---|---|---|
| Product or business data scientist | SQL, experimentation, probability, metrics, product cases, communication, behavioral | Advanced ML systems and calculus |
| Generalist data scientist | Statistics, SQL, Python, ML, experimentation, projects | Deep infrastructure unless stated |
| ML-focused or applied scientist | Probability, linear algebra, calculus, modeling, evaluation, feature engineering, ML systems | Dashboard-focused topics |
| Data analyst | SQL, spreadsheets, descriptive statistics, visualization, business cases | Advanced modeling |
| New graduate | Fundamentals, coding fluency, projects, clear reasoning | Large-scale deployment detail |
| Experienced candidate | Impact, trade-offs, deployment constraints, stakeholder management, measurement | Basic syntax review only |
Turn the job description into an assessment matrix: copy each requirement, link it to sheets, and assign a practice task. “Data scientist” is not a standardized syllabus; inspect the interview stages, tools, domain, and whether the team explores data, predicts outcomes, estimates causal effects, or operates models.
The 21 cheat sheets
1. Interview map and role calibration
Know: role type, stages, tools, domain, mathematics level, and whether work is exploratory, predictive, causal, or operational. Trap: assuming every interview is an ML interview. Practice: make a one-page matrix from one real job description.
#1 Best Overall
2. Descriptive statistics
Review mean, median, mode, range, variance, standard deviation, quantiles, percentiles, IQR, covariance, correlation, skew, outliers, and population versus sample statistics.
mean = (1/n)Σxᵢ; s² = (1/(n−1))Σ(xᵢ−x̄)². The mean is outlier-sensitive; the median is more robust. Standard deviation describes dispersion, not uncertainty in the mean, and correlation does not prove causation. Practice: explain which center and spread you would report for a skewed income distribution.
3. Probability fundamentals
Be able to use conditional probability, independence, Bayes’ theorem, expected value, variance, conditional expectation, and total probability:
P(A|B)=P(A∩B)/P(B); P(A|B)=P(B|A)P(A)/P(B); E[X]=ΣxP(X=x). Practice: calculate the positive predictive value of a fraud alert when the base rate is low; explain the base-rate error.
4. Distributions and sampling
Know when Bernoulli, binomial, Poisson, uniform, normal, exponential, and t distributions fit; also know the central limit theorem, law of large numbers, and sampling bias. The central limit theorem concerns sample means under suitable conditions; it does not make every raw dataset normal. Practice: choose a model for daily support-ticket counts and defend the assumptions.
Rank #2
- Careercup, Easy To Read
- Condition : Good
- Compact for travelling
5. Confidence intervals and hypothesis tests
Review null and alternative hypotheses, Type I and II errors, power, significance level, p-values, confidence intervals, one- versus two-sided tests, practical significance, multiple comparisons, and minimum detectable effect. A p-value is not the probability that the null is true; a 95% interval is not a 95% probability statement about a fixed parameter. Repeated peeking and many unplanned metrics raise false-positive risk. Practice: explain a non-significant result with its interval and business implications.
6. A/B testing and experimentation
- Define the decision and hypothesis.
- Choose primary and guardrail metrics.
- Specify the experimental unit, sample size, and duration.
- Randomize and check instrumentation, balance, and sample-ratio mismatch.
- Analyze treatment effects, uncertainty, and carefully pre-specified segments.
- Decide whether to ship, iterate, or stop.
Check contamination, novelty, seasonality, interference, network effects, multiple testing, peeking, sequential methods, and ratio-metric variance. Practice: design an experiment for a checkout change and list what would invalidate it.
7. Regression
For linear regression, explain coefficients, intercepts, residuals, linearity, independence, homoscedasticity, multicollinearity, regularization, and when residual normality matters for inference. For logistic regression, know log-odds, odds ratios, thresholds, imbalance, calibration, ROC-AUC, and precision-recall. Check assumptions against the use case rather than reciting them.
Free tools Windows power users keep installed
One-click scans. No signup required.
8. Machine-learning algorithm selection
| Problem | Candidate methods | Discuss |
|---|---|---|
| Continuous prediction | Linear regression, random forest, gradient boosting | Error metric, interpretability, nonlinearities |
| Binary classification | Logistic regression, trees, boosting | Thresholds, imbalance, calibration |
| Multiclass | Multinomial models, ensembles, neural networks | Macro versus micro metrics |
| Clustering | K-means, hierarchical, density methods | Scaling, distance, validation |
| Ranking or recommendation | Learning-to-rank, collaborative filtering, two-stage systems | Offline versus online metrics |
| Time series | Naive baselines, classical, boosted, deep models | Temporal splits and leakage |
Choose using data volume, feature types, latency, interpretability, retraining, robustness, label quality, and error costs—not a universally “best” algorithm.
9. Bias, variance, overfitting, and regularization
Review underfitting, train-validation-test separation, learning curves, L1 and L2 penalties, early stopping, cross-validation, augmentation, and simple baselines. If training is excellent but validation is poor, investigate leakage, distribution shift, complexity, validation design, duplicates, hyperparameter overfitting, label noise, and preprocessing done before splitting.
Rank #3
10. Model-evaluation metrics
Classification: accuracy, precision, recall, F1, ROC-AUC, PR-AUC, log loss, calibration, confusion matrix. Regression: MAE, MSE, RMSE, R², limited use of MAPE, quantile loss. Ranking: precision@k, recall@k, NDCG, MAP, coverage, diversity. Tie the metric to the decision and false-positive or false-negative cost; accuracy can mislead on imbalanced data.
11. Feature engineering and leakage
Cover missing values, categorical encoding, scaling, bins, interactions, dates, text, aggregates, selection, and dimensionality reduction. Leakage includes future fields, full-dataset aggregates, post-outcome status, target-derived features, random splits of temporal data, and transformations fitted on validation or test data.
- Split according to the real prediction setting.
- Fit transformations on training data only.
- Apply identical transformations to validation and test data.
- Use pipelines and audit feature availability at prediction time.
12. Python fundamentals
Review lists, tuples, sets, dictionaries, comprehensions, iterators, generators, scope, exceptions, classes, sorting keys, and time and space complexity. Practice frequency counts, deduplication, grouping, two pointers, sliding windows, hash maps, stacks, queues, normalization, and nested-data parsing. Data-science screens often test practical manipulation, not only algorithm puzzles; DataCamp’s preparation guide identifies Python, R, and SQL as common requirements.
13. NumPy and vectorization
Know array shapes, broadcasting, boolean masks, axes, reshaping, aggregation, matrix multiplication, random sampling, missing values, views versus copies, and numerical stability. Practice replacing loops with vectorized operations and explain why elementwise multiplication differs from matrix multiplication.
14. pandas manipulation
Practice selection, filtering, sorting, groupby, aggregation, joins, concatenation, pivots, melt, missing values, datetimes, ranks, rolling calculations, and duplicate handling.
Rank #4
df.groupby("category", as_index=False)["revenue"].sum()
df.merge(dim_users, on="user_id", how="left")
df["event_time"] = pd.to_datetime(df["event_time"])
State join cardinality and expected grain before coding; a one-to-many join can silently duplicate revenue.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →15. SQL foundations
Review SELECT, filters, grouping, HAVING, CASE, nulls, joins, subqueries, CTEs, dates, strings, and set operations. A left join can become an inner join when right-table filters sit in WHERE; null is not zero or an empty string; aggregate only after confirming the grain.
16. SQL windows and advanced patterns
Practice ROW_NUMBER, RANK, DENSE_RANK, LAG, LEAD, running totals, moving averages, first and last events, sessionization, deduplication, cohorts, retention, and percent-of-total.
ROW_NUMBER() OVER (PARTITION BY user_id ORDER BY event_time)
Before writing a window function, identify the partition key, ordering key, frame, output grain, and tie behavior. DataCamp’s interview guide specifically highlights joins, windows, subqueries, date logic, and optimization reasoning.
17. Visualization and communication
Match chart to question: distributions, trends, category comparisons, relationships, and uncertainty. Use accessible colors, honest axes, denominators, annotations, and intervals. Answer in this order: what happened, why it might have happened, how certain you are, what the business should do, and what you would investigate next.
Best Value
18. Product metrics and business cases
Know north-star, input/output, leading/lagging, guardrail, funnel, retention, churn, engagement, conversion, revenue, marketplace, quality, and reliability metrics.
- Clarify the product goal, user, and use case.
- Map the journey and choose a primary metric plus guardrails.
- Segment results and diagnose causes.
- Recommend an action and define success measurement.
Do not recite a generic metric list; tie every metric to the decision.
19. Machine-learning system design
Structure answers around objective and constraints, then labels, training data, features, offline and online evaluation, serving, batch versus real-time inference, latency, monitoring, drift, retraining, feedback loops, fairness, safety, and rollback. Practice ranking, fraud, search, recommendations, forecasting, and moderation prompts.
20. Project and resume walkthrough
- Problem and why it mattered.
- Data source and limitations.
- Baseline and approach.
- Evaluation design and results.
- Deployment or operational use.
- Failure modes, changed decisions, and what you would do differently.
Expect follow-ups on metric choice, model choice, leakage, production behavior, uncertainty, and stakeholder influence.
21. Behavioral and interview execution
Prepare STAR examples—Situation, Task, Action, Result, and reflection—for conflict, failure, ambiguity, trade-offs, bad news, competing priorities, learning, and influence without authority. In live rounds, clarify assumptions, state SQL grain, compare alternatives, test edge cases, check results, admit uncertainty precisely, and summarize in business terms.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to turn sheets into preparation
- Read a sheet once, then hide it.
- Reproduce formulas or syntax from memory.
- Solve two or three interview-style questions.
- Explain one concept aloud.
- Record the mistake, correct rule, similar question, and retest date.
- Repeat with untimed learning, timed practice, and verbal practice.
Rereading PDFs alone will not build coding fluency or communication. Do not use notes or external tools in an assessment unless the employer explicitly permits them.
Seven-day crash plan
- Day 1: SQL foundations, joins, grain, and windows.
- Day 2: Python, NumPy, and pandas.
- Day 3: Descriptive statistics, probability, and distributions.
- Day 4: Testing, confidence intervals, and A/B experiments.
- Day 5: ML selection, leakage, regularization, and metrics.
- Day 6: Product case, project walkthrough, and visualization.
- Day 7: Full mock, behavioral answers, and error-log review.
Thirty-day plan
- Week 1: Statistics, probability, SQL fundamentals, and role calibration.
- Week 2: Timed SQL, Python, NumPy, and pandas practice.
- Week 3: Modeling, experimentation, metrics, product cases, and system design as relevant.
- Week 4: Mock interviews, project stories, behavioral practice, and targeted retesting.
Cheat sheets versus courses and paid platforms
Sheets are best for final-week revision, syntax recall, formula comparison, and checklists. They are poor substitutes for learning statistics from scratch, building coding fluency, understanding assumptions, handling ambiguous cases, or receiving feedback.
| Product | Strongest use | Main limitation | Pricing signal |
|---|---|---|---|
| DataLemur | Targeted SQL and data-science questions | Not a complete fundamentals course | $15/month; $60/year; $300 coaching/book/lifetime package, shown on the official page on August 18, 2026 |
| StrataScratch | Realistic SQL/Python practice and data cases | Numeric price not established here | Verify the live pricing page |
| LeetCode Premium | Algorithms and data structures | Not a complete data-science curriculum | $35/month; $159/year shown in USD on the cited page |
| DataCamp | Structured courses, projects, and broad fundamentals | Less targeted than an interview bank | Free basic; paid annual-billing displays varied between $14 and $27.50/month |
| Interview Query | Data-science question bank and explanations | Numeric price not established here | Verify the live pricing page |
| Interview books | Portable, linear reference | No adaptive feedback or live simulation | Varies by edition and retailer |
Prices, counts, promotions, and features change by date, geography, tax, and eligibility. Vendor testimonials and claims about interview relevance are not independent evidence. LeetCode is useful when algorithms are explicitly tested, but it should supplement—not replace—statistics, experimentation, product, and data-manipulation practice. Realistic platforms such as StrataScratch describe SQL, Python, statistics, product sense, and system-design practice on their official site, but no platform guarantees hiring outcomes.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Final readiness checklist
- Can I explain my project in two minutes and defend every major choice?
- Can I state the grain before writing SQL?
- Can I select and justify a metric?
- Can I explain a p-value and confidence interval accurately?
- Can I identify leakage and a flawed validation split?
- Can I discuss model trade-offs and operational constraints?
- Can I communicate uncertainty to a nontechnical stakeholder?
- Can I answer behavioral questions with specific results and reflection?
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




