StatQuest statistics and machine learning library
Use for statistics, model evaluation, machine learning, neural network, and AI explainers.
A practical machine learning path covering problem framing, baselines, regression, classification, ensembles, unsupervised learning, neural networks, validation, leakage, fairness, and deployment monitoring.
A practical machine learning path covering problem framing, baselines, regression, classification, ensembles, unsupervised learning, neural networks, validation, leakage, fairness, and deployment monitoring.
This track is organized as a mastery loop: source study, sequenced checks, Khan report evidence, then tutor handoff only when the data shows a real stuck point.
Use linked OER and companion sources before attempting checks.
Move through numbered PeerTutor problems without skipping failed gates.
Enter Khan report evidence and let the adaptive plan rank repair units.
Bring exact misses, notes, and one sharp question to a tutor.
The bank is intentionally mixed across facets and difficulty so high scores cannot come from one narrow question style.
Use these linked courses for video instruction and mastery practice, then return here for PeerTutor original checks and tutor help on the exact unit that got stuck.
External video content is linked or embedded through provider-hosted players, not copied. PeerTutor practice is original and cites each source path separately.
Watch the strongest public video path inside PeerTutor, then use the unit checks below to prove the student can actually do the work.
Use for statistics, model evaluation, machine learning, neural network, and AI explainers.
Use for long-form project tutorials in Python, JavaScript, data science, machine learning, and web development.
Khan Academy progress has to be entered from the student or tutor report. Khan does not provide a supported public progress API, so this mirror stores unit status locally and uses it to target PeerTutor checks.
Khan Academy does not provide a supported public progress API or external API keys. PeerTutor stores student-provided report evidence and maps it to original practice instead.
0 repair units and 5 practice units need attention before extension.
No Khan mirror data yet; this is a normal practice candidate, not a proven weakness.
Complete 12 sequenced checks and advance only after misses are corrected.
No Khan mirror data yet; this is a normal practice candidate, not a proven weakness.
Complete 12 sequenced checks and advance only after misses are corrected.
No Khan mirror data yet; this is a normal practice candidate, not a proven weakness.
Complete 12 sequenced checks and advance only after misses are corrected.
No Khan mirror data yet; this is a normal practice candidate, not a proven weakness.
Complete 12 sequenced checks and advance only after misses are corrected.
No Khan mirror data yet; this is a normal practice candidate, not a proven weakness.
Complete 12 sequenced checks and advance only after misses are corrected.
Rebuild the unit: do 12 PeerTutor checks, log every miss, then ask a tutor from the error log.
Khan evidence to mirror: percent/mastery for "Problem framing and data boundaries", missed skill, last activity date, and the next Khan item assigned by the teacher report.
Rebuild the unit: do 12 PeerTutor checks, log every miss, then ask a tutor from the error log.
Khan evidence to mirror: percent/mastery for "Baselines, regression, and error", missed skill, last activity date, and the next Khan item assigned by the teacher report.
Rebuild the unit: do 12 PeerTutor checks, log every miss, then ask a tutor from the error log.
Khan evidence to mirror: percent/mastery for "Classification and probability metrics", missed skill, last activity date, and the next Khan item assigned by the teacher report.
Rebuild the unit: do 12 PeerTutor checks, log every miss, then ask a tutor from the error log.
Khan evidence to mirror: percent/mastery for "Trees and ensemble methods", missed skill, last activity date, and the next Khan item assigned by the teacher report.
Rebuild the unit: do 12 PeerTutor checks, log every miss, then ask a tutor from the error log.
Khan evidence to mirror: percent/mastery for "Unsupervised learning and representations", missed skill, last activity date, and the next Khan item assigned by the teacher report.
Rebuild the unit: do 12 PeerTutor checks, log every miss, then ask a tutor from the error log.
Khan evidence to mirror: percent/mastery for "Neural networks and optimization", missed skill, last activity date, and the next Khan item assigned by the teacher report.
Rebuild the unit: do 12 PeerTutor checks, log every miss, then ask a tutor from the error log.
Khan evidence to mirror: percent/mastery for "Validation, overfitting, and fairness", missed skill, last activity date, and the next Khan item assigned by the teacher report.
Rebuild the unit: do 12 PeerTutor checks, log every miss, then ask a tutor from the error log.
Khan evidence to mirror: percent/mastery for "Capstone and deployment monitoring", missed skill, last activity date, and the next Khan item assigned by the teacher report.
Built for independent progress first, then tutor support where the student gets stuck.
Practice: Turn a vague model idea into a supervised-learning spec with target, features, split rule, and baseline.
Start with Principles of Data Science and 6.0002 Introduction to Computational Thinking and Data Science. Take notes until you can explain: Define prediction target, label timing, features, unit of analysis, baseline, and decision use.
Do the first sequenced checks until the definitions, vocabulary, and setup are correct without hints.
Produce the assignment artifact, then pass the application and analysis checks tied to: Turn a vague model idea into a supervised-learning spec with target, features, split rule, and baseline.
Any miss becomes an error-log entry, a clean redo, and one nearby transfer problem before advancing.
Bring your attempted work, the exact missed check, and one question about: Detect leakage and ambiguous labels before modeling.
Move in order: foundation, transfer, analysis, timed readiness, then metacognitive repair.
Start the next sequenced check: #1 Trace check. Do not jump ahead until this one is correct.
Weakest facet: Concept
A loop for Problem framing and data boundaries runs once for each value 0 through 3. How many times does it run?
Which implementation is easiest to test?
Which target best matches the Machine Learning unit "Problem framing and data boundaries"?
Which practice artifact should you produce for "Problem framing and data boundaries" before asking a tutor for help?
Which target best proves readiness for "Problem framing and data boundaries"?
Which second target belongs to "Problem framing and data boundaries"?
Practice: Compare a regression model to a baseline and explain whether the improvement is meaningful.
Start with Principles of Data Science and 6.0002 Introduction to Computational Thinking and Data Science. Take notes until you can explain: Fit mean, linear, and regularized regression baselines.
Do the first sequenced checks until the definitions, vocabulary, and setup are correct without hints.
Produce the assignment artifact, then pass the application and analysis checks tied to: Compare a regression model to a baseline and explain whether the improvement is meaningful.
Any miss becomes an error-log entry, a clean redo, and one nearby transfer problem before advancing.
Bring your attempted work, the exact missed check, and one question about: Use MAE, RMSE, residuals, and calibration-style checks.
Move in order: foundation, transfer, analysis, timed readiness, then metacognitive repair.
Start the next sequenced check: #1 Trace check. Do not jump ahead until this one is correct.
Weakest facet: Concept
A loop for Baselines, regression, and error runs once for each value 0 through 4. How many times does it run?
Which implementation is easiest to test?
Which target best matches the Machine Learning unit "Baselines, regression, and error"?
Which practice artifact should you produce for "Baselines, regression, and error" before asking a tutor for help?
Which target best proves readiness for "Baselines, regression, and error"?
Which second target belongs to "Baselines, regression, and error"?
Practice: Choose a classification threshold from a confusion matrix and defend the tradeoff.
Start with Principles of Data Science and 6.0002 Introduction to Computational Thinking and Data Science. Take notes until you can explain: Use logistic regression, confusion matrices, precision, recall, specificity, ROC-AUC, and threshold tradeoffs.
Do the first sequenced checks until the definitions, vocabulary, and setup are correct without hints.
Produce the assignment artifact, then pass the application and analysis checks tied to: Choose a classification threshold from a confusion matrix and defend the tradeoff.
Any miss becomes an error-log entry, a clean redo, and one nearby transfer problem before advancing.
Bring your attempted work, the exact missed check, and one question about: Choose metrics based on false-positive and false-negative costs.
Move in order: foundation, transfer, analysis, timed readiness, then metacognitive repair.
Start the next sequenced check: #1 Trace check. Do not jump ahead until this one is correct.
Weakest facet: Concept
A loop for Classification and probability metrics runs once for each value 0 through 5. How many times does it run?
Which implementation is easiest to test?
Which target best matches the Machine Learning unit "Classification and probability metrics"?
Which practice artifact should you produce for "Classification and probability metrics" before asking a tutor for help?
Which target best proves readiness for "Classification and probability metrics"?
Which second target belongs to "Classification and probability metrics"?
Practice: Explain how a tree split improves prediction and how validation prevents overfitting.
Start with Principles of Data Science and 6.0002 Introduction to Computational Thinking and Data Science. Take notes until you can explain: Use decision trees, random forests, gradient boosting concepts, feature importance, and pruning.
Do the first sequenced checks until the definitions, vocabulary, and setup are correct without hints.
Produce the assignment artifact, then pass the application and analysis checks tied to: Explain how a tree split improves prediction and how validation prevents overfitting.
Any miss becomes an error-log entry, a clean redo, and one nearby transfer problem before advancing.
Bring your attempted work, the exact missed check, and one question about: Recognize when tree models overfit.
Move in order: foundation, transfer, analysis, timed readiness, then metacognitive repair.
Start the next sequenced check: #1 Trace check. Do not jump ahead until this one is correct.
Weakest facet: Concept
A loop for Trees and ensemble methods runs once for each value 0 through 6. How many times does it run?
Which implementation is easiest to test?
Which target best matches the Machine Learning unit "Trees and ensemble methods"?
Which practice artifact should you produce for "Trees and ensemble methods" before asking a tutor for help?
Which target best proves readiness for "Trees and ensemble methods"?
Which second target belongs to "Trees and ensemble methods"?
Practice: Interpret a clustering result with one useful insight and one reason it may be misleading.
Start with Principles of Data Science and 6.0002 Introduction to Computational Thinking and Data Science. Take notes until you can explain: Use clustering, dimensionality reduction, embeddings, and similarity measures.
Do the first sequenced checks until the definitions, vocabulary, and setup are correct without hints.
Produce the assignment artifact, then pass the application and analysis checks tied to: Interpret a clustering result with one useful insight and one reason it may be misleading.
Any miss becomes an error-log entry, a clean redo, and one nearby transfer problem before advancing.
Bring your attempted work, the exact missed check, and one question about: Validate unsupervised patterns without pretending they are ground truth.
Move in order: foundation, transfer, analysis, timed readiness, then metacognitive repair.
Start the next sequenced check: #1 Trace check. Do not jump ahead until this one is correct.
Weakest facet: Concept
A loop for Unsupervised learning and representations runs once for each value 0 through 7. How many times does it run?
Which implementation is easiest to test?
Which target best matches the Machine Learning unit "Unsupervised learning and representations"?
Which practice artifact should you produce for "Unsupervised learning and representations" before asking a tutor for help?
Which target best proves readiness for "Unsupervised learning and representations"?
Which second target belongs to "Unsupervised learning and representations"?
Practice: Sketch a small neural network for tabular or image input and explain loss, output, and overfitting controls.
Start with Principles of Data Science and 6.0002 Introduction to Computational Thinking and Data Science. Take notes until you can explain: Use layers, activation, loss, gradient descent, backpropagation intuition, and regularization.
Do the first sequenced checks until the definitions, vocabulary, and setup are correct without hints.
Produce the assignment artifact, then pass the application and analysis checks tied to: Sketch a small neural network for tabular or image input and explain loss, output, and overfitting controls.
Any miss becomes an error-log entry, a clean redo, and one nearby transfer problem before advancing.
Bring your attempted work, the exact missed check, and one question about: Understand what neural networks buy and what they make harder.
Move in order: foundation, transfer, analysis, timed readiness, then metacognitive repair.
Start the next sequenced check: #1 Trace check. Do not jump ahead until this one is correct.
Weakest facet: Concept
A loop for Neural networks and optimization runs once for each value 0 through 8. How many times does it run?
Which implementation is easiest to test?
Which target best matches the Machine Learning unit "Neural networks and optimization"?
Which practice artifact should you produce for "Neural networks and optimization" before asking a tutor for help?
Which target best proves readiness for "Neural networks and optimization"?
Which second target belongs to "Neural networks and optimization"?
Practice: Design a validation protocol that blocks leakage and reports at least one subgroup performance check.
Start with Principles of Data Science and 6.0002 Introduction to Computational Thinking and Data Science. Take notes until you can explain: Use train/validation/test splits, cross-validation, leakage audits, drift checks, and fairness metrics.
Do the first sequenced checks until the definitions, vocabulary, and setup are correct without hints.
Produce the assignment artifact, then pass the application and analysis checks tied to: Design a validation protocol that blocks leakage and reports at least one subgroup performance check.
Any miss becomes an error-log entry, a clean redo, and one nearby transfer problem before advancing.
Bring your attempted work, the exact missed check, and one question about: Separate model selection from final evaluation.
Move in order: foundation, transfer, analysis, timed readiness, then metacognitive repair.
Start the next sequenced check: #1 Trace check. Do not jump ahead until this one is correct.
Weakest facet: Concept
A loop for Validation, overfitting, and fairness runs once for each value 0 through 9. How many times does it run?
Which implementation is easiest to test?
Which target best matches the Machine Learning unit "Validation, overfitting, and fairness"?
Which practice artifact should you produce for "Validation, overfitting, and fairness" before asking a tutor for help?
Which target best proves readiness for "Validation, overfitting, and fairness"?
Which second target belongs to "Validation, overfitting, and fairness"?
Practice: Write a model card with intended use, data limits, metrics, fairness checks, and monitoring plan.
Start with Principles of Data Science and 6.0002 Introduction to Computational Thinking and Data Science. Take notes until you can explain: Build a small end-to-end ML project with reproducible data prep, model comparison, evaluation, and documentation.
Do the first sequenced checks until the definitions, vocabulary, and setup are correct without hints.
Produce the assignment artifact, then pass the application and analysis checks tied to: Write a model card with intended use, data limits, metrics, fairness checks, and monitoring plan.
Any miss becomes an error-log entry, a clean redo, and one nearby transfer problem before advancing.
Bring your attempted work, the exact missed check, and one question about: Plan monitoring for data drift, error shifts, and retraining triggers.
Move in order: foundation, transfer, analysis, timed readiness, then metacognitive repair.
Start the next sequenced check: #1 Trace check. Do not jump ahead until this one is correct.
Weakest facet: Concept
A loop for Capstone and deployment monitoring runs once for each value 0 through 10. How many times does it run?
Which implementation is easiest to test?
Which target best matches the Machine Learning unit "Capstone and deployment monitoring"?
Which practice artifact should you produce for "Capstone and deployment monitoring" before asking a tutor for help?
Which target best proves readiness for "Capstone and deployment monitoring"?
Which second target belongs to "Capstone and deployment monitoring"?
These are source links, not scraped course copies. Licenses differ, so the label tells students how each source is used.
View the full citation indexData science cycle, Python, statistics, modeling, visualization, and ethics.
Python-based modeling, simulation, optimization, and data analysis course.
Statistics, machine learning, neural network, and AI explainers organized from fundamentals through advanced topics.
Long-form programming, computer science, data, math, and project tutorials from the freeCodeCamp nonprofit channel.
Harvard CS50 sections and course materials for programming, algorithms, data structures, Python, SQL, and web foundations.