Skip to content
All courses

Course 03

Everything Data

The modern data stack for AI and ML - SQL and batch processing to streaming, vector databases, feature stores, and data pipelines for LLMs.

For ML engineers who need to build reliable data pipelines, feature stores, and training data workflows at scale.

12
modules
85
lessons
90
curated videos
26h 36m
of video
01

The Modern Data Stack

Understand how modern data platforms are organized for analytics, machine learning, real-time systems, and AI applications.

7 lessons · 9 videos · Free
02

SQL for ML Engineers

Use SQL to create reliable datasets, features, labels, and diagnostics for machine learning workflows.

7 lessons · 8 videos
03

Data Modeling

Design data models that support analytics, feature engineering, training, and serving without leakage or excessive complexity.

7 lessons · 7 videos
04

Batch Processing at Scale

Build scalable batch data processing jobs for large training datasets and offline feature computation.

7 lessons · 3 videos
05

Streaming & Real-Time

Design and implement streaming data systems for low-latency features, real-time analytics, and event-driven ML applications.

7 lessons · 8 videos
06

Warehouses & Lakehouses

Use modern warehouses and lakehouse technologies to store, query, optimize, and govern data for AI systems.

7 lessons · 9 videos
07

Orchestration

Build reliable, observable, and maintainable data pipelines using modern workflow orchestration patterns.

7 lessons · 9 videos
08

Data Quality & Monitoring

Validate, monitor, and debug data and ML pipelines before bad data reaches models or users.

7 lessons · 4 videos
09

Vector Databases & RAG Pipelines

Build retrieval systems that transform documents into embeddings, store them in vector databases, and serve them to LLM applications.

7 lessons · 8 videos
10

Data Versioning & ML Reproducibility

Make datasets, features, experiments, and model outputs reproducible across time, teams, and environments.

7 lessons · 6 videos
11

Feature Engineering & Stores

Design, compute, store, serve, and monitor features for offline training and online inference.

7 lessons · 4 videos
12

Data for LLMs & Foundation Models

Build data pipelines for pretraining, fine-tuning, evaluation, synthetic data generation, and human feedback loops for foundation models.

8 lessons · 15 videos

Curated from 26 channels

Every video is a public YouTube video. We pick the single clearest explanation for each lesson and credit the channel that made it.