Everything Data - full curriculum
The modern data stack for AI and ML - SQL and batch processing to streaming, vector databases, feature stores, and data pipelines for LLMs.
- 01
The Modern Data Stack
Understand how modern data platforms are organized for analytics, machine learning, real-time systems, and AI applications.
- 01.01From Data Warehouse to AI PlatformVideo
- 01.02Data Mesh and Domain Ownership2 videos
- 01.03Lakehouse Architecture OverviewVideo
- 01.04Tooling Landscape for ML EngineersCurating
- 01.05Data Platform Reference Architecture3 videos
- 01.06Batch, Stream, and Serving BoundariesVideo
- 01.07Data Contracts and GovernanceVideo
- 02
SQL for ML Engineers
Use SQL to create reliable datasets, features, labels, and diagnostics for machine learning workflows.
- 03
Data Modeling
Design data models that support analytics, feature engineering, training, and serving without leakage or excessive complexity.
- 04
Batch Processing at Scale
Build scalable batch data processing jobs for large training datasets and offline feature computation.
- 05
Streaming & Real-Time
Design and implement streaming data systems for low-latency features, real-time analytics, and event-driven ML applications.
- 06
Warehouses & Lakehouses
Use modern warehouses and lakehouse technologies to store, query, optimize, and govern data for AI systems.
- 06.01Warehouse vs Lakehouse TradeoffsVideo
- 06.02Snowflake for ML DataVideo
- 06.03BigQuery for Large-Scale AnalyticsCurating
- 06.04Databricks and Unified Analytics3 videos
- 06.05Parquet and Columnar StorageVideo
- 06.06Iceberg, Delta Lake, and Table FormatsCurating
- 06.07Catalogs, Lineage, and Access Control3 videos
- 07
Orchestration
Build reliable, observable, and maintainable data pipelines using modern workflow orchestration patterns.
- 08
Data Quality & Monitoring
Validate, monitor, and debug data and ML pipelines before bad data reaches models or users.
- 09
Vector Databases & RAG Pipelines
Build retrieval systems that transform documents into embeddings, store them in vector databases, and serve them to LLM applications.
- 10
Data Versioning & ML Reproducibility
Make datasets, features, experiments, and model outputs reproducible across time, teams, and environments.
- 11
Feature Engineering & Stores
Design, compute, store, serve, and monitor features for offline training and online inference.
- 11.01Feature Engineering LifecycleVideo
- 11.02Offline vs Online FeaturesCurating
- 11.03Feast Feature Store WalkthroughCurating
- 11.04Tecton and Managed Feature PlatformsVideo
- 11.05Point-in-Time Feature RetrievalCurating
- 11.06Online Serving and Low-Latency AccessVideo
- 11.07Feature Reuse, Discovery, and GovernanceVideo
- 12
Data for LLMs & Foundation Models
Build data pipelines for pretraining, fine-tuning, evaluation, synthetic data generation, and human feedback loops for foundation models.
- 12.01The LLM Data LifecycleCurating
- 12.02Data Curation for Foundation Models3 videos
- 12.03Deduplication and Contamination ControlCurating
- 12.04Tokenization and Dataset Packing2 videos
- 12.05Instruction Tuning Datasets2 videos
- 12.06Synthetic Data Generation Pipelines3 videos
- 12.07RLHF and Preference Data Pipelines2 videos
- 12.08LLM Evaluation Data Management3 videos