Skip to content
Abstract cover artwork for Vector Databases and Semantic Search at Scale, a Data Engineering course Data Engineering
Data Engineering intermediate 10 hours 4.8

Vector Databases and Semantic Search at Scale

Taught by Daniel Okafor

How approximate nearest-neighbour indexes actually work (HNSW, IVF, product quantisation) and what you trade for speed. Metadata filtering, hybrid dense plus sparse retrieval, reciprocal rank fusion, re-indexing strategies for embedding migrations, and keeping a vector store in sync with a relational system of record.

More in Data Engineering

Abstract cover artwork for Data Engineering with dbt and Modern SQL, a Data Engineering course Data Engineering
Data Engineering intermediate

Data Engineering with dbt and Modern SQL

Treat analytics code like software. Layered model architecture, incremental materialisations, snapshots for slowly changing dimensions, data tests as contracts, and documentation that stays current because it is generated. Includes CI that blocks a merge when a data test fails.

NH Nadia Hussain 4.8

$79.00

14 hours

Abstract cover artwork for Analytics at Speed with DuckDB, a Data Engineering course Data Engineering
Data Engineering beginner

Analytics at Speed with DuckDB

An entire analytics engine in a single process, and often faster than the cluster you were about to provision. Columnar execution, querying Parquet in place, larger-than-memory workloads, and where DuckDB stops being the right answer.

EV Elena Vasquez 4.7

$49.00

8 hours

Abstract cover artwork for Streaming Data with Kafka and Flink, a Data Engineering course Data Engineering
Data Engineering advanced

Streaming Data with Kafka and Flink

Event-time versus processing-time, watermarks, windowing, and what exactly-once actually guarantees. You will build a stateful stream processor with checkpointing, handle out-of-order and late events correctly, and reason about the throughput and latency trade-offs under real partitioning.

KM Kwame Mensah 4.6

$129.00

19 hours

Abstract cover artwork for Apache Airflow: Orchestration You Can Debug at 3am, a Data Engineering course Data Engineering
Data Engineering intermediate

Apache Airflow: Orchestration You Can Debug at 3am

Idempotent tasks, correct backfills, sensible retry semantics and the execution-date confusion that causes most Airflow incidents. You will build a pipeline that is safe to re-run, alerts on the right failures, and does not silently skip a day when a DAG is paused.

NH Nadia Hussain 4.5

$69.00

12 hours