Data Science & Engineering

SQL and NoSQL, data modelling, pipelines and orchestration, big-data engines, statistics, analytics, and the Python and R toolchains.

0 of 18 chapters passed

Pass mark is 75% on the chapter quiz.

Start track

1 · Foundations

How data systems evolved, and the relational model that still anchors them.

  1. 1. Genesis of Data SystemsFiles to relational databases to the lakehouse, and the roles in a modern data team.8 min · 3 MCQs

2 · SQL & databases

Querying, indexing, transactions, NoSQL families and warehouses.

  1. 2. SQL FundamentalsSELECT mechanics, joins, aggregation, logical execution order and NULL semantics.11 min · 3 MCQs
  2. 3. Advanced SQL: Windows, CTEs and Query PlansWindow functions, recursive CTEs, indexing and reading EXPLAIN output.11 min · 3 MCQs
  3. 4. Transactions, Consistency and DistributionACID, isolation levels, MVCC, replication and the CAP trade-off.10 min · 3 MCQs
  4. 5. NoSQL Families and When to Use ThemKey-value, document, wide-column, graph and search engines — and polyglot persistence.10 min · 3 MCQs

3 · Data modelling

Normalisation, dimensional models, lakehouse layers and quality contracts.

  1. 6. Relational Modelling and NormalisationKeys, normal forms, constraints and controlled denormalisation.9 min · 3 MCQs
  2. 7. Dimensional Modelling for AnalyticsFacts, dimensions, star schemas, grain and slowly changing dimensions.10 min · 3 MCQs
  3. 8. Lakehouse Layers, Governance and Data QualityBronze/silver/gold, table formats, catalogues, contracts and tests.10 min · 3 MCQs

4 · Data engineering

Ingestion, ELT, orchestration, streaming and distributed processing.

  1. 9. Ingestion, ETL and ELTBatch and incremental loading, CDC, idempotency and backfills.10 min · 3 MCQs
  2. 10. Orchestration and Pipeline ReliabilityDAGs, scheduling, dependencies, retries, SLAs and observability.9 min · 3 MCQs
  3. 11. Big Data Engines and Distributed ProcessingMapReduce to Spark, partitions, shuffles, skew and query engines.11 min · 3 MCQs
  4. 12. Streaming Data and Real-Time AnalyticsKafka, event time, windows, watermarks and exactly-once semantics.10 min · 3 MCQs

5 · Analytics & data science

Statistics, Python, R, visualisation, experimentation and delivery.

  1. 13. Statistics for PractitionersDistributions, sampling, confidence intervals, hypothesis testing and common traps.10 min · 3 MCQs
  2. 14. Python for Data WorkNumPy, pandas, vectorisation, Polars and reproducible environments.11 min · 3 MCQs
  3. 15. R for Statistics and AnalysisThe tidyverse, data frames, statistical modelling and R Markdown reporting.9 min · 3 MCQs
  4. 16. Visualisation and Analytical StorytellingChart choice, perceptual accuracy, dashboards and communicating uncertainty.8 min · 3 MCQs
  5. 17. Experimentation and Causal InferenceA/B testing, power, peeking, and quasi-experimental methods.10 min · 3 MCQs
  6. 18. Delivering Data Products in ProductionFeature stores, semantic layers, serving, governance and a readiness checklist.10 min · 3 MCQs

Other tracks