The Data Engineering Handbook: A Data Engineering Roadmap
The Data Engineering Handbook is a link-and-list repo maintained by DataExpert.io, pointing to books, bootcamps, communities, and vendor blogs rather than teaching data engineering itself. Reach for it when you want a map of where to learn next — a books list, community Discords, a companies directory by category. Skip it if you want structured lessons, since the repo holds almost no original content of its own.
What This Resource Offers
The Data Engineering Handbook is a GitHub repository of links for learning data engineering, not a codebase you install. It bundles a getting-started roadmap, intro pages for a 4-week beginner and 6-week intermediate boot camp, a books list, a communities list, interview advice, a companies directory by category, company blogs, whitepapers, and a list of content creators.
Who Benefits from This Handbook?
The Data Engineering Handbook suits people just starting out — the beginner and intermediate boot camp tracks are aimed at complete newcomers — as well as engineers prepping for interviews via the interviews.md notes. It also works as a page of data engineering resources for engineers who want a shortlist of vendors by category (which company handles orchestration vs. data quality) rather than a Google search. It's less useful if you already know the landscape and want deep technical tutorials, since most entries are one-line links out to books, blogs, or company sites rather than write-ups of their own.
Key Sections and Curated Content
- ✓A getting-started section built around a data engineering roadmap link, plus intro and software pages for a free 4-week beginner boot camp and a free 6-week intermediate boot camp — DataExpert.io's own data engineering bootcamps.
- ✓A books.md list of over 25 books, headlined by three named must-reads: Fundamentals of Data Engineering, Designing Data-Intensive Applications, and Designing Machine Learning Systems.
- ✓A communities.md list of over 10 communities, split into data engineering picks (DataExpert.io Community Discord, Data Talks Club Slack, Data Engineer Things Community) and ML picks (AdalFlow Discord, Chip Huyen MLOps Discord).
- ✓A companies directory grouped into categories ranging from orchestration (Airflow, Dagster, Prefect, Mage, Astronomer, Kestra) to data lake/cloud (Apache Iceberg, Delta Lake, Databricks, Onehouse), data warehousing (Snowflake, Firebolt, Databend), data quality (dbt, Great Expectations, Soda), data integration/ETL tools (Fivetran, Airbyte, dlt, Meltano), analytics/visualization, semantic layers, modern OLAP, and data lineage.
- ✓Separate projects.md and interviews.md pages for hands-on practice and interview prep, both linked from the getting-started section.
- ✓Links to data engineering blogs from Netflix, Uber, Databricks, Airbnb, AWS, Microsoft, Oracle, and Meta, plus a whitepapers list covering papers like The Google File System and MapReduce.
- ✓A social-accounts table of YouTube, LinkedIn, and X/Twitter data engineering creators, gated at a self-reported 5k-follower minimum per the README.
Strengths
- ✓The companies directory sorts vendors into specific functional categories (orchestration, data warehouse, data quality, semantic layers) instead of dumping every data-pipeline vendor into one undifferentiated list.
- ✓The books list narrows down to three named must-reads instead of leaving you to sort through all 25+ entries yourself.
- ✓The getting-started section links directly to two structured, free boot camps (4-week beginner, 6-week intermediate) rather than a vague 'learn more' pointer.
- ✓It covers both breadth (books, blogs, whitepapers, communities) and applied prep (projects.md, interviews.md) in one place.
Considerations and Scope
- △The license field on GitHub is unset ('?'), so reuse or redistribution terms for the repo's own content aren't defined.
- △Almost every entry is a bare link with no annotation beyond a title — the README doesn't explain why a given company or book made the cut versus another, so you're trusting the curator's taste sight unseen.
- △The social-accounts tables (YouTube, LinkedIn, X/Twitter) list follower counts with no stated check-date, and the README's own truncated table formatting suggests the file isn't rigorously maintained.
- △GitHub tags the repo's primary language as Jupyter Notebook, but the README shows no notebooks, code, or exercises of its own — everything routes out to boot camp pages or outside sites.
Other Data Engineering Learning Paths
Frequently Asked Questions
The Data Engineering Handbook is a public GitHub repository, free to browse, and it links to boot camps described as free (a 4-week beginner track and a 6-week intermediate track).
The Data Engineering Handbook covers a range: its beginner boot camp targets people new to data engineering, the intermediate boot camp and books list suit engineers building depth, and the interview-prep and companies directory suit working engineers already on the job.
The Data Engineering Handbook links out to a separate projects.md page described as hands-on examples, but the repo's own README is a set of links rather than exercises you complete inline.
The Data Engineering Handbook has a dedicated interviews.md page the README describes as advice on how to pass data engineering interviews, linked from the getting-started section.
Update frequency isn't clearly documented in the README; the boot camp dates and links suggest active maintenance, but there's no changelog or last-updated marker to confirm a cadence.
The Data Engineering Handbook's companies directory spans orchestration, data lakes and cloud storage, data warehousing, data quality, analytics and visualization, data integration, semantic layers, modern OLAP databases, and real-time data — organized by vendor category rather than by tutorial.
Best use cases
- •Bookmarking a starting roadmap before enrolling in one of DataExpert.io's free beginner or intermediate boot camps.
- •Picking a first data engineering book from the books.md shortlist instead of browsing Amazon blind.
- •Looking up which vendor category a tool belongs to — checking whether Dagster is orchestration or data quality — via the companies directory.
- •Prepping for interviews using the interviews.md notes before technical rounds.
- •Finding a Discord or Slack community to ask data engineering questions in, via communities.md.
