Role roadmaps / Data Engineer Role flowchart
Move data that other people can trust: ingest, transform, orchestrate, and keep the warehouse bill from becoming a surprise.
Also called: Analytics Platform Engineer, Pipeline Engineer
You ship: Pipelines with tests, an on-call story, and a warehouse someone else can query.
4–6 months to ship independently
Roadmap Coding Interview Question bank Switch roleData Analyst BI Analyst Data Engineer Analytics Engineer Data Scientist ML Engineer MLOps Engineer AI Engineer Data Architect Nearby roles:Analytics Engineer , Data Architect , MLOps Engineer
Checked skills
0% · click a node for concepts, stacks, access, and what to produce · drag to pan · Ctrl-scroll to zoom
− Fit + Clear checks
Foundations Daily craft How you ship Career Done
Start · Data Engineer Foundations SQL as a production language Windows, anti-joins, late data. Not just SELECT for a dashboard. Python for jobs, not notebooks Packaging, retries, and memory. Pandas is a tool, not an identity. Warehouse modeling Medallion as a communication tool. Grain, SCD, and why gold is expensive. Daily craft Ingest without folklore Batch vs CDC, schema drift, and idempotent loads. Orchestration Airflow, Dagster, or Prefect — one clock, data-aware if you can. One warehouse, cost first Snowflake, BigQuery, or Databricks. Query Profile before Cortex. How you ship Tests that page someone Row counts, uniqueness, accepted values. Fail the job on grain. Streaming when batch is wrong Kafka, watermarks, and exactly-once as a contract, not a slogan. Open tables Iceberg or Delta, time travel, and compaction as an ops job. Career moves The DE interview SQL, Python, a platform sketch, and an incident story. Certs as vocabulary SnowPro, Databricks, AWS — useful after you have shipped, not instead of. Next role Analytics engineer toward metrics. Architect toward platform. ML data toward features. Keep practicing Study this node
SQL as a production language Windows, anti-joins, late data. Not just SELECT for a dashboard.
DE interviews still open with SQL. The pad is not optional.
SQL as a production language: windows, anti-joins, late data — not SELECT for a dashboard.
Time to learn 4–8 weeks of production SQL, not dashboard SQL.
Prerequisites Windows, anti-joins, late data as vocabulary A pad you will actually open What you need All SQL tracks on this site A warehouse (Snowflake / BigQuery / Databricks / Redshift — one is enough) dbt or a reviewed SQL repo Query Profile / EXPLAIN Concepts to study Window frames (ROWS vs RANGE) and running totals Anti-joins and reconciliation (orders vs payments) Idempotent loads: rerunning must not double revenue Late data and watermarks as a warehouse problem too QUALIFY / ROW_NUMBER for latest-per-key JSON/semi-structured in the warehouse Tech stacks Snowflake / BigQuery / Databricks SQL / Redshift — Production dialectdbt — How SQL gets reviewedThis site’s SQL tracks — The padHow to study Finish Warehouse SQL Core, then Windows, then Quality & Hard SQL. Rewrite one dashboard query as a tested dbt model. Practice talking through EXPLAIN / Query Profile. What to produce A CDC-into-gold sketch with where you dedup and what you page on A window query you can write from memory Practice on this site All SQL tracks; coding sandbox; DE interview SQL Pitfalls Notebook SQL that nobody can rerun Windows used as GROUP BY because the tutorial said so Ignoring late-arriving facts Interview prompts Design CDC into gold. Where do you dedup, and what do you page on? Links
Mark as done All Articles Stacks Practice Learn Pages
No results. Try analyst, numpy, snowflake, or roadmap.
↑↓ navigate↵ openesc closeAliases: sf, numpy, analyst, mlops