AWS for data engineers cheat sheet
S3, Glue, Athena, Kinesis, and what to skip on day one.
Storage and query
s3://bucket/bronze/dt=2026-08-01/- Hive-style prefixes so Athena/Glue can prune.
CREATE TABLE t ... PARTITIONED BY (dt) LOCATION s3://...- Glue catalog + Athena SQL on the lake.
MSCK REPAIR TABLE / Glue crawler- Register new partitions; prefer explicit ADD PARTITION in prod.
Ingest and transform
Kinesis Data Streams / Firehose → S3- Streams when you need sub-minute; Firehose for managed batching.
Glue job (Spark) or Glue Python shell- Serverless Spark for batch; skip EMR until you outgrow Glue.
Lambda for <15 min transforms- Great for small files; terrible as a warehouse.
Orchestration
MWAA or Step Functions- MWAA if the team already writes Airflow; Step Functions for AWS-native graphs.
From DataLane — tutorials at/blog, practice SQL live in theplayground.