AWS Certified Data Engineer – Associate (DEA-C01) Cheat Sheet
A practical DEA-C01 study reference covering the AWS data engineering services, comparisons, and exam traps that come up most — ingestion, transformation, storage, analytics, and governance in one place.
Free Study Resource · DEA-C01 Exam Prep · 25-Question Practice ExamWhat This Cheat Sheet Covers
DEA-C01 tests whether you can build, secure, and troubleshoot real data pipelines on AWS. Domain 1 (Data Ingestion and Transformation) carries the most weight at 34%, and Domains 1 and 2 together account for 60% of the scored content, so that's where this page puts the most detail.
High-Yield DEA-C01 Concept Comparisons
DEA-C01 scenarios almost always come down to picking the right service for a stated constraint. These cards match that format: the cue in the question, and the service that fits.
Medallion Architecture — Bronze, Silver, Gold
See: raw, unprocessed data stored as-is, typically in Amazon S3 → Bronze layer. See: validated, deduplicated, converted to Parquet/ORC → Silver layer. See: aggregated, query-optimized, partitioned → Gold layer (Redshift or Athena on Iceberg).
Zero-ETL Integrations
See: near-real-time replication from supported operational databases into Amazon Redshift with minimal pipeline management → Zero-ETL integration. Note: Complex multi-source transformations or heavy business logic still require AWS Glue ETL pipelines.
Kinesis Data Streams vs Amazon Data Firehose vs MSK
See: custom consumers, sub-second latency, manual shard management → Kinesis Data Streams. See: fully managed, auto-scaling delivery to S3, Redshift, or OpenSearch with optional format conversion → Amazon Data Firehose. See: open-source Apache Kafka compatibility → Amazon MSK.
AWS Lambda — Execution Limits, Memory & Cold Starts
See: event-driven processing with max execution time → 15 minutes (hard limit). See: memory allocation → 128 MB to 10 GB, with CPU scaling proportionally. See: mitigating cold starts for latency-sensitive workloads → Provisioned Concurrency.
Step Functions — Standard vs Express
See: long-running, durable workflows requiring auditability and execution history → Standard Workflows. See: high-volume, short-duration event-driven data processing → Express Workflows.
EBS Volume Persistence & Instance Store
See: persistent block storage for EC2 → EBS. See: ultra-high IOPS ephemeral scratch space → Instance Store (data is lost permanently when the instance stops or terminates).
EBS Volume Types — gp3 vs io2 vs st1
See: primary cost-effective SSD choice with independent IOPS/throughput configuration → gp3. See: mission-critical databases requiring highest IOPS and durability → io2. See: throughput-intensive sequential workloads → st1.
S3 Storage Classes & Glacier Tiers
See: frequent access pipeline data → S3 Standard. See: unpredictable or changing access patterns → S3 Intelligent-Tiering. See: infrequent access with rapid retrieval needs → S3 Standard-IA. See: long-term compliance archives at lowest cost → Glacier Deep Archive.
RDS Multi-AZ vs Read Replicas
See: synchronous replication for high availability and automatic failover → Multi-AZ deployments. See: asynchronous replication to scale read-heavy workloads → Read Replicas.
DynamoDB Core Features
See: microsecond latency read acceleration → DAX. See: automatic deletion of expired items → TTL. See: continuous item-level change capture → DynamoDB Streams. See: multi-region active-active tables → Global Tables.
DynamoDB GSI vs LSI
See: index that can be created or modified anytime after table creation → Global Secondary Index (GSI). See: index that must be defined exclusively at table creation time → Local Secondary Index (LSI).
Redshift Distribution Styles
See: automatic distribution choice → AUTO. See: co-locating matching join keys → KEY. See: small dimension tables replicated across all compute nodes → ALL. See: uniform round-robin distribution → EVEN.
AWS Glue — Crawlers, Catalog, Bookmarks & Streaming
See: automatic schema discovery → Glue Crawlers updating the Data Catalog. See: tracking processed data to prevent duplicate runs → Job Bookmarks. See: ETL capabilities → Glue supports batch ETL as well as streaming ETL workloads.
Athena — Serverless SQL & Storage Formats
See: ad hoc SQL queries directly on S3 → Amazon Athena. Why: Converting raw files to columnar formats (Parquet/ORC) and partitioning significantly reduces data scanned and query costs.
Lake Formation vs IAM for Data Lakes
See: broad AWS service and resource-level authorization → IAM policies. See: fine-grained table, column, and row-level access control on S3 data lakes → AWS Lake Formation.
Security Groups vs Network ACLs
See: instance-level stateful firewall → Security Groups. See: subnet-level stateless firewall → Network ACLs.
VPC Peering vs Transit Gateway
See: point-to-point connection between two VPCs → VPC Peering. See: central hub connecting dozens or hundreds of VPCs → AWS Transit Gateway.
Gateway Endpoints vs Interface Endpoints
See: private access to Amazon S3 or DynamoDB without requiring a NAT Gateway → Gateway Endpoints. See: private connection powered by AWS PrivateLink for other AWS services → Interface Endpoints.
IAM Policy Evaluation & SCPs
See: explicit Deny in any applicable policy → overrides all Allows unconditionally. See: AWS Organizations guardrails → Service Control Policies (SCPs) define maximum allowed permissions but do not grant permissions directly.
Data Migration Service Matrix
See: moving existing databases with minimal downtime → AWS DMS. See: fast online file transfer to S3/EFS → AWS DataSync. See: offline device-based petabyte-scale data transfer → AWS Snowball Edge.
DEA-C01 Exam Traps
S3 provides strong read-after-write consistency for PUTs and DELETEs — "eventual consistency" is no longer accurate.
Traditional Multi-AZ standby instances do not serve read traffic; use Read Replicas for read scaling.
Instance Store data does not survive an instance stop or termination; use EBS volumes for persistent storage.
Kinesis Data Streams requires manual shard management; Amazon Data Firehose auto-scales delivery.
Lambda has a strict 15-minute maximum execution timeout; long-running batch processing belongs on Glue, EMR, or ECS.
A DynamoDB LSI can only be defined when the table is created; GSIs can be added or removed anytime.
Quick Reference — Ingestion, Processing, Storage & Analytics
Data Ingestion
- Kinesis Data Streams: custom real-time consumers, manual shard scaling
- Amazon Data Firehose: managed delivery, auto-scales
- Amazon MSK: Apache Kafka-compatible
- AWS DMS: database migration and CDC
Data Processing
- AWS Glue: serverless batch/streaming ETL
- Amazon EMR: big data Spark/Hadoop clusters
- AWS Lambda: event-driven processing, 15m timeout
- Step Functions: orchestrates workflows
Storage & Analytics
- Amazon S3: primary object storage
- Amazon Redshift: analytical data warehouse
- Amazon Athena: serverless SQL on S3
- Amazon DynamoDB: low-latency NoSQL
How to Use This Cheat Sheet
Review core concepts, focusing on Domain 1 and Domain 2.
Memorize key service distinctions.
Review exam traps to spot distractors quickly.
Utilize quick-reference tables for summary review.
Take the free practice exam to assess readiness.
Ready to Test Your DEA-C01 Knowledge?
Test yourself with CloudExamPro's free 25-question DEA-C01 practice exam.
Frequently Asked Questions
What is the AWS Certified Data Engineer – Associate exam?
It validates expertise in building, securing, and operating data pipelines on AWS.
How many questions are on the exam?
65 questions, 130 minutes, 720 passing score.