AWS Certified Data Engineer – Associate (DEA-C01) Cheat Sheet

A practical DEA-C01 study reference covering the AWS data engineering services, comparisons, and exam traps that come up most — ingestion, transformation, storage, analytics, and governance in one place.

Free Study Resource · DEA-C01 Exam Prep · 25-Question Practice Exam
65Questions
130 minDuration
720/1000Passing Score
$150Exam Cost
2–3 yrsRecommended Exp.
25 QsFree Practice Exam
Exam Domains

What This Cheat Sheet Covers

DEA-C01 tests whether you can build, secure, and troubleshoot real data pipelines on AWS. Domain 1 (Data Ingestion and Transformation) carries the most weight at 34%, and Domains 1 and 2 together account for 60% of the scored content, so that's where this page puts the most detail.

34%Data Ingestion & Transformation
26%Data Store Management
22%Data Operations & Support
18%Data Security & Governance
Core Reference

High-Yield DEA-C01 Concept Comparisons

DEA-C01 scenarios almost always come down to picking the right service for a stated constraint. These cards match that format: the cue in the question, and the service that fits.

Architecture

Medallion Architecture — Bronze, Silver, Gold

See: raw, unprocessed data stored as-is, typically in Amazon S3 → Bronze layer. See: validated, deduplicated, converted to Parquet/ORC → Silver layer. See: aggregated, query-optimized, partitioned → Gold layer (Redshift or Athena on Iceberg).

Architecture

Zero-ETL Integrations

See: near-real-time replication from supported operational databases into Amazon Redshift with minimal pipeline management → Zero-ETL integration. Note: Complex multi-source transformations or heavy business logic still require AWS Glue ETL pipelines.

Streaming

Kinesis Data Streams vs Amazon Data Firehose vs MSK

See: custom consumers, sub-second latency, manual shard management → Kinesis Data Streams. See: fully managed, auto-scaling delivery to S3, Redshift, or OpenSearch with optional format conversion → Amazon Data Firehose. See: open-source Apache Kafka compatibility → Amazon MSK.

Compute

AWS Lambda — Execution Limits, Memory & Cold Starts

See: event-driven processing with max execution time → 15 minutes (hard limit). See: memory allocation → 128 MB to 10 GB, with CPU scaling proportionally. See: mitigating cold starts for latency-sensitive workloads → Provisioned Concurrency.

Orchestration

Step Functions — Standard vs Express

See: long-running, durable workflows requiring auditability and execution history → Standard Workflows. See: high-volume, short-duration event-driven data processing → Express Workflows.

Storage

EBS Volume Persistence & Instance Store

See: persistent block storage for EC2 → EBS. See: ultra-high IOPS ephemeral scratch space → Instance Store (data is lost permanently when the instance stops or terminates).

Storage

EBS Volume Types — gp3 vs io2 vs st1

See: primary cost-effective SSD choice with independent IOPS/throughput configuration → gp3. See: mission-critical databases requiring highest IOPS and durability → io2. See: throughput-intensive sequential workloads → st1.

Storage

S3 Storage Classes & Glacier Tiers

See: frequent access pipeline data → S3 Standard. See: unpredictable or changing access patterns → S3 Intelligent-Tiering. See: infrequent access with rapid retrieval needs → S3 Standard-IA. See: long-term compliance archives at lowest cost → Glacier Deep Archive.

Database

RDS Multi-AZ vs Read Replicas

See: synchronous replication for high availability and automatic failover → Multi-AZ deployments. See: asynchronous replication to scale read-heavy workloads → Read Replicas.

Database

DynamoDB Core Features

See: microsecond latency read acceleration → DAX. See: automatic deletion of expired items → TTL. See: continuous item-level change capture → DynamoDB Streams. See: multi-region active-active tables → Global Tables.

Database

DynamoDB GSI vs LSI

See: index that can be created or modified anytime after table creation → Global Secondary Index (GSI). See: index that must be defined exclusively at table creation time → Local Secondary Index (LSI).

Analytics

Redshift Distribution Styles

See: automatic distribution choice → AUTO. See: co-locating matching join keys → KEY. See: small dimension tables replicated across all compute nodes → ALL. See: uniform round-robin distribution → EVEN.

Analytics

AWS Glue — Crawlers, Catalog, Bookmarks & Streaming

See: automatic schema discovery → Glue Crawlers updating the Data Catalog. See: tracking processed data to prevent duplicate runs → Job Bookmarks. See: ETL capabilities → Glue supports batch ETL as well as streaming ETL workloads.

Analytics

Athena — Serverless SQL & Storage Formats

See: ad hoc SQL queries directly on S3 → Amazon Athena. Why: Converting raw files to columnar formats (Parquet/ORC) and partitioning significantly reduces data scanned and query costs.

Governance

Lake Formation vs IAM for Data Lakes

See: broad AWS service and resource-level authorization → IAM policies. See: fine-grained table, column, and row-level access control on S3 data lakes → AWS Lake Formation.

Networking

Security Groups vs Network ACLs

See: instance-level stateful firewall → Security Groups. See: subnet-level stateless firewall → Network ACLs.

Networking

VPC Peering vs Transit Gateway

See: point-to-point connection between two VPCs → VPC Peering. See: central hub connecting dozens or hundreds of VPCs → AWS Transit Gateway.

Networking

Gateway Endpoints vs Interface Endpoints

See: private access to Amazon S3 or DynamoDB without requiring a NAT Gateway → Gateway Endpoints. See: private connection powered by AWS PrivateLink for other AWS services → Interface Endpoints.

Security

IAM Policy Evaluation & SCPs

See: explicit Deny in any applicable policy → overrides all Allows unconditionally. See: AWS Organizations guardrails → Service Control Policies (SCPs) define maximum allowed permissions but do not grant permissions directly.

Migration

Data Migration Service Matrix

See: moving existing databases with minimal downtime → AWS DMS. See: fast online file transfer to S3/EFS → AWS DataSync. See: offline device-based petabyte-scale data transfer → AWS Snowball Edge.

Watch For These

DEA-C01 Exam Traps

✕

S3 provides strong read-after-write consistency for PUTs and DELETEs — "eventual consistency" is no longer accurate.

✕

Traditional Multi-AZ standby instances do not serve read traffic; use Read Replicas for read scaling.

✕

Instance Store data does not survive an instance stop or termination; use EBS volumes for persistent storage.

✕

Kinesis Data Streams requires manual shard management; Amazon Data Firehose auto-scales delivery.

✕

Lambda has a strict 15-minute maximum execution timeout; long-running batch processing belongs on Glue, EMR, or ECS.

✕

A DynamoDB LSI can only be defined when the table is created; GSIs can be added or removed anytime.

Memorize This

Quick Reference — Ingestion, Processing, Storage & Analytics

Data Ingestion

  • Kinesis Data Streams: custom real-time consumers, manual shard scaling
  • Amazon Data Firehose: managed delivery, auto-scales
  • Amazon MSK: Apache Kafka-compatible
  • AWS DMS: database migration and CDC

Data Processing

  • AWS Glue: serverless batch/streaming ETL
  • Amazon EMR: big data Spark/Hadoop clusters
  • AWS Lambda: event-driven processing, 15m timeout
  • Step Functions: orchestrates workflows

Storage & Analytics

  • Amazon S3: primary object storage
  • Amazon Redshift: analytical data warehouse
  • Amazon Athena: serverless SQL on S3
  • Amazon DynamoDB: low-latency NoSQL
Exam Strategy

How to Use This Cheat Sheet

Review core concepts, focusing on Domain 1 and Domain 2.

Memorize key service distinctions.

Review exam traps to spot distractors quickly.

Utilize quick-reference tables for summary review.

Take the free practice exam to assess readiness.

Ready to Test Your DEA-C01 Knowledge?

Test yourself with CloudExamPro's free 25-question DEA-C01 practice exam.

Download the DEA-C01 Cheat Sheet

Get the PDF version for offline revision.

FAQ

Frequently Asked Questions

What is the AWS Certified Data Engineer – Associate exam?

It validates expertise in building, securing, and operating data pipelines on AWS.

How many questions are on the exam?

65 questions, 130 minutes, 720 passing score.

Limited-time offer · ends in --Days:--Hrs:--Min:--Sec
Get Instant Access