Partitioning organizes data in S3 into folders by a key — often date, like year/month/day — so query engines scan only the relevant slices instead of everything. It’s one of the highest-impact cost and performance techniques on the exam.
Why it saves money and time
Athena and Redshift Spectrum charge by data scanned. If logs are partitioned by date, a query for one day reads only that partition instead of the whole dataset — often a huge cost and speed win. When a question asks how to cut Athena costs or speed up queries on large S3 data, partitioning (plus columnar formats) is the answer.
Test yourself
Athena queries on a large S3 log dataset are slow and expensive because each query scans everything, even when filtering by date. What’s the best fix?
- Store the data as one large CSV file
- Partition the data in S3 by date
- Increase the Athena query timeout
- Copy the data to DynamoDB
👉 Click to reveal the answer & explanation
Correct answer: B. Partitioning by date lets Athena scan only the relevant partitions, cutting data scanned, cost, and runtime. One big CSV (A) forces full scans; a longer timeout (C) doesn’t reduce cost; DynamoDB (D) isn’t for analytical scans.
Related topics
Parquet vs ORC vs CSV · Amazon Athena · AWS Glue
Ready to pass the AWS Data Engineer Associate (DEA-C01)?
Stop guessing whether you’re ready. Our full-length, exam-realistic practice exams put you through the exact question style you’ll face — with a detailed explanation behind every answer, so you learn why, not just what.
- ✓ 6 full-length practice exams
- ✓ A detailed explanation for every single question
- ✓ Realistic, scenario-based questions — not memory dumps
- ✓ Lifetime access, kept current for 2026
Get the DEA-C01 Practice Exams →or try 25 free questions first