File format is a recurring DEA-C01 topic because it directly affects query cost and speed. The short version: columnar formats beat row formats for analytics.

Columnar vs row formats

CSV and JSON are row-based and human-readable but slow and expensive to query at scale — engines read every column. Parquet and ORC are columnar and compressed, so a query reads only the columns it needs, scanning far less data. For analytics on S3 with Athena or Spectrum, converting to Parquet is the standard cost-cutting move.

Test yourself

Practice question

An analytics team queries only a few columns from huge S3 datasets and wants to minimize data scanned and cost. Which format should they use?

  1. CSV
  2. JSON
  3. Parquet
  4. Plain text
👉 Click to reveal the answer & explanation

Correct answer: C. Parquet is columnar and compressed, so queries read only the needed columns and scan far less data — cutting cost and time. CSV (A), JSON (B), and plain text (D) are row-based and force reading whole rows.

Related topics

Data partitioning · Amazon Athena · Amazon Redshift

CloudExamPro Premium

Ready to pass the AWS Data Engineer Associate (DEA-C01)?

Stop guessing whether you’re ready. Our full-length, exam-realistic practice exams put you through the exact question style you’ll face — with a detailed explanation behind every answer, so you learn why, not just what.

  • ✓  6 full-length practice exams
  • ✓  A detailed explanation for every single question
  • ✓  Realistic, scenario-based questions — not memory dumps
  • ✓  Lifetime access, kept current for 2026

Get the DEA-C01 Practice Exams →or try 25 free questions first

Limited-time offer · ends in --Days:--Hrs:--Min:--Sec
Get Instant Access