File format is a recurring DEA-C01 topic because it directly affects query cost and speed. The short version: columnar formats beat row formats for analytics.
Columnar vs row formats
CSV and JSON are row-based and human-readable but slow and expensive to query at scale — engines read every column. Parquet and ORC are columnar and compressed, so a query reads only the columns it needs, scanning far less data. For analytics on S3 with Athena or Spectrum, converting to Parquet is the standard cost-cutting move.
Test yourself
An analytics team queries only a few columns from huge S3 datasets and wants to minimize data scanned and cost. Which format should they use?
- CSV
- JSON
- Parquet
- Plain text
👉 Click to reveal the answer & explanation
Correct answer: C. Parquet is columnar and compressed, so queries read only the needed columns and scan far less data — cutting cost and time. CSV (A), JSON (B), and plain text (D) are row-based and force reading whole rows.
Related topics
Data partitioning · Amazon Athena · Amazon Redshift
Ready to pass the AWS Data Engineer Associate (DEA-C01)?
Stop guessing whether you’re ready. Our full-length, exam-realistic practice exams put you through the exact question style you’ll face — with a detailed explanation behind every answer, so you learn why, not just what.
- ✓ 6 full-length practice exams
- ✓ A detailed explanation for every single question
- ✓ Realistic, scenario-based questions — not memory dumps
- ✓ Lifetime access, kept current for 2026
Get the DEA-C01 Practice Exams →or try 25 free questions first