Parquet: An open source columnar file format for big data analytics
Apache Parquet is an open source, column-oriented data file format designed for efficient data storage and retrieval. It provides efficient data compression and encoding schemes with enhanced performance to handle complex data in bulk. Apache Parquet is designed to be a common interchange format for both batch and interactive workloads. It is similar to other columnar-storage file formats available in Hadoop, namely RCFile and ORC.
Apache Parquet is implemented using the record-shredding and assembly algorithm, which accommodates the complex data structures that can be used to store the data. Parquet is optimized to work with complex data in bulk and features different ways for efficient data compression and encoding types. This approach is best especially for those queries that need to read certain columns from a large table. Parquet can only read the needed columns therefore greatly minimizing the IO.
CSV is a simple and common format that is used by many tools such as Excel, Google Sheets, and numerous others. Even though the CSV files are the default format for data processing pipelines it has some disadvantages:
Parquet has helped its users reduce storage requirements by at least one-third on large datasets, in addition, it greatly improved scan and deserialization time, hence the overall costs. The following table compares the savings as well as the speedup obtained by converting data into Parquet from CSV.
| Dataset | Size on Amazon S3 | Query Run Time | Data Scanned | Cost |
| Data stored as CSV files | 1 TB | 236 seconds | 1.15 TB | $5.75 |
| Data stored in Apache Parquet Format | 130 GB | 6.78 seconds | 2.51 GB | $0.01 |
| Savings | 87% less when using Parquet | 34x faster | 99% less data scanned | 99.7% savings |
The open source Delta Lake project builds upon and extends the Parquet format, adding additional functionality like ACID transactions on cloud object storage, time travel, schema evolution, and simple DML commands (CREATE/UPDATE/INSERT/DELETE/MERGE). Delta Lake implements many of these important features through the use of an ordered transaction log that makes data warehousing functionality possible on cloud object storage. Learn more in the Databricks blog post Diving into Delta Lake: Unpacking the Transaction Log.
Parquet integrates well with interactive and serverless technologies such as AWS Athena, Amazon Redshift Spectrum, Google BigQuery, and Google Dataproc. These engines take advantage of Parquet's columnar layout to scan only the columns a query needs, which reduces the amount of data read and lowers compute costs. This selective-read behavior fits the large-batch, analytical querying these platforms are built for.
Yes, Parquet is designed to store big data of any kind, including structured data tables, images, videos, and documents. This flexibility comes from its support for complex, nested data structures rather than a rigid, table-only schema.
Parquet is similar to other columnar-storage file formats used in the Hadoop ecosystem, specifically RCFile and ORC. Like Parquet, both formats organize data by column rather than by row so that analytics queries can skip over irrelevant data more quickly.
In one comparison, converting a 1 TB CSV dataset to Parquet shrank it to 130 GB, cut query run time from 236 seconds to 6.78 seconds, and reduced the data scanned per query from 1.15 TB to 2.51 GB. Altogether, that amounted to roughly 87% less storage, a 34x faster query time, and about 99.7% savings in query cost.
Parquet uses what's called the record-shredding and assembly algorithm to accommodate complex data structures. This approach lets Parquet store nested data efficiently while still allowing queries to read only the specific columns they need, which keeps I/O to a minimum.
Subscribe to our blog and get the latest posts delivered to your inbox.