Skip to main content

What is Parquet?

Parquet: An open source columnar file format for big data analytics

by Databricks Staff

  • Parquet is an open source columnar file format designed for efficient storage and retrieval of big data so analytics engines can scan less data and run queries faster than with row based formats like CSV.
  • By storing data column by column and supporting powerful compression and encoding, Parquet reduces cloud storage footprints and the amount of data scanned, often delivering major cost and performance gains on large datasets.
  • Parquet underpins many modern data lakes and lakehouse architectures and is extended by Delta Lake to add capabilities like ACID transactions, time travel and schema evolution on cloud object storage.

What is Parquet?

Apache Parquet is an open source, column-oriented data file format designed for efficient data storage and retrieval. It provides efficient data compression and encoding schemes with enhanced performance to handle complex data in bulk. Apache Parquet is designed to be a common interchange format for both batch and interactive workloads. It is similar to other columnar-storage file formats available in Hadoop, namely RCFile and ORC.

What are the characteristics of Parquet?

  • Free and open source file format.
  • Language agnostic.
  • Column-based format - files are organized by column, rather than by row, which saves storage space and speeds up analytics queries.
  • Used for analytics (OLAP) use cases, typically in conjunction with traditional OLTP databases.
  • Highly efficient data compression and decompression.
  • Supports complex data types and advanced nested data structures.

What are the benefits of Parquet?

  • Good for storing big data of any kind (structured data tables, images, videos, documents).
  • Saves on cloud storage space by using highly efficient column-wise compression, and flexible encoding schemes for columns with different data types.
  • Increased data throughput and performance using techniques like data skipping, whereby queries that fetch specific column values need not read the entire row of data.

Apache Parquet is implemented using the record-shredding and assembly algorithm, which accommodates the complex data structures that can be used to store the data. Parquet is optimized to work with complex data in bulk and features different ways for efficient data compression and encoding types. This approach is best especially for those queries that need to read certain columns from a large table. Parquet can only read the needed columns therefore greatly minimizing the IO.

What are the advantages of storing data in a columnar format?

  • Columnar storage like Apache Parquet is designed to bring efficiency compared to row-based files like CSV. When querying, columnar storage you can skip over the non-relevant data very quickly. As a result, aggregation queries are less time-consuming compared to row-oriented databases. This way of storage has translated into hardware savings and minimized latency for accessing data.
  • Apache Parquet is built from the ground up. Hence it is able to support advanced nested data structures. The layout of Parquet data files is optimized for queries that process large volumes of data, in the gigabyte range for each individual file.
  • Parquet is built to support flexible compression options and efficient encoding schemes. As the data type for each column is quite similar, the compression of each column is straightforward (which makes queries even faster). Data can be compressed by using one of the several codecs available; as a result, different data files can be compressed differently.
  • Apache Parquet works best with interactive and serverless technologies like AWS Athena, Amazon Redshift Spectrum, Google BigQuery and Google Dataproc.
REPORT

The agentic AI playbook for the enterprise

What is the difference between Parquet and CSV?

CSV is a simple and common format that is used by many tools such as Excel, Google Sheets, and numerous others. Even though the CSV files are the default format for data processing pipelines it has some disadvantages:

  • Amazon Athena and Spectrum will charge based on the amount of data scanned per query.
  • Google and Amazon will charge you according to the amount of data stored on GS/S3.
  • Google Dataproc charges are time-based.

Parquet has helped its users reduce storage requirements by at least one-third on large datasets, in addition, it greatly improved scan and deserialization time, hence the overall costs. The following table compares the savings as well as the speedup obtained by converting data into Parquet from CSV.

DatasetSize on Amazon S3Query Run TimeData ScannedCost
Data stored as CSV files1 TB236 seconds1.15 TB$5.75
Data stored in Apache Parquet Format130 GB6.78 seconds2.51 GB$0.01
Savings87% less when using Parquet34x faster99% less data scanned99.7% savings

How does Parquet work with Delta Lake?

The open source Delta Lake project builds upon and extends the Parquet format, adding additional functionality like ACID transactions on cloud object storage, time travel, schema evolution, and simple DML commands (CREATE/UPDATE/INSERT/DELETE/MERGE). Delta Lake implements many of these important features through the use of an ordered transaction log that makes data warehousing functionality possible on cloud object storage. Learn more in the Databricks blog post Diving into Delta Lake: Unpacking the Transaction Log.

Additional resources on Parquet


Frequently Asked Questions

What tools and platforms work well with Parquet files?

Parquet integrates well with interactive and serverless technologies such as AWS Athena, Amazon Redshift Spectrum, Google BigQuery, and Google Dataproc. These engines take advantage of Parquet's columnar layout to scan only the columns a query needs, which reduces the amount of data read and lowers compute costs. This selective-read behavior fits the large-batch, analytical querying these platforms are built for.

Can Parquet be used for anything besides structured data tables?

Yes, Parquet is designed to store big data of any kind, including structured data tables, images, videos, and documents. This flexibility comes from its support for complex, nested data structures rather than a rigid, table-only schema.

What other file formats are similar to Parquet?

Parquet is similar to other columnar-storage file formats used in the Hadoop ecosystem, specifically RCFile and ORC. Like Parquet, both formats organize data by column rather than by row so that analytics queries can skip over irrelevant data more quickly.

How much can switching from CSV to Parquet actually save?

In one comparison, converting a 1 TB CSV dataset to Parquet shrank it to 130 GB, cut query run time from 236 seconds to 6.78 seconds, and reduced the data scanned per query from 1.15 TB to 2.51 GB. Altogether, that amounted to roughly 87% less storage, a 34x faster query time, and about 99.7% savings in query cost.

What method does Parquet use to handle complex, nested data structures?

Parquet uses what's called the record-shredding and assembly algorithm to accommodate complex data structures. This approach lets Parquet store nested data efficiently while still allowing queries to read only the specific columns they need, which keeps I/O to a minimum.

Get the latest posts in your inbox

Subscribe to our blog and get the latest posts delivered to your inbox.