Run R programs at scale using Apache Spark's distributed computing engine with familiar R syntax
SparkR is a tool for running R on Spark. It follows the same principles as all of Spark’s other language bindings. To use SparkR, we simply import it into our environment and run our code. It’s all very similar to the Python API except that it follows R’s syntax instead of Python. For the most part, almost everything available in Python is available in SparkR.
SparkR is an R package that lets R users run distributed data processing on Apache Spark™ using familiar R syntax, so they can scale analysis beyond what fits in local memory. It follows the same principles as Spark's other language bindings, exposing core Spark capabilities through an R interface. This makes it possible for R data scientists to work directly with big data on Databricks clusters without switching languages.
SparkR works much like the Python API for Spark, except it follows R's syntax instead of Python's. Most features available to Python users on Spark are also available in SparkR. This parity means R users can access nearly the same distributed processing capabilities as PySpark users when working on Databricks clusters.
You start using SparkR by importing the package into your existing R environment and then running your code as you normally would. Because SparkR follows the same principles as Spark's other language APIs, the workflow will feel familiar if you've used Spark in another language before. For setup details, see the SparkR Overview Documentation.
Almost everything available in the Python API is also available in SparkR. SparkR exposes Spark's core capabilities through an R package, following the same design principles used across Spark's other language bindings. This means R users generally don't have to give up functionality to work with Spark.
SparkR lets you scale R-based analysis beyond what fits in local memory by distributing the processing across a Spark cluster. Rather than being limited to a single machine's RAM, you can process larger datasets while keeping familiar R syntax. This makes it straightforward for R data scientists to work with big data on Databricks clusters.
Subscribe to our blog and get the latest posts delivered to your inbox.