Skip to main content
Nextdoor

CUSTOMER
STORY

Nextdoor powers a faster, leaner data platform with Spark Declarative Pipelines

2–3 minutes

Event ingestion latency reduced from 1 hour, delivering 95%+ faster data access for real-time analysis

75%

Reduction in event ingestion compute costs achieved by modernizing data pipelines with SDP

60%+

Faster pipeline setup by automating workflows on Databricks Lakeflow

Nextdoor is a neighborhood-focused social networking platform where neighbors can connect, share information and build community. Its 46 million weekly active users across 11 countries, including the US, Canada, and the United Kingdom, generate an immense amount of data. As Nextdoor scaled, its existing ingestion pipeline needed to evolve. Delays in data could mean slower analysis and decision-making across the business. 

By modernizing with Apache Spark™ Declarative Pipelines (SDP) on Databricks Lakeflow, Nextdoor now ingests and processes data in near real time—enabling faster decisions, better user experiences and a stronger foundation for growth.

Scaling data to meet growing community demands

Nextdoor connects over 105 million neighbors — and growth at that scale brought challenges. The company was collecting up to 400,000 events per second — from mobile clicks to server side events — using a batch-based system which caused data delays of up to an hour, impeding timely analysis. “In our previous architecture we’d have to wait to add a partition for anyone to query the data,” says Mayur Makwana, Software Engineer at Nextdoor. “We wanted to make it near real-time so we could enable everyone within the company to query the data sooner.”

Beyond that goal, Nextdoor data engineers also wanted “exactly once” delivery rather than “at least once,” clear insights into ingestion, and to achieve all of this without cost increase. Databricks Lakeflow made all of this achievable.

Streamlining Nextdoor’s data pipeline with Spark Declarative Pipelines

Bringing SDP into its data architecture set Nextdoor up to achieve its goals of real-time event ingestion, streamlined schema management and reduced costs—all while maintaining high performance. “We didn’t want to spend time on infrastructure setup, maintaining checkpoints for our applications or running maintenance on these jobs separately,” says Makwana. “The cost benefits we get with SDP compared to doing everything in batch Spark make it worth it. You save time and infrastructure effort, with added benefits like auto-upgrades.”

By implementing SDP on Lakeflow, Nextdoor engineers were able to automate many of the manual processes that were causing inefficiencies. They used Databricks Auto Loader to efficiently manage and partition incoming event data, leveraged schema evolution for dynamic adjustments and implemented tagging for clear visibility into resource attribution. Lakeflow’s optimized write feature also automatically compacts files, improving query performance. Plus, SDP automatically retries if it fails and integrates notifications, reducing engineer intervention and increasing pipeline reliability. “SDP is self-reliant in its operation,” notes Makwana.

Following their initial Lakeflow implementation, Nextdoor engineers continued improving with further refinements such as tuning the autoloader and optimizing compute clusters. In the near future they plan to transition to serverless architecture and leverage more of Lakeflow’s automated observability tools.

The impact of near real-time data

Nextdoor’s efforts paid off. With SDP on Lakeflow, setting up a pipeline is significantly faster—it can be completed in hours instead of days. This improvement not only streamlines the setup process but also greatly reduces the effort required for ongoing infrastructure maintenance.

Nextdoor’s ingestion latency decreased dramatically. Insights that took up to an hour with prior systems now arrive in a span of two to three minutes. “It’s easier, faster, and less setup processing with Lakeflow SDP. Now time is spent on understanding what the processing needs to be, instead of everything else that comes with it,” notes Makwana.

Embracing these advancements also prompted the company to improve schema evolution iterations by moving from JavaScript Object Notation (JSON) to Protobuf (Protocol Buffers) for schema definitions. Having Apache Spark as the driving force for their pipeline enables Nextdoor engineers to conduct robust unit testing as they continue to evolve. ”Our processing latency also decreased from four minutes to forty seconds—an 83 percent improvement. Spark Declarative Pipelines is great to get started with anything you want to process in Databricks,” says Makwana.

With Lakeflow’s ability to handle both structured and unstructured data, along with the overall flexibility and (automated) reliability of Databricks, Nextdoor has reduced its event ingestion compute costs by 75 percent. The company has significantly reduced operational overhead, streamlined its data pipeline and achieved near real-time event ingestion with minimal infrastructure management. The cost savings, enhanced performance and ease of use have enabled the team to focus on driving innovation rather than dealing with complex maintenance, setting the stage for continued growth and more efficient data-driven decision-making.

Frequently asked questions

Explore more