🔒 Cookie policy must be accepted. External Content is blocked from info.ness.com. This content requires you to accept our cookie policy. Accept Cookie Policy to view this content.

For years, most organizations have relied on a familiar cloud-native stack for big data analytics, with managed clusters, serverless ETL, and a query engine on object storage. But is it still the fastest or most cost-effective for scale, complexity, and data formats required for modern workloads?

A recent benchmark by a Map & Location Software Provider for their Traffic Analytics platform put this question to the test. They faced challenges with their legacy stack:

  • Slow Queries: Processing highly nested, compressed JSON GPS probe data took ‘minutes to hours’ .
  • Exorbitant Costs: The cost of running user queries was prohibitively high.
  • Query Limitations: Several complex queries were simply ‘not feasible’ on the legacy platform.
  • Cumbersome Data Sharing: Delivering data securely to external partners was inefficient.

They compared their existing cloud-native architecture to a Databricks-based modern data Lakehouse solution, and the results were striking.

The Benchmark: What 400X means

The team ran 8 queries (3 simple, 1 medium, 4 large) on the raw JSON probe data for the client.

Legacy Stack:

  • Compute: EMR 6.15.0 (Spark 3.4.1)
  • Capacity: 1 Primary + 8 r5.xlarge workers (32 GiB RAM, 4 vCore)
  • Query: Serverless query engine on object storage

Modern Lakehouse Stack:

  • Compute: Databricks 16.4 LTS (Spark 3.5.2)
  • Workers: 3x i3.xlarge (30.5 GB RAM, 4 Cores)
  • SQL: Serverless Warehouse (Size: Small)
  • Format: Tables converted to Databricks Delta and Apache Iceberg file formats

The Results: Speed, Cost, and Scalability

Performance (Speed)

The modern Lakehouse solution on Databricks is ~20 times faster than the legacy serverless SQL engine and over 400 times faster than the baseline EMR setup.

Crucially, the two largest queries (“Large C” and “Large D”) could not run on the baseline platform but ran in 15 and 61 seconds , respectively, on Databricks.

  • Total Baseline Managed Cluster Time: 69,110 seconds
  • Total Serverless SQL Engine (Parquet) Time: 3,262 seconds
  • Modern Lakehouse (Delta) on Databricks Time 145 seconds

Price/Performance (Cost)

The Databricks Solution was 400 times lower in costs than the original Baseline Cluster .

Even with a slightly higher cost per query than the lightweight SQL engine, the 20X performance gain means an entirely different business experience—from “run overnight” to “run interactively.”

  • Total Baseline Managed Cluster Cost: $149.00
  • Total Serverless SQL (Iceberg) Cost: $0.11
  • Total Modern Lakehouse on Databricks (Delta) Cost: $0.34

Concurrency & Scalability

The Databricks serverless warehouse scaled from 5 to 200 concurrent queries with near-linear performance, autoscaling clusters as per requirement, without impacting other users. In the legacy stack, however, heavy queries often caused unpredictable slowdowns.

Key Technical Lessons & Best Practices

The case study provided key technical best practices learned during the engagement:

  • Flatten for Speed: For deeply nested JSONs, it’s better to flatten or de-normalize the structure into Silver tables. Storage is cheap, latency isn’t.
  • Use Databricks Unity Catalog Volumes: For the best I/O performance, mount S3 object storage to DBFS as a Volume in Unity Catalog.
  • Optimize Storage: Both Delta and Iceberg formats were used with predictive optimization and liquid clustering to reduce TCO and engineering effort.
  • DLT Pipelines: Three “Traffic Analytics” DLT Pipelines handle streaming and incremental data, loading 33.4 million records (1.60 GB) in just 21 minutes.
  • Simplify Sharing: Efficient and secure Delta Sharing replaced older, fragile sharing patterns and drastically simplified external delivery.

For the customer, migrating from a traditional cloud-native stack to a Databricks platform solved their core challenges:

  • Speed: Queries ran in seconds instead of hours, enabling real-time analysis and insights.
  • Cost: Significant TCO reduction compared to the client’s existing baseline setup.
  • Capability: Running complex queries that were previously impossible is now routine.
  • Modernization: Real-time streaming, simplified external data-sharing, and ML-driven data gap-filling.

What This Meant For The Business

For Business/Client:

  • Queries now run in seconds instead of hours, enabling real-time analysis and a huge boost in Team productivity.
  • Teams gain interactive exploration for varied workloads like analytics, BI, and ML.
  • Dashboards can be refreshed during live discussions, accelerating decision-making.
  • Databricks Unity Catalog provides cleaner, safer, and unified data governanceacross business units and teams.
  • Operational overhead is reduced with a projected decrease in compute costs.
  • Composable architecture stacks that can be blended with other leading tools or frameworks with ease. This minimizes the risk of vendor lock-in.

For Client Customers And End Users:

  • Fresher, more accurate traffic insights for navigation and mobility products.
  • More reliable routing and congestion detection due to higher-frequency processing.
  • Faster rollout of new analytics features, benefiting automotive and logistics partners.
  • Better city and mobility intelligence, driven by richer probe coverage.
  • Ability to scale to more regions and higher data volumes without performance trade-offs.

The Bottom Line

There is a compounding cost when decisions are delayed, insights arrive too late, and strategic opportunities are missed. Legacy architectures still have their place. But for complex, large-scale, high-performance, and scalable analytics, modern Lakehouse platforms on Databricks deliver transformative outcomes.

Ready to see what your team could do with 400X faster analytics? Ness, with a proven track record across industries, is happy to run your benchmark.



Let’s Engineer What’s Next. Together.

Partner with us to build intelligent solutions faster and smarter — we’re ready when you are.

Our "Contact Us" webform relies on a tracking cookie. Your current cookie preferences do not permit these cookies. To contact us through our "Contact Us" webform, please ["Allow All"] cookies in Manage Cookie Settings option in our Cookie policy. Alternatively, you can email us directly at [email protected].