Pioneers Insight Method Research Author
Databricks: From Data to Decisions [Business Breakdowns Episode 238]
Back to Episodes

Databricks: From Data to Decisions [Business Breakdowns Episode 238]

Summary

  • The host frames Databricks as more mysterious to the general public than companies such as Stripe, while Tu’s simplest frame is data processing at scale: the pain of spending “80 to 90% of your time” unifying messy data before you can even run an average, multiplied across unstructured logs, video, and clickstream. Every headline use case — fraud detection, movie recommendations, pricing, inventory — traces back to pipelines that make raw data analyzable, then feed machine-learning models built on top.
  • The founding DNA is seven Berkeley AMPLab academics circa 2009 who made three bets — “cloud was going to be big, data was going to be big, and open source would be a good way to build a business” — and “in hindsight, all three of those bets were very good bets.” Their outsider status was an edge: unaware of the Red Hat services precedent, they reasoned from first principles that monetizing open source means building “a better product that is worth paying for,” even if the community feels betrayed — “you need to be willing to be a villain.”
  • Tu describes a pivotal moment around WCM’s investment as Databricks proved it could cross from the data-engineer/data-scientist persona into Snowflake’s data-warehouse turf: the SQL product announced earlier this year is “on pace to be a billion dollars in revenue,” dwarfing Snowflake’s reverse move into data engineering. Tu calls the multi-product-to-multi-persona expansion “a tremendous TAM expansion in its own right” and the proof Databricks is “successfully becoming a true platform.”
  • Current scale: over $4B ARR, roughly $1B of it AI-related, net dollar expansion above 140%, free cash flow positive, and a capital-light model — core data-processing workloads are CPU-based and are not compute-intensive in the same way as AI-native companies, though model serving can entail some GPU costs. The AI tailwind is framed as durable rather than spiky: “you don’t have an AI strategy without a data strategy,” a driver “not dependent on whether we achieve AGI or not.”
  • Marketing and category creation are underrated strengths: the “lakehouse” coinage was ridiculed as “almost too clever” and is now a very real, defined category that industry observers have coalesced around, and Delta launched with free T-shirts reading “Delta Lake is Spark on ACID.” Strategic pricing follows the same savvy — the “address book” analogy for giving away strategic products like governance for free, and refusing to charge for storage via open formats to architecturally undercut Snowflake’s data-move-in model.
  • The repeated multi-billion fundraises aren’t funding operations — they largely offset employee stock compensation and the IRS tax bill once RSU liquidity is provided, part of a “private MAG 7” dynamic where going public is now “a discretionary decision.” Tu calls IPO timing “literally a trillion-dollar question,” noting staying private through the 2022 growth-tech cycle let Databricks “play offense” in ways public peers couldn’t.
  • Key risks: sustaining R&D execution at scale — “we’ve seen examples in this space of companies that have taken their eye off the ball” — and repeating the lakehouse playbook for agentic AI with Agent Bricks and Lakebase, where the category is still being defined. Tu’s core lesson is long-termism with visible trade-offs: never shipping an on-prem product despite potentially giving up near-term monetization, because “every single time you make more of a short-term-oriented decision that inevitably opens you up to some sort of vulnerability down the road.”

Deep dive

1. Databricks demystified: the Excel headache at industrial scale

  • Tu’s opening frame for a company whose use cases sound unrelated — “recommending movies or pricing strategy or fraud detection”: everyone has received a spreadsheet where prices sit in wrong columns, mixed currencies, some written as text, and spent “80 to 90% of your time just going through the process of unifying all the data” before computing a simple average. That pain point is data processing; Databricks is that “at a completely different scale,” across unstructured log files, images, video, and clickstream data.
  • The e-commerce illustration carries the value chain: deciding how many T-shirts to stock could draw on digital-ad performance, competitor sales, credit-card data — and if you can get that data into shape, you can feed it into a model that answers the question. The host’s addition rings true: the upfront data workload often “stops you and limits ultimately what you’re doing.”

2. Seven Berkeley academics and three prescient bets

  • The founders came out of Berkeley’s AMPLab around 2009, with early cloud-computing research literally one floor below — and unlike AI today, “the idea of cloud was still somewhat controversial.” One co-founder had created Apache Spark, which applied distributed, scale-out compute to data processing; scale-out architecture also enabled the proliferation of compute, storage, and data processing that contributed to the data explosion.
  • Ali’s telling, per Tu: three bets — cloud will be big, data will be big, and open source will be a good way to build a business — made amid the optimism of the Twitter/Airbnb/Facebook era. “In hindsight, it turned out all three of those bets were very good bets.” Tu’s through-line: “the culture and the DNA of the organization — you can trace a lot of the decisions back to this founding story.”

3. Monetizing open source: two home runs and the willingness to be a villain

  • Ali’s framework as Tu relays it: you must hit two home runs — mainstream adoption of the open-source technology, then building a business on top of it, where “the open-source technology ends up becoming one of the business’s main competitors” and better-distributed rivals can monetize it better than you.
  • The academic outsider advantage: not knowing the Red Hat services-and-support precedent, they reasoned from first principles to a simple answer — “you need to create a better product that is worth paying for.” The host’s quip: “Simple answer, maybe not simple execution.” The hard part is social: after being celebrated for gifting technology to the world, “all of a sudden you need to be willing to be a villain” by withholding bells and whistles from the free version.
  • On what goes behind the paywall, Tu sharpens the host’s freemium-LLM analogy: conventional wisdom monetizes ancillary enterprise features like single sign-on, but those aren’t the core product and won’t command as much. Databricks built a fully proprietary Spark implementation with enterprise-grade performance, reliability, and scalability — the equivalent of “the better model that is smarter and will give you better answers, you do have to pay for.”

4. Why “Databricks,” not “Spark” — and the ladder of products

  • Unlike Docker or MongoDB, which named the company after the technology, Databricks forwent Spark’s brand equity because “from day one they always felt like it was going to be more than just Spark” — many bricks applied to the broader data problem. Tu reads the name itself as “a reflection of this long-termism.”
  • The expansion sequence served the same personas first: MLflow (open-sourced again) extended into the machine-learning toolchain for data engineers and data scientists; Delta Lake was a first step toward addressing ACID requirements — atomicity, consistency, isolation, and durability — needed for transactional workloads where data inconsistency can’t be tolerated, unlike an inventory analysis where a slightly stale source may not break the answer.
  • Tu’s aside on core competency: “one of their core competencies is marketing” — Delta Lake launched with free T-shirts reading “Delta Lake is Spark on ACID.”

5. The platform moment: a $1B data warehouse and the lakehouse land-grab

  • A pivotal proof point around WCM’s investment was Databricks laddering from Delta Lake to full data-warehouse workloads and shipping a SQL product directly competitive with Snowflake, crossing from data engineers and scientists to traditional SQL data analysts. “To expand further to multi-persona… was just a tremendous TAM expansion in its own right.” Announced earlier this year: the warehouse product is “on pace to be a billion dollars in revenue.”
  • Empirically, the unstructured-to-structured direction won — that $1B “dwarfed the analogous revenues that Snowflake has had around moving to data engineering” — though Tu credits execution as much as architecture, “not being able to run an A/B test in different versions of the world.”
  • The “lakehouse” coinage — data lake plus data warehouse — drew “quite a bit of ridicule… almost too clever” at launch; today it’s “a very real, defined category that industry observers have all coalesced around.” Tu: educating the market on why lakehouse architecture is “the best of all worlds” is “an incredible piece of the story that Databricks probably doesn’t get enough credit for.”

6. Snowflake, the TAM that big data failed to unlock, and stickiness

  • The market reality is multi-vendor: a classic pattern is Databricks processing data upstream, then storing that data in a Snowflake warehouse — Snowflake now moving upstream, Databricks downstream. Snowflake was the next-generation cloud data warehouse, an existing market; the data-lake side had “never been as well-established for basically lack of good enough technology.” Hadoop predated Spark, and companies such as Cloudera were built on it, but the technology was not good enough. The early-2010s big-data boom then hit Gartner’s “trough of disillusionment”: companies stored volumes of data “because big data, why not?” then found it “very difficult to get anything out of it.” Solving that was “massively TAM expansionary.”
  • On stickiness beyond the disclosed 140%+ net dollar expansion: use cases like a streamer’s next-movie recommendation are revenue-generating and mission-critical, not back-office analysis; “data gravity” from cataloged data adds another layer; and processed data gets reused across products — even if one product sunsets, another leans on the same pipelines.
  • The fraud-alert walkthrough: Databricks provides the pipelines, processing, model fine-tuning, and post-hoc model evaluation; typically a separate customer-built application takes the action, with “the model output from Databricks informing that action.”

7. AI: a durable tailwind, a native-customer base, and the agentic bet

  • Quantitatively: $4B+ ARR, about a quarter ($1B) AI-related. Tu’s favorite property as an investor is “multiple ways to win,” starting with consensus that “you don’t have an AI strategy without a data strategy” — models “can only do so much” without clean, cataloged data. That makes the growth “perhaps not as spiky on the upside but also not as volatile on the downside” if AI sentiment turns, and “not dependent on whether we achieve AGI.”
  • Second and third ways to win: AI-native companies including the largest AI labs use Databricks internally, and the product bet — Agent Bricks and Lakebase — builds the stack for enterprises’ agentic applications, automating labor with a TAM “probably just as infinite” as data itself. The right to win sits in everything around the model: RAG, vector databases, embeddings, and model evaluation to quantify whether unpredictable LLM agents are “doing what we think that they should be doing.”
  • On hyperscaler risk, Tu corrects the host’s premise — the clouds do have competing offerings; Databricks’ positioning just obscures it. The pattern since the first strategic Microsoft partnership, Azure Databricks, is deliberate co-opetition: the partnership was important to jump-start Databricks’ modernization, and customers using Databricks consume hyperscaler compute and storage. Databricks has “never positioned themselves in such a way that the hyperscalers are 100% incented to kill them.”

8. The financial model, why they keep raising, and what could break

  • Usage-based pricing on compute — but Tu argues they monetize more than compute via strategic giveaways. Ali’s “address book” analogy: no handset maker charges for the address book despite its importance; likewise, governance — a “single pane of glass” over metadata — is an example of a strategic layer that can be given away to drive adoption. Offense too: embracing open formats and refusing to charge for storage attacked Snowflake’s move-your-data-in model architecturally, not just on price. Pricing is evaluated through total cost of ownership relative to performance. The model is capital-light and core workloads are CPU-based rather than compute-intensive in the same way as AI-native companies, though model serving — hosting inference endpoints — introduces some GPU cost at “a very different order of magnitude,” and Jensen’s view that these workloads may shift to GPUs could change things.
  • The fundraising puzzle answered: proceeds mostly offset employee stock compensation and the tax bill once employees have been given enough opportunity for RSU or option liquidity to trigger IRS treatment — part of the “private MAG 7” dynamic where late-stage capital infrastructure makes going public “a discretionary decision.” Tu: “literally a trillion-dollar question.” Staying private through the 2022 cycle let Databricks “continue to play offense” on sales and R&D while many public peers could not — raising the bar any IPO rationale must clear.
  • Risks and lessons converge on culture: continued R&D execution at scale can’t be assumed (“companies that have taken their eye off the ball… it can really show up in the numbers”), and the agentic category-creation challenge must be executed “in the same way that they executed on the lakehouse.” Tu’s takeaway he now screens other companies for: long-termism with identifiable trade-offs — e.g., never shipping an on-premises product despite potentially giving up near-term monetization — because “there are certain times when people talk about being long-term where it isn’t clear what the trade-off is.”