Pioneers Insight Method Research Author
Back to Pioneers
Arvind Jain
Founders 4 Curated Dialogues

Arvind Jain

Glean · Founder & CEO

Frontier Insights

Frontier Thesis: LLMs are commoditizing into interchangeable infrastructure as open-source achieves parity for 90%+ of enterprise workloads. Front-line ROI belongs not to frontier model providers, but to the application layer governing proprietary enterprise context, workflows, and execution.

Strategic Decisions: Glean is expanding beyond enterprise search into autonomous agents, leveraging deep SaaS APIs, authoritative information filtering, and multi-model routing to automate core workflows and capture service-labor spend.

Risks & Warnings: Value capture is threatened by unreliable agents, brittle governance, complex permission indexing exposing data vulnerabilities, and the friction of human behavioral adoption.

Key Views & Dialogues

⁠Why OpenAI and Anthropic Won’t Win the App Layer | Glean Founder

  • 🗓️ Date2026-07-11 | 🎙️ Show:20VC

Open-source models now handle 90%+ of enterprise use cases, with GLM 5.2 within 3 months of frontier capability and trusted for most of Glean’s workloads. That cost advantage could shift most enterprise workloads open within 3 years, while falling token prices and consumption pricing pressure frontier labs and Microsoft’s bundling.

View Dialogue Notes & Key Takeaways
  • Open source just hit its enterprise inflection. Arvind Jain says 90%+ of enterprise use cases can now be fully handled by many models including open source, and GLM 5.2 — arriving within 3 months of frontier capability, “literally a month back” — is the first open model Glean’s own team trusts with “majority of our workloads.” His call: majority of enterprise workloads run on open source within 3 years “for sure,” and the only gate is comfort with Chinese models — “it’s not open source versus closed source.”

  • The frontier model business may be mispriced. Standalone, it’s “probably not as lucrative as everybody believes”: fierce competition even in a three-way lab race plus open source at “an order of magnitude” cheaper, and Jain has “heard rumors that OpenAI was going to drastically reduce their model prices.” Meanwhile every model bizarrely raised per-token prices in the last 6-9 months; Stebbings’ jab — “they needed to prove that they were good businesses before they went public” — and his warning that if AI gets much cheaper, these loss-making labs “which prop up our entire global economy” are very threatened.

  • The labs’ app-layer push is shallow. Anthropic’s vertical packs (Figma, legal, finance) are “quite shallow” — net-new usage that expands the market, not workload displacement — while enterprises are “terrified” of operational dependence: institutional learning accumulates inside the agent doing the work, so enterprises need control of the agent and its compounding learnings. Jain’s advice to founders: treat the labs as “a huge asset, not a competition.”

  • Jain argues teams should get bigger — the episode’s sharpest disagreement. Glean is 1,000+ people and Jain wants 5,000 in five years: with symmetric AI access, the competitor that keeps headcount ships a 10x better product and “they’re going to beat you.” On the labor-versus-tokens framing, Jain says technology costs should fall: “we’ve not put technology cost and labor cost in the same sentence ever before… this is not how technology works” — inference costs will fall by orders of magnitude.

  • AI ROI is a throughput problem, not a model problem. Nearly 100% of Glean’s code is AI-written, yet across companies “the actual shipping speed of products has not increased.” His triage agent resolves 95% of production issues for a 15-person on-call team — at $1M/month, a cost he questioned against the humans. Advanced use cases sit with only ~5% of employees; the fix is investing in context rather than letting models “brute force their way” through raw MCP connections.

  • Consumption pricing can break bundling. Microsoft is a significant competitor today (“hard to compete with free”), but “once you move towards consumption, there’s no inherent bundling advantage” — enterprises pay per unit of work wherever users choose to do it, weakening the Copilot lock-in argument Stebbings pushed.

  • China has the leading open models; the US is playing catch-up. On OpenRouter the top six models by usage are Chinese, with Anthropic the first US entry at seventh; Jain says China is the only country producing models outside the US, while Stebbings mentions perhaps a little activity in France. Jain’s explanation: model training needs upfront capital that skunkworks open source can’t fund. The US needs to build its own open models, with Nvidia among the motivated funders.

  • 🔗 Original source & video: ⁠Why OpenAI and Anthropic Won’t Win the App Layer | Glean Founder

Listen to full conversation →


AI Enterprise - Databricks & Glean | BG2 Guest Interview

  • 🗓️ Date2025-12-23 | 🎙️ Show:BG2

Ali Ghodsi argues that AGI already exists and LLMs are commodities, shifting durable value toward proprietary data, business processes, and applications rather than model providers. The 95% project-failure rate reflects healthy experimentation, but frozen models and computer use remain unresolved; enterprise adoption, agent revenue, and Glean’s move toward a proactive personal work companion are the catalysts to monitor amid a clear startup bubble.

View Dialogue Notes & Key Takeaways
  • Ali Ghodsi’s central claim: “I think we have AGI. We really have it” — by the definition his 2009 Berkeley AMP Lab used, it’s already satisfied, and the industry is just “moving the goalpost.” He sorts the field into three camps: the superintelligence quest (frontier labs, most of the capital, “I would be very worried there”), the Turing-Award researchers (Sutton, LeCun — sober, 20 years out, “probably the ones that are right, unfortunately”), and camp three — Databricks and Glean — extracting economic value from the AGI we already have.

  • “The LLM is a commodity” — interchangeable like gas stations, “just compare price,” with users switching models in a day unlike any prior platform battle. Model companies can still be valuable (“TSMC is very valuable”) but as fabs; the real moat is proprietary data and business process — “there’s not an AI out there that understands your secret sauce and your data. That’s not a commodity.”

  • On the capex math — ~$250B to Nvidia implying ~$500B capex needing ~$1T of AI revenue vs a $400B total software industry — Arvind Jain’s resolution: AI isn’t extending software, it’s converting services dollars, an industry “25 times larger than software.” Ali’s answer is camp-dependent: if superintelligence lands, “any of your cost equations pale in comparison”; camp three doesn’t need it.

  • Is there a bubble? Yes, but not binary: “there are startups with zero revenue worth 10, 20, 30 billion. That’s a bubble.” Yet both call OpenAI and Anthropic up over 12 months — ChatGPT and Gemini “on fire,” coding having “only eaten into a small portion of that market.”

  • Ali reframes the MIT 95%-failure stat as healthy: “that’s actually what you want” from an experimentation phase, and hopes for similar stats next year. The working 5%: RBC agents producing equity research notes 15 minutes after an earnings call vs a 2-hour industry standard, Merck’s “Teddy” transformer for gene-regulatory drug discovery, and 7-Eleven’s fully agent-automated marketing stack.

  • Value accrual call: Arvind thinks the intelligence layer stays thick — “maybe half of enterprise value” — while Ali says most value goes to apps, “I just don’t know which apps”, invoking 1998: everyone bet on Cisco routers and portals, the winners were Facebook/Airbnb/Uber. Software isn’t dead (Salesforce is “a full ecosystem of workflows,” not a database), but data entry is the wedge — “Zoom is really the perfect data entry application.”

  • Longs and shorts: Ali is long agents and speech (“as long as you’re using a keyboard, we haven’t nailed speech” — keyboards “basically going to disappear”), Arvind calls coding and customer-service automation “a little bit over hyped.” Brad is long proactive AI that comes to the user — the shift that takes “5% power users to 100%.” Glean, fresh off a $200M revenue run rate, is building toward a privileged personal work companion.

  • 🔗 Original source & video: AI Enterprise - Databricks & Glean | BG2 Guest Interview

Listen to full conversation →


Arvind Jain on building Glean and the future of enterprise AI

  • 🗓️ Date2025-08-05 | 🎙️ Show:Gradient Dissent

Glean’s early BERT-based enterprise-search bet became a generative-AI wedge, combining customer-specific retrieval and permissions-aware indexing with GPT, Gemini, or Claude for synthesis and reasoning. The strongest ROI signal is reasoning across unstructured data, but stale or missing knowledge and retrieval failures remain larger risks than hallucinations as Glean expands toward an AI operating layer for every employee.

View Dialogue Notes & Key Takeaways
  • Glean’s founding thesis was enterprise search, but its early transformer bet made the 2019 product unusually well positioned for generative AI. Arvind Jain began with a universal pain point—research suggested employees spend one-third of their working time finding information—and used BERT-based models to match concepts rather than keywords. As generation and reasoning improved, Glean evolved from “a Google for you in your work life” into “ChatGPT for work life.”

  • Glean uses frontier models where broad capabilities already exist while building enterprise-specific retrieval. It trains small models on a customer’s corpus for custom embeddings, and separately fine-tunes small open-domain models for narrow search tasks such as spell-checking, synonym handling, and acronym expansion. It uses GPT, Gemini, or Claude for synthesis and multi-step reasoning. Jain’s rule is blunt: “Do not reinvent things that have been invented already.”

  • Permissions and data freshness—not merely model quality—are the central enterprise-AI constraints. Glean imports governance from systems including Google Drive, Slack, and Salesforce, bakes permissions into its index, and retrieves only documents the signed-in user may access. Because embeddings themselves might leak restricted information, customer-specific models are trained only on subsets Glean judges safe.

  • Jain’s repeat-founder pattern is a contrarian bet on universal problems in markets others have abandoned. Lukas Biewald pressed him on whether Rubrik and Glean reflected exceptional execution rather than novel ideas; Jain agreed enterprise search had produced “only failures” and become a “dead area” where investors did not want to invest. His conviction came from two changes: SaaS made fragmented data much worse but also more accessible, while transformers made semantic understanding technically viable. His Google-derived operating model is to put innovation first, hire smart engineers, and largely let them build.

  • A concrete ROI example comes from reasoning across the 95% of enterprise data that is unstructured. A non-engineer in Glean’s finance team asked an agent to combine Salesforce customer lists, shared Slack sentiment, and product usage into green/yellow/red account-risk profiles after churn appeared. Jain says the result was better than a conventional dashboard because it could incorporate subjective textual evidence.

  • Glean evaluates the complete answer pipeline, while conceding that enterprise AI cannot eliminate errors. It derives “golden” question-answer sets from real interactions such as well-received Slack replies, tests retrieval and model changes against them, and uses LLMs as judges. For hallucinations, it checks answers “line by line” against supplied source material and may suppress unsupported claims or abstain—but Jain says stale, missing, or poorly retrieved knowledge causes more failures than fabrication alone.

  • Jain rejects labor reduction as the most valuable AI strategy and instead wants every employee surrounded by a scalable “dream team.” He imagines assistants, coworkers, and coaches helping each person do up to 90% of the work they need to do, while companies retain and even expand teams whose members can do “ten times more work.” The transformation may feel incremental—one task at a time—until workers discover they are fundamentally different from two years earlier.

  • 🔗 Original source & video: Arvind Jain on building Glean and the future of enterprise AI

Listen to full conversation →


AI is Making Enterprise Search Relevant, with Arvind Jain of Glean

  • 🗓️ Date2025-05-15 | 🎙️ Show:No Priors

SaaS APIs, cloud infrastructure, and transformers turned enterprise search from a graveyard market into a platform that understands questions and documents conceptually across massive private corpora. Glean combines permissioned internal data with world knowledge, then extends search into assistants and agents that can answer questions or perform work in connected systems. Security remains both the gating risk and an adjacent opportunity: better search exposed salaries and sensitive M&A material, pushing Glean toward AI-readiness while employee adoption still requires training.

View Dialogue Notes & Key Takeaways
  • Jain’s account is that enterprise search became tractable as SaaS APIs, cloud infrastructure, and transformers arrived. He began thinking about Glean in late 2018, founded it in early 2019, and used Google’s BERT model plus customer-specific embeddings built on business content in version one. One of Glean’s largest customers has more than one billion documents—the size of the entire internet in 2004, by Jain’s comparison.

  • Good enterprise search requires more than vector search or an ever-larger context window alone. Enterprise systems must distinguish current, authoritative information from decades of obsolete material and present it coherently; dumping “one million documents” into a model out of chronological order still creates a reasoning problem. Jain says finding the right source—or discovering that nobody documented the answer—is often harder than hallucination.

  • Glean has expanded from a “Google in your work life” into an assistant and agent platform on top of enterprise data and knowledge. Glean Assistant combines world knowledge with permissioned internal data, while function-specific apps and agents can restrict sources, specify tone, and perform work in connected systems. The sharpest example is HR: employees can ask about benefits or PTO, but answers should use only content “authorized or blessed” by the people team.

  • Security is simultaneously the gating factor for enterprise AI and an adjacent product opportunity for Glean. Jain estimates 90% of company knowledge is private in some form, so all platform access must honor source permissions. Good search exposed existing governance failures—including salaries and a sensitive M&A document—creating the paradox that “we can’t sell because the product is so good” and pushing Glean toward AI-readiness and security.

  • Employee adoption is a behavior-change problem even when the interface is only one box. Twenty years of Google trained users to enter one or two keywords, leaving many unsure how to use a conversational assistant; Jain’s conclusion is that “AI is actually very unintuitive.” Beyond near-term ROI, he argues companies should train an AI-first workforce now for the organization they want three years from today.

  • Glean’s top-down sales motion was dictated by product architecture, despite Jain’s original PLG ambition. Even one employee’s search requires indexing the entire company corpus, so small-seat deployment is expensive and company-wide rollout makes the economics work. His advice, when possible, is to start PLG and enterprise sales together, using PLG as lead generation rather than waiting years to build the commercial motion.

  • Management is staying focused on the assistant and agent platform because its own promise remains largely unsolved. Jain’s pitch—ask any question or assign any task, and Glean will safely use public and internal knowledge to complete it—is still “a long, long way” from reality. The stated destination is a personal team of assistants, coworkers, and coaches that “does ninety percent of your work” and “make[s] us all 10Xers,” clearly framed as ambition rather than current capability.

  • 🔗 Original source & video: AI is Making Enterprise Search Relevant, with Arvind Jain of Glean

Listen to full conversation →