AI is Making Enterprise Search Relevant, with Arvind Jain of Glean
Summary
- Jain’s account is that enterprise search became tractable as SaaS APIs, cloud infrastructure, and transformers arrived. He began thinking about Glean in late 2018, founded it in early 2019, and used Google’s BERT model plus customer-specific embeddings built on business content in version one. One of Glean’s largest customers has more than one billion documents—the size of the entire internet in 2004, by Jain’s comparison.
- Good enterprise search requires more than vector search or an ever-larger context window alone. Enterprise systems must distinguish current, authoritative information from decades of obsolete material and present it coherently; dumping “one million documents” into a model out of chronological order still creates a reasoning problem. Jain says finding the right source—or discovering that nobody documented the answer—is often harder than hallucination.
- Glean has expanded from a “Google in your work life” into an assistant and agent platform on top of enterprise data and knowledge. Glean Assistant combines world knowledge with permissioned internal data, while function-specific apps and agents can restrict sources, specify tone, and perform work in connected systems. The sharpest example is HR: employees can ask about benefits or PTO, but answers should use only content “authorized or blessed” by the people team.
- Security is simultaneously the gating factor for enterprise AI and an adjacent product opportunity for Glean. Jain estimates 90% of company knowledge is private in some form, so all platform access must honor source permissions. Good search exposed existing governance failures—including salaries and a sensitive M&A document—creating the paradox that “we can’t sell because the product is so good” and pushing Glean toward AI-readiness and security.
- Employee adoption is a behavior-change problem even when the interface is only one box. Twenty years of Google trained users to enter one or two keywords, leaving many unsure how to use a conversational assistant; Jain’s conclusion is that “AI is actually very unintuitive.” Beyond near-term ROI, he argues companies should train an AI-first workforce now for the organization they want three years from today.
- Glean’s top-down sales motion was dictated by product architecture, despite Jain’s original PLG ambition. Even one employee’s search requires indexing the entire company corpus, so small-seat deployment is expensive and company-wide rollout makes the economics work. His advice, when possible, is to start PLG and enterprise sales together, using PLG as lead generation rather than waiting years to build the commercial motion.
- Management is staying focused on the assistant and agent platform because its own promise remains largely unsolved. Jain’s pitch—ask any question or assign any task, and Glean will safely use public and internal knowledge to complete it—is still “a long, long way” from reality. The stated destination is a personal team of assistants, coworkers, and coaches that “does ninety percent of your work” and “make[s] us all 10Xers,” clearly framed as ambition rather than current capability.
Deep dive
1. Three shifts made a graveyard market tractable
Jain’s thirty-year search arc frames the discontinuity: keyword systems found matching words, while LLMs can “deeply understand” both a question and a document, match them conceptually, and answer directly. His categorical conclusion is that the paradigm “has completely shifted” and become less brittle.
Glean’s timing was unusually favorable. Jain began exploring it in late 2018, founded the company in early 2019, and used transformers in version one: Google’s BERT model, which had been trained on internet data, followed by customer-specific embeddings built on business content. “Vector search,” “RAG,” and “generative AI” were not yet the vocabulary; internally, it was “embedding search.”
The biggest practical unlock, however, was SaaS. Pre-SaaS vendors had to locate servers and storage across private data centers; SaaS standardized versions and exposed content through APIs. At Rubrik—the origin of Glean—information sprawled across 300 SaaS systems, employees complained they could find nothing, and Jain discovered “there’s nothing to buy.”
Cloud supplied the required scale: one of Glean’s largest customers has more than one billion documents, equal to what Jain says the entire internet contained in 2004. Enterprises also lack the behavioral signal generated by a billion web users, making techniques such as semantic understanding more necessary.
2. Larger context windows do not eliminate retrieval engineering
Elad asked whether traditional information retrieval survives as models improve. Jain’s answer: embeddings are only one component; enterprise search must prefer information that is correct today, current, and written by someone with authority, while discarding obsolete material.
Jain is skeptical that near-infinite context resolves this soon. Give a model one million documents mixing material from today, four months ago, three years ago, and two days ago, and even human-like intelligence with exceptional memory and speed will struggle. Ordering and organizing evidence remain important for good reasoning.
The first product progression turned a “Google in your work life” into Glean Assistant, which looks more like ChatGPT. It combines world knowledge with internal company context, identifies the user, and limits answers to information that person can access—effectively a permission-aware personal “sidekick.”
Functional demand then pulled Glean from answers into workflows. HR teams wanted benefits, PTO, and vacation responses restricted to “authorized or blessed” people-team content, with prescribed behavior and tone. Glean called these experiences apps before agents became fashionable; once they began doing work in connected systems, “agents” fit.
3. Better search exposed the enterprise’s hidden security debt
Jain estimates that 90% of company knowledge is private in some form. An enterprise therefore cannot dump all internal data into one model and expose it company-wide: Glean indexes each document or Slack conversation alongside the identities permitted to access it and enforces those controls for every signed-in user.
The unexpected rollout problem was fear of success. Customers said, “I don’t want a good search product” because Glean surfaced governance gaps that already existed; employees found other people’s salaries and, at one company, a sensitive M&A document before the transaction had happened.
Sarah’s proposed remedy—use LLMs to classify potentially sensitive documents—matched Glean’s direction. The company went beyond mirroring source permissions to consider who is asking, what they are asking, and whether a returned result feels safe to show. Jain says Glean consequently “ended up becoming a security product” that companies buy to become AI-ready.
4. AI adoption requires education, not merely an empty prompt box
Jain assumed a product with “no UI”—just one box—would require no instruction. Search validated that assumption; Assistant did not. Users remained trained to enter one or two keywords, while the more adventurous asked unanswerable questions such as, “What should I do with my life?”
His conclusion is blunt: “AI is actually very unintuitive.” Adoption improves when capabilities appear incrementally inside recognizable work—for an engineer, perhaps an offer to produce a two-page tutorial on unfamiliar technology—rather than expecting employees to invent sophisticated paragraph-long prompts unaided.
Businesses naturally demand efficiency gains or top-line improvement from AI spending, but Jain thinks that framing neglects education. Three years from now, a CEO should want an AI-first workforce expert at exploiting a technology that is “not perfect, it’s not easy, it makes mistakes, and it hallucinates, but it’s powerful.” Today’s objective is building that expertise through daily use.
5. Product structure dictated the sales motion, while conviction beat the priors
Rubrik entered an established replacement market with buyers and budgets; Glean entered a “graveyard” with no search line item and a product dismissed as a vitamin. Jain’s response to the bad priors was unusually simple: persistent pain means nobody is solving it yet, and every idea offers ten reasons to quit—sometimes the answer is “just do it.”
Jain candidly says his “dream was to do PLG,” but one employee’s search still requires indexing the entire company’s knowledge. That makes small-seat deployments expensive and company-wide rollout natural. Given a choice, he would launch PLG and enterprise sales simultaneously, using PLG as a lead-generation funnel rather than postponing enterprise sales for three years.
The operating scope remains deliberately narrow: Glean Assistant for every worker and an agent platform for every business process. Yet the promise—“ask it any question, or give it any task”—is far from fulfilled because information may be missing, obsolete, or simply the wrong needle from the haystack. Jain says, “We’re a long, long way from solving even the pitch.”
Jain’s personal uncertainty remains visible alongside that ambition: after running R&D at Rubrik, he is learning to be Glean’s CEO and says, “I don’t think I’ve learned it yet.” The long-term product vision is equally personal—a team of assistants, coworkers, and coaches around every worker, doing 90% of the work and helping even a new graduate enter that world, with the goal of making “all of us” 10Xers.