Pioneers Insight Method Research Author
AI in Investing with Daloopa's founder Thomas Li
Back to Episodes

AI in Investing with Daloopa's founder Thomas Li

Summary

  • Thomas Li’s core distinction is that AI generates plausible objects, while fundamental investors often need exact numerical processing. It can synthesize a Croatia itinerary or compare two bodies of text spectacularly, but financial models expose its weakness: “I don’t need you to seem correct… I just want correct.” Treating a “really, really shiny hammer” as the tool for every task invites hallucinated data and fragile analysis.

  • The highest-value near-term use case is a “smart black-lining function” that compares an investor’s prior notes with new transcripts and disclosures. Feed the system the latest earnings call, a conference transcript such as Morgan Stanley TMT, and notes written before those events; then ask where management’s language contradicts the thesis. Many buy-side firms cannot do this in public ChatGPT because they are not allowed to upload their notes. Modern models need little note formatting, but Li draws a hard boundary around numbers, where the logic “isn’t quite there yet.”

  • For sophisticated funds, the decisive advantage is context rather than owning the single best foundational model. Internal notes, analyst models, licensed transcripts, consensus, filings, and structured historical data can turn a general model into a finance-specific tool; public ChatGPT lacks much of that context and cannot reliably obtain every required document. Li’s wager: “A second-tier foundational model with the right context fed in will solve more problems than the best foundational model with no access to data.”

  • AI adoption is being driven less by age or firm type than by senior sponsorship and organizational agency. Li finds senior investors more likely than juniors to see AI as an Excel-, Google-, or iPhone-scale transition, while juniors may distrust technology that touches the craft they are still learning. Banks, pod shops, long-onlys, and small funds can all move quickly when one empowered person commits to iterating for 12 to 36 months.

  • Automation is likely to increase Wall Street’s output before it reduces its hours or headcount. Excel did not end 100-hour weeks; it enabled more detailed models, and better PowerPoint tools produced more presentation work. Firms therefore ask how associates can cover more companies, react faster, or do more research—not how to preserve revenue while employing four fewer people.

  • Human alpha remains because finance is an “industry of corner cases,” and the relevant decision is rarely just whether a company will grow. Durable advantages can come from a retail investor’s long holding horizon or a multi-manager’s ability to strip out risk factors, increase trade velocity, and leverage residual alpha. Alternative data and channel checks that once created alpha in 2017-2019 are now widespread; the harder edge is interpreting who already knows what and who will buy next.

  • Short-term public-market investing increasingly resembles explicit game theory with crowding treated as a modeled risk factor. The key questions become: who is at the poker table, who is in the hand, and who might buy in five minutes, four days, or one month? Even a pristine earnings beat can produce a 10% decline when pod shops are already full, but sophisticated market-neutral books seek to hedge crowding by balancing crowded longs against crowded shorts.

  • The most powerful institutional application may be a center book that learns when each analyst is genuinely skilled. Funds can consolidate forecast-versus-actual performance, normalize variance by sector, separate numerical accuracy from positioning, and identify narrow patterns—such as an analyst excelling in mid-cap internet names into earnings but struggling with large caps during conference season. Scale matters because more proprietary workflow data allows the system to distinguish repeatable judgment from luck.

Deep dive

1. AI generates plausibility, but financial analysis demands correctness

  • Li’s level set begins with the mechanism: foundational models predict the next object—first a word, then a sentence, pixel set, paragraph, or regenerated answer checked against prior work. That generation feels human because people also create new language while thinking, but it is not synonymous with every kind of cognition.

  • Much of an analyst’s job is processing rather than generation: extracting numbers, reconciling inconsistencies, comparing cost growth with revenue, tracking customer-acquisition cost for software, or connecting hotel occupancy with the broader picture. Those tasks require structured understanding of relationships, not merely a fluent continuation.

  • The contrast is clearest in Li’s travel example. ChatGPT can compress reviews, Reddit posts, and travel reports into a useful one-page Croatia itinerary; asked to capture every intricacy of a public-company model, it starts producing material that “seems correct.” His objection: “This is a public company with public disclosures. I just want correct.”

  • Li’s warning is the familiar tool bias intensified: “When you’re a hammer, everything looks like a nail. When you’re a really, really shiny hammer, everything definitely looks like a nail.” AI can succeed and fail spectacularly; the workflow must follow what the model was built to do.

2. Thesis blacklining is a better use case than autonomous modeling

  • Li’s sharpest practical application is to compare language-based objects. Give an AI an analyst’s existing notes, the latest earnings transcript, and a subsequent conference appearance, then ask where management’s new statements conflict with what the analyst previously believed.

  • Walker turns that into a falsification exercise: could an investor write a thesis, load it into the system, and ask whether the past six months disproved it? Li says to make that test more specific—provide defined earnings and conference transcripts plus notes written before those events, then identify the inconsistencies.

  • The notes themselves need not be polished. Li says that around “ChatGPT-3” some structuring helped, but newer model iterations can generate multiple items, check their own output, and loop across varying spans of words; for language, “how you structure it doesn’t matter to the AI anymore.”

  • His hedge is categorical around numerical work: this flexibility applies “so long as it’s words.” Ask the system to compare an analyst’s model with reported financials, and Li says the requisite number logic is not there yet—or is not there at all.

3. Institutional AI is constrained by missing context and document plumbing

  • The comparison workflow needs data that general-purpose ChatGPT may not possess: private notes, purchased earnings and conference transcripts, actual analyst models, consensus, and company disclosures. Most buy-side firms also cannot upload their notes to ChatGPT under internal restrictions. Li says even public filings are not reliably available as source documents; a model may surface a blog linking to EDGAR without actually ingesting the filing.

  • Andrew’s pushback—worth keeping—is that Google can immediately locate an SEC filing or Nvidia investor-relations deck. Li’s answer is that document acquisition remains an old-school systems problem: finding the page and downloading the document are difficult but not large-language-model problems.

  • That distinction explains why large funds build internal tools. They can combine analyst-generated notes, purchased transcripts, structured historical data from vendors such as Daloopa, internal estimates, and a chosen foundational model so the system can answer the real question: “How has my company changed?”

  • One example captures the potential: a firm could compare each analyst’s estimates with company historicals and ask whether accuracy is improving. With the right context, Li says a one-sentence query can replace what might historically have required a research associate to spend two weeks grinding through Excel.

4. The fastest adopters are strategic sponsors, not necessarily young analysts

  • Walker expects a generational split among a 50-year-old PM, a 35-year-old senior analyst, and a 25-year-old AI-native junior. Li instead sees strategic orientation as the separator: senior people are required to think more strategically, while a new graduate is often focused on completing today’s assigned activities.

  • The analogies senior investors invoke are institutional missed turns: refusing Microsoft Excel, avoiding Google as a research analyst, or remaining on BlackBerry while everyone moved to iPhones. Li’s observation is counterintuitive: “The more senior the person… the more likely they are to want to be AI adopters,” while juniors can be more skeptical.

  • Walker supplies a plausible reason. A senior PM already works by asking an analyst two questions and making a yes-or-no decision, so querying AI feels familiar; the junior who was trained to read every 10-K and cultivate special insight may feel the technology threatens both learning and identity.

  • Firm category is similarly inconclusive. Li sees big banks, pod shops, long-only managers, and very small funds moving fastest when “a single person” has the agency to build a process, fund the product, and say, “Let’s go.”

5. Productivity gains expand ambition before they shrink headcount

  • Asked whether AI means fewer investment-banking juniors or more researchers gathering bespoke information, Li rejects a simple labor-substitution frame. Wall Street firms see themselves as growth engines and ask how saved time can expand coverage, responsiveness, competitiveness, and revenue.

  • His historical analogy is the migration from Lotus Notes to Microsoft Excel. Easier, more detailed modeling should theoretically have ended punishing schedules, yet “people still work 100 hours a week”; improved PowerPoint and logo-alignment shortcuts likewise raised the expected output rather than returning time to employees.

  • The same logic applies to AI. Time no longer spent updating models, summarizing things, or transcribing earnings calls can be redeployed into calling franchises, analyzing more companies, or getting to market faster. The typical question is not, “Now I need four fewer associates.”

6. Automate the grind, but preserve the judgment built through doing it

  • Walker’s central training concern is the bodyguard who memorizes 100 speeches but cannot answer a novel question. Building a key model by hand can force an investor to confront assumptions; outsourcing every intermediate step may create the illusion of understanding until an edge case arrives.

  • Li says he does not see much wholesale outsourcing of thinking. “Finance is also the industry of corner cases”: history may rhyme, but unusual facts, human assumptions, and judgment dominate—and that is also the enjoyable part of the work. Agents are better directed at the hours spent “spinning wheels.”

  • A 15-page earnings transcript illustrates process leverage. An investor studying tariff effects across a peripheral industry might otherwise spend roughly 16 hours reading every transcript; a summary system plus 30 minutes of targeted Q&A can provide the needed map without replacing the intellectual decision.

  • Walker’s echo-problem example is the failure mode: ChatGPT supplies $500 million of 2023 earnings for a business with only $200 million of revenue, and downstream work compounds the false premise. Li’s blunt observation is that he does not see many firms worrying about this echo problem, leaving Walker’s demand for source checking unresolved.

7. Context and compliance—not model selection—separate strong deployments

  • Li calls the key institutional challenge “the context problem.” Differences between the best and second-best algorithms may be small, while the quality jump from adding relevant data and guardrails can be enormous; teams should frame a product problem before framing an AI problem.

  • This is also why supposedly surprising professional use cases remain limited: Li tells Walker the real surprise is “how little AI adoption there is.” The appetite exists, but firms hesitate to let a general enterprise model observe enough queries to infer their next major investment or potentially expose trade intent.

  • The compliance concern is not imaginary. If a backend operator asked what a fund was preparing to buy based on all its prompts, Li thinks the model might make “a really, really good guess.” Funds therefore prefer an internal system or a vendor trusted in the way they trust Bloomberg despite Bloomberg’s visibility into substantial trading activity.

  • Li expects foundational models to remain competitive, more like cloud infrastructure than a winner-take-all search monopoly. “Where the rubber meets the road” is the application layer: assembling context, guardrails, and workflows on top of relatively accessible models, much as internet and mobile infrastructure enabled differentiated software businesses.

8. Human alpha migrates toward risk architecture and market game theory

  • Li defines a true alpha source as something systematically valid and protectable for a reasonable period—not forever. Buffett’s classic edge is holding horizon; retail investors can share that advantage because professional managers face daily, weekly, or monthly marks.

  • Thomas points to multi-managers’ risk modeling as a different source. By using tools such as PCA to identify and hedge major factors, they seek residual alpha, increase the number and velocity of trades, apply cheap leverage, and deliver the uncorrelated returns LPs want—up 15% in good markets or 20% in bad ones.

  • Li does not think foundational models can simply manufacture those hedged slivers of alpha. Nor does he still view alternative data, satellite imagery, credit-card data, channel checks, or even proprietary checks as durable standalone edges; they may have been powerful in 2017-2019, but high-volume adoption made them closer to reading a 10-K.

  • Walker resists smoothing that conclusion: surely information generated in the field remains proprietary. Li’s resolution is that the enduring analyst task is not merely obtaining the fact but determining how it changes supply and demand for the stock relative to what every other player knows.

9. The next buyer matters as much as the fundamental result

  • Li’s short-term mechanism is simple: stocks rise when there are many buyers and few sellers. The analyst must ask who has access to the same information, who is five minutes or two days behind, and who might take the other side in 30 minutes, four days, two weeks, or one month.

  • His conference example makes the chain tradeable. Channel checks suggest a structural shift in a sub-piece of the business because of potential retaliatory tariffs; management appears at a conference in seven days and may disclose or be asked about it. The edge therefore has a seven-day catalyst window and requires identifying the likely buyer afterward.

  • Walker pushes the game theory deeper: if every pod reaches the same tariff conclusion, the better trade might be to buy because crowded shorts will cover when management sounds less negative than feared. Li agrees funds go “pretty deep”: “Who’s sitting at a poker table is the most important thing to figure out. And once you figure out who’s sitting at the poker table, you need to figure out who’s in the hand.”

  • Crowding itself is an explicit risk factor. A company can beat pristine expectations and fall 10% because every pod already owns it; sophisticated books try to neutralize that exposure with crowded shorts against crowded longs. Walker’s DeepSeek-day specimen is that Nvidia was not necessarily the worst casualty—the hidden crowding unwound hardest in AI-adjacent power names such as Talen and Vistra.

10. Proprietary data and center-book learning compound institutional scale

  • Walker asks whether a general-purpose ChatGPT that becomes 10 times better will render today’s internal applications wasted. Li is “very confident” it will not: model improvement does not create access to proprietary notes, purchased datasets, analyst histories, or internal workflow data.

  • Large firms consequently possess a real data advantage. They buy more external information and generate more workflow data, allowing bespoke systems to solve internal problems that a context-free frontier model cannot. Li’s formulation is decisive: the best builder will be “the people with the most access to data.”

  • Center books show what scale can unlock. A fund can test whether Walker is unusually good at mid-cap internet names into earnings but weak at large caps during conference season, then double, halve, or hedge exposures accordingly; it can also uncover unconscious preferences such as favoring high-dividend stocks in low-rate environments.

  • The evaluation layer separates forecasting from positioning and skill from luck. Consolidated forecast-versus-actual data, normalized for sector-specific variance, can show that an analyst’s assessment was precise even when all four positions were wrong because a macro shock rocked the book—or that the analyst got the stock right despite missing the numbers. What remains after stripping out observable factors may be positioning.

  • Li’s closing condition is agency and patience: buying AI will not transform a firm tomorrow or eliminate ten associates. Building now, iterating with users, and improving research, coverage, and reaction speed “might take 12 months, might take 36 months,” but waiting risks the painful catch-up that followed earlier platform shifts.