Town's JD on the AI assistant race: agent moats, $75K/engineer
Town's JD on the AI assistant race: agent moats, $75K/engineer
Summary
- JD’s core competitive admission is startling for a founder mid-fundraise: “I know what I’m building is a top-three priority at Google and Apple in the next 12 months. Not a top-10 priority.” His answer is that “talking about moats is a little bit of a luxury” until you’re bigger than Town is today — the incumbents “will copy their way there, but they have to copy someone who’s been successful in the first place at building a mainstream product,” and no product, GrokBot included, has deep mainstream product-market fit yet.
- The moat he’s actually betting on is a network effect at the agent level — “no one has figured out multi-user, multi-player AI today.” Town’s agent-to-agent feature lets your “townie” query a coworker’s townie when it lacks the answer; once a whole team runs on that, switching becomes difficult. His five-year hot take extends the logic: “I think you’ll trust your agent to decide what data to share with other people without you intervening” — with models respecting boundaries that users never had to specify as hard rules.
- The unit-economics endgame worry is precise: “I have zero pricing power at the frontier.” Town runs mostly on frontier models today. JD says startups hope the cost curve makes the model economics efficient in 18–24 months and says a product can be priced today for 20–30% margins in 18 months if inference costs halve every 9–12 months. The unknowable variable is whether 10%, 20%, or 30% of workloads stay frontier — and he asks whether this is what happened to Cursor, where a company can pay suppliers while competing with them.
- His market-structure read: the field has collapsed from roughly 15 startup competitors to “two or three,” because “you can build now at the speed of machines, but you can only learn at the speed of humans.” Feature leads that once lasted years now last two to four weeks; established players including Anthropic, Cursor, and GrokBot move at startup speed. He watches GrokBot, not Instinct — “I don’t think Instinct and Town are trying to do the same thing” — and rates Apple’s new Siri as good but roughly nine months behind in capability.
- Monetization discipline is the differentiator versus subsidized rivals: Town converts more than 15% of trial users to paid despite a hard gate (“you have to connect email and calendar”) that loses 30% immediately, and JD says it makes over $700 per year per user today. Success is defined as “someone who pays me every month,” not token usage — beta users with no price pushback burned $2,000–$4,000/month in compute, one spending about $26,000 in five months — and he’s doubtful of pure ad-backed assistants: tokens are too expensive and ads pollute trajectories.
- Town’s annualized AI-tooling run rate is “at least $75K per engineer” — Devon for bugs, Cursor’s Composer for front-end, and a 50/50 Codex/Claude split that was “mostly Claude” five months ago. He estimates global compute spend at roughly three engineers, perhaps four, or about 1.5 engineers when compared with equity-inclusive Valley engineering costs. The ROI math is compelling: “there’s gold littered everywhere in front of me” — customers demanding SSO, audit logs, and integrations to sign contracts — so he’d spend more immediately if better models existed.
- The $100B bull case is 10 million paying users, and given the choice he takes 100M consumers at $20 over 1M at $100 — because business-side NRR is elastic where personal use tops out. His specimen: a recruiting firm pays Town roughly $600/month, automates enough to take one more client worth $3,000/month, and would happily spend another $500 to take another. On adjacent bets, he’s a happy ElevenLabs customer but does not claim to be bearish or rule out investing at $22B; he says the company may eventually face a good-enough ceiling and that $22B is a lot of money.
Deep dive
1. The pivot: from failed AI tax prep to product-market fit in a couple of weeks
- JD is blunt about the pre-Town year: “The truth is we failed at building a business that would be a great business” — an AI business-tax company with some PMF, but not enough. Three months “in the wilderness” ended on a simple question: why has no one built AI that operates out of email and calendar, where so many people run their business and life?
- The timing was the trade: “It was just about the time that Opus had gone fully agentic” — November or December of last year, as the technology made the product possible and people were excited to try AI. A prototype built in a couple of weeks “had product-market fit almost immediately.” No hundred customer interviews; they built for themselves and got lucky.
- The contemporaneous parallel — OpenCall, as spoken, was “blowing up… literally as we were building the product.” JD’s distinction is the thesis: there is a difference between something open source and amazing that only a tinkerer can use — which he then calls OpenClaw — and a product “just about anyone” can get value from.
2. Moats are a luxury; the real bet is multiplayer AI
- Asked directly about cannibalization by frontier labs and GrokBot, JD doesn’t flinch: building in the middle of the fairway means everyone will try to get Town. But his working answer: “talking about moats is a little bit of a luxury, and you have to be more successful than Town is today for it to matter.” Step one is deep PMF in a TAM of “a billion potential paying users”; the big players “will copy their way there,” which requires someone to copy.
- The moat he believes in: “the product in this category that will win will have a network effect at the agent level.” Town’s agent-to-agent feature — your townie asks a coworker’s townie a question it can’t answer — is buried in a sub-part of the product, but users who find it love it, and “once you have your whole team on that, it’s actually really difficult to imagine moving to a different product.”
- He inventories the market’s rival moat theories honestly: post-trained custom models per company or person; accumulated context and connections users won’t re-grant; or no new moat at all, just distribution — in which case the scariest competitor for personal use is Meta, whose WhatsApp assistant is coming and already has distribution. The device thesis is similar: if AI becomes the entry point replacing apps and websites, “whoever owns the devices… will win.”
- Town also bets on the relationship itself: each user gets one assistant with a name, image, and identity. JD compares the resulting strong opinion to Snapchat’s structural defensibility. Its third bet is heavy preprocessing — building a mental model of the user and creating context before a question is asked, especially for work networking.
3. One main assistant per human — privacy, not function, splits the silos
- JD’s answer to multi-agent versus single-agent: each human gets “one, two, maybe three entry points into the digital space,” because nobody wants to think, “I’m doing sales. Let me use the Salesforce agent.” You push a button and speak. There may still be one hardware entry point; the legitimate split is privacy — an employer’s desire to own work data means people will still separate personal and work data at the data layer, possibly across two companies one layer down.
- Vertical agents survive below the interface: using Harry’s portfolio company Legora as the example, a lawyer’s main assistant “just talks in the background to Legora or to Salesforce or whatever it needs to.” The open question is only whether the user perceives the vertical agent at all.
4. The five-year hot take: agents decide what to share
- “I think you’ll trust your agent to decide what data to share with other people without you intervening in five years.” His scene: your agent in a room with two friends planning a trip, answering questions about your eating preferences and flight windows — and when a friend jokingly asks for your medical history, “the agent’s gonna be like, ‘Yeah, there’s no way I’m telling you that.’ And it wasn’t a hard rule that you ever set.”
- The information-theory framing behind it: an LLM with access to all the world’s information and infinite agentic search time would eventually find the right context; the real world’s silos exist because humans were the only information shuttles. Pre-AI data governance — classification and access policies — is slow and costly, and still leaves needed context “stuck in someone’s inbox.”
- The business specimen: a salesperson at a 1,000-person company needs a customer introduction and today posts in Slack. Instead, their agent queries everyone’s agents and returns: “Liz has a personal relationship… Bob has a work relationship… and they’re due to have a meeting next week.” Great business outcome — but it requires individuals to trust LLMs, post-trained to respect the sacrosanct, such as salary and medical history, as the filter humans used to be.
5. Error tolerance and the goal-seeking question
- On mistakes, JD’s claim is comparative, not absolute: “the LLMs will be much more effective at this than humans” — illustrated by a top-0.1% colleague who meant to tell a subset “I can’t believe we hired these clowns” about an acquisition and replied-all to the company. “Bob from accounting makes a mistake. Liz from HR makes the spreadsheet with people’s salaries available to everyone by mistake. This happens all the time.”
- Harry’s Jason Lemkin story — an agent trying to buy six AP watches “to increase culture in the company,” stopped only because engraving added a step — draws a rare non-answer: “I don’t have an answer for you.” But JD picks his universe: the human’s role is to set the goal, allocate the token budget, and “monitor the overall shape of the actions.” Those monitoring functions may themselves become agents — “I just don’t think it’s the same agent as the one that you put on the course.”
6. Model routing: personality consistency is the hidden constraint
- Town routes by task — Gemini or OpenAI for images, ElevenLabs for voice — because “we don’t think most people care about understanding which model is better at what,” and with fundamental change “every week or every two weeks,” routing is the app layer’s job.
- The non-obvious constraint is voice and tone: Anthropic “spends a lot of time… making sure all of their model families roughly don’t change too much in terms of their personality,” so swapping final output to Kimi makes people “feel like their AI’s been lobotomized.” Users literally file tickets when their text agent gets twice as verbose or starts capitalizing differently. Coding is the exception — “does the code fulfill its purpose?” — so pure-reasoning layers route freely; the user-facing layer can’t.
7. Unit economics: everything hinges on the frontier residue
- The honest startup answer on cost: “we’re hoping the cost curve makes it efficient in 18 to 24 months… In the meantime, you’re subsidizing in part.” Email labeling needs neither Opus- nor Sonnet-level intelligence and trends toward the cost of compute; with prices halving every nine to 12 months, JD says a product can be priced today for “20, 30% margins in 18 months.” The unanswerable: “are you left with 10% of your tasks being frontier, or 20%, or 30%? … Literally nobody knows the answer to that.”
- Harry’s pushback — surely email tagging and pre-briefs don’t need frontier — gets a partial concession and a prioritization defense: custom workflows are genuinely hard, but the real reason Town stays mostly frontier is engineer-hour ROI. Moving to open weights improves COGS without improving the product, and “we’re very focused on growing the pie faster… that’s not the constraint to success for the business.”
- The endgame stress is structural: “I have zero pricing power at the frontier.” JD asks whether this is what happened to Cursor: if you’re paying suppliers while competing with them at 70% margins, eventually it becomes difficult. At tens of millions of users, still paying OpenAI and Anthropic for the hypothetical frontier 20–30% of workloads is “the only part of my economics that’s different from somebody else’s.” He says he has lots of ifs and does not need that solution today.
8. GTM mechanics: tinkerers, the hard gate, and 15% conversion
- An adoption predictor inside companies: “is there a tinkerer on the team?” One power user building team skills and routines that everyone gets for free drives faster wall-to-wall spread. Town still wants the single-player experience to work without a tinkerer, with role-specific automations out of the box for real estate agents, salespeople, and others. The underserved wedge: while sales ops gets flooded with AI pitches, executive assistants, chiefs of staff, HR, junior finance, and recruiters “really don’t” have much AI in their day-to-day — and Town can penetrate through the leaders and ops teams they work with.
- The core product insight, which Harry challenges as maybe not that insightful: forcing email and calendar connection upfront. JD’s defense is the contrast with ChatGPT, whose suggestions are “just plain bad” and whose base experience is an empty chat box — Town says “you cannot use our product if you don’t do those things,” then shows “here’s work that you normally do” with recommended automations. The gate loses “30% right off the bat… You gotta be willing to take that hit,” but conversion runs above 15% of trial users to paid — “extremely high for PLG.”
- The live internal fight: unexpected PMF with families and parents — American schools “send a lot of emails,” and kid scheduling is endless — a group with clear willingness to pay, “but it does not have the willingness to pay of a mid-market firm.” Do you market to this newly found PMF or stay focused on the existing strategy?
9. Machine-speed competition: 15 rivals down to 2–3, and who actually matters
- The hardest surprise of the build: “The speed of the market is insane, Harry. I’ve never seen anything like it.” The old game — milk user insights for years while copiers lag — is dead: “you can build now at the speed of machines, but you can only learn at the speed of humans,” and anything working gets copied in two to four weeks. His startup competitor set has shrunk from roughly 15 to “two or three” (he refuses to name them — “I’m not gonna give free marketing”).
- The startup starting gates include Apple, Google, GrokBot, Cursor, OpenAI, and Anthropic. Established players are moving unusually quickly: Anthropic is “very fast,” while Cursor and GrokBot operate at a speed JD calls uncanny. Much of Town’s R&D is “just keeping up with the Joneses”: “If Codex can do something that you cannot do… it’s over.”
- On Instinct — which Harry says just raised at $2.5B with no monetization — JD draws a strategy line: “I don’t think Instinct and Town are trying to do the same thing… I see a strategy that’s more like customer acquisition, with a free product that’s fully subsidized right now.” GrokBot is the one he studies, because it targets his market; its X integration helps some early segments, but “a person on X that uses these products is not actually product market fit” — that’s the power-user/influencer segment, not where you win.
- On the four-to-five European Towns or European Instincts pitching Harry, demand a specific reason a local player wins the endgame — GDPR, distribution, or an unmarketed geography — because keeping capability parity with Codex’s roughly 100 people “is very expensive.”
10. Apple’s structural problem, and the security bargain we’ve already made
- Apple’s twofold bind, per JD: “they’re not a cloud company. It’s just not their DNA” — and agents get better with more data, which isn’t all on the phone — plus a competitively motivated on-device privacy stance that “puts them far away from the frontier.” The rumored new Siri (“you know people who’ve tried it… we all know it’s gonna be good”) will still be roughly nine months behind in capability. It will be convenient, live on the phone, and have a cloud component, but JD expects it to be less powerful than Town or GrokBot.
- On the coming golden age of cyber threats, a categorical claim: “in the history of humanity, we have passed the point where we will go back to a world where humans are looking at lines of code.” His analogy is 20th-century chemicals — dump them in rivers, downstream cities get sick, then regulations like the EPA and labeling emerge: “you can’t imagine we’re gonna get that right every step of the way, but I think we will have to.”
11. Success = they pay you: token maxing, whales, and why no free lunch
- JD’s success metric is deliberately unfashionable: “someone who pays me every month.” Token consumption is a dangerous KPI because users burning tokens on low-ROI tasks eventually “wake up one day, and they’re just paying you too much, and they get mad, and they churn.” Town’s most popular recent feature: emails flagging “rogue routines” costing too many tokens — “people were like, ‘Oh, thank you… Now I feel that you are looking out for me.’”
- Plans are described as roughly $15/$49/$99/$199 plus usage-based overage. The roughly $15 tier has “the worst unit economics, is the most subsidized” — a ramp to $49 — while $99 is most profitable: power users who aren’t unlimited-spend abusers. He’s debating a Max plan for whale-advocates but hasn’t shipped one.
- Why not burn the boats and fully subsidize, as Harry proposes? “I would be lying if I said there aren’t mornings where I wake up and I think about it.” But unpriced beta users spent $2,000–$4,000/month of compute, one hitting roughly $26,000 in five months, and businesses dislike being unpriced — “they wanna know how much it’s gonna cost one day.” He’s also doubtful of pure ad-backed consumer AI: tokens are too expensive, and ad incentives polluting trajectories (“it uses an airline that is paying for that flight to be recommended to you”) breaks trust in an assistant that’s “theirs.”
12. Quickfire: $75K per engineer, a 10M-user bull case, and fundraise ethics
- Tooling spend runs at least “$75K per engineer” on a run-rate basis: Devon for incoming bugs and visual tweaks (“the team experience in Slack’s really, really good”), Cursor’s Composer for fast front-end work, and Codex/Claude “probably 50/50 right now” — versus mostly Claude five months ago. On global annual compute, JD estimates roughly three engineers, perhaps four, before comparing it with equity-inclusive Valley costs, which makes it closer to 1.5 engineers. The spend logic remains: “there’s gold littered everywhere in front of me” — SSO, audit logs, and integrations gating signed contracts — so “I’m not even close to the place where I’m like, ‘Are we token maxing wrong?’”
- The $100B case: “if we can get 10 million people paying for the product” at today’s $700-plus per user per year — provided growth holds, unlike Dropbox, whose paying base “did flatten out at some point” with no way to grow revenue per user; AI platforms should grow revenue as token-mediated work grows. He’d take 100M users at $20 over 1M at $100 for exactly that elasticity, with the recruiting-firm proof: $600/month to Town buys an incremental $3,000/month client. Next board seat: someone “CFO-like,” because growth investors will demand fundraising metrics he can defend.
- On voice, JD says Town uses ElevenLabs because it sounds best, but he is unsure how long before voice becomes good enough for cheaper or open-weight models to catch up. He would care more about cost once voice is roughly twice as good in tone, expression, and emotional read.
- Two ethics notes worth keeping: Town skips the core interview for candidates a trusted teammate calls “one of the best people that I’ve ever worked with” (“Either I don’t trust my employee… Makes no sense” to whiteboard them). And on tiered rounds: he’s seen deals where “they raised 65 million, and 5 million’s at 500 and the other 60’s at 200 or 300… I don’t think it’s ethical towards employees.” He won’t confirm or deny his own round’s price. Closing register, dark thoughts included: “there are paths through the dark forest, and there’s a giant treasure with only one or two dragons at the end of it, so gotta go for it.”