Lovable CEO, Anton Osika: The State of Foundation Models, Grok vs OpenAI, and Replit vs Bolt
Lovable CEO, Anton Osika: The State of Foundation Models, Grok vs OpenAI, and Replit vs Bolt
Summary
- Asked to allocate across the labs — OpenAI at 380, Anthropic at 180, Grock at ~$100 — Osika goes long Grock, short OpenAI (after first saying Anthropic, then correcting himself). The reason is “the slope on the Grock team”: they “hire missionaries for the data curation part” and “the morale is super high,” while “OpenAI has gone through all this mess.” The raw captions also say “Dropbox” has good morale and is growing faster on the enterprise side, likely referring to Anthropic.
- The next leading model “has not been created yet” — “Yes. From China.” He puts it at “a 50/50 chance they will have the best model” and says “we’ll be using a Chinese model at some point,” subject to checking whether it receives data Lovable does not want to share and whether there are other negatives — a striking admission from Lovable.
- Unit-economics honesty: of a paid-usage dollar today, the share passed through to Anthropic/OpenAI “is majority. It’s not everything.” The plan is subscription value plus token-margin optionality — Lovable-built apps were already pushing >$10M in AR through model providers months ago — but he deliberately indexes on mindshare over Revolut-style payback optimization: “you need to look at the weights in my neural network.”
- Defensibility doctrine: “AI startups are like chickens shot out of a cannon… it’s all about flapping fast” — don’t worry about moats on day one. The endgame moat is a platform you can’t leave, with Lovable graduating from “your technical co-founder” to “your co-founder in general.” Among the labs, OpenAI — not Anthropic — is the more serious competitor over 12 months.
- GPT-5 verdict: “oftentimes too ambitious for our users” and “the model is still too ambitious.” Harry’s capability-wise read was that it “hasn’t been a step function improvement.” Lovable still uses Anthropic for code writing, GPT-5 for hard debugging; Osika sees plateauing on nuance but still-exponential sigmoid curves in science and bioengineering.
- Revenue mix at $100M ARR in 7 months: 80% of revenue from people “building real complex applications,” ~10% enterprise (a fast-growing segment — a Google product leader: “we’re never again writing a document about a product”), ~10% hobbyists. By 2035, the vision is “the mostly used interface for humans to AI.”
- Figma is the competitor he respects most, but Figma Make’s design-first entry may “slow you down too much” — his thesis is that detailed design work is increasingly replaced by high-level design direction and AI implementation, not that all design disappears. And on Harry’s charge that “all of you guys suck at security”: “Uh, yes” — followed by the claim that Lovable’s security-review process gives it a lower chance of vulnerability than the average human developer, with a target of 0% vulnerability.
- Talent is #1, brand #2, and capital “not a constraint at all” for Lovable at the application layer. Europe is “hard mode” — a thin network of operators who’ve scaled before — but Lovable is “the biggest talent magnet in Stockholm,” a position “much much more difficult” to hold in San Francisco.
Deep dive
1. The lab trade: long Grock, short OpenAI — and the wildcard is Chinese
- Harry’s forced allocation — OpenAI at 380, Anthropic at 180, Grock at ~$100: “I’d invest in Grock and I would probably short Anthropic. No, I would short OpenAI, let’s say.” The self-correction is on tape and worth keeping — his first instinct went the other way.
- The reasoning is team slope, not current capability: Grock is “hiring missionaries for the data curation part and they call it AI tutoring… the morale is super high,” while “OpenAI has gone through all this mess.” The raw captions render the other lab’s name as “Dropbox” — likely Anthropic in context — and say it “has good morale as well” and is growing faster on the enterprise side from what he hears.
- He rejects the tidy duopoly map — OpenAI wins consumer, Anthropic wins developer/enterprise: “No, I think it’s going to be unknown. There’s going to be something else happening that we don’t know what it is.”
- The something else may be Chinese. Will a leading model come from a lab that doesn’t exist yet? “Yes. From China.” He gives “a 50/50 chance they will have the best model” and says Lovable could use one — “I would have to look into the details… do we give them data we don’t want to give them?” — but “we just want to do what’s best for our customers.” Harry’s addendum: four new Chinese models a week, “the speed of distillation is just insane.” On open vs closed: “the best ones will always be closed,” though open may be what most people choose for flexibility.
2. Defensibility: flap faster than the other chickens
- The opening question — is this a capital arms race? No: “it’s an arms race to build the best team and then… the best brand and trust from your users… for us, [capital] is not a constraint at all.” Capital might constrain foundation-model training because compute is so large. His brand model is the Apple ecosystem: “they obsess about details, maybe too much so they move slowly, but that’s what builds up trust.”
- His friend’s analogy, as told: “AI startups are like chickens shot out of a cannon… it’s all about flapping fast as a chicken, because there are new chickens shot out from cannons every day.” The advice to founders is categorical: don’t worry about defensibility from day one — “execute fast, grow faster,” and think about moats only “when you’re starting to get up there.”
- The endgame moat is switching cost through accumulated value: “Lovable today is your technical co-founder. We want it to be your co-founder in general” — handling admin, finance, operations — “and if you’re on a platform like that, you probably don’t want to leave.”
- On the labs invading the way Claude Code came for Cursor: “in the long term it just comes down to execution of a team.” Lovable’s bet is being “the gateway for humans… the best user experience for AI — and so far OpenAI is doing that better than Anthropic. So I see them as a more serious competitor in 12 months.”
3. Unit economics: the pass-through admission and the patience trade
- Harry’s blunt question — if I give you a dollar, how much goes straight to Anthropic and OpenAI? — gets a straight answer: “If you look at the paid usage, it’s majority. It’s not everything.” Today “you’re paying to build”; the plan is to shift revenue toward subscription once users hit “I love this platform, I’m never leaving,” with AI compute becoming a small part of cost.
- Why not route simple jobs to cheap models now? His analogy: “you’re driving a car and you’re not thinking about what you’re doing. When you’re in a new situation… your brain really goes on fire. We’re not close to being there yet.” The AI is doing “completely different things” every month, so it’s too early to optimize — you “build for what tomorrow’s model can do, not what we have today” (Harry’s phrase, endorsed “to quite large extent”).
- The hidden margin lever: months ago they measured more than $10M in AR flowing through the AI from Lovable applications, all requiring users to wire up model providers themselves. “We’re simplifying… and if we can reduce the underlying cost, maybe we can take a margin there as well.”
- On when to optimize margins, two advisors live in his head: Nick of Revolut (“compute the payback time, then do super hard performance optimization”) versus mindshare — “as many users who just love the brand as possible right now.” Which wins? “You need to look at the weights in my neural network, but it’s some combination of the two” — though he indexes on mindshare. Harry maps the arc: brand first, then funnel optimization, then back to brand — “we have to sponsor race cars.”
4. GPT-5: smart consolidation, too ambitious, no step function
- Before shipping GPT-5, Lovable checked latency, ran quantitative evals, and “vibe checked it in many different ways.” Conclusion: “it’s oftentimes too ambitious for our users” — great “when you have to solve a really really hard problem,” so they shipped it to everyone and watched. “The model is still too ambitious.”
- Collapsing five models into one was “really smart, obvious” and “executed pretty well” — but “it’s inevitably going to fall short in some dimensions… it just is a disappointment that you can’t improve in all the directions at the same time.” Harry’s read — no step function improvement capability-wise — goes unchallenged.
- The production stack today: a “very complex agentic chain” with fast small models doing routing, “for code writing we usually use Anthropic,” and GPT-5 selectable — “better when you’re solving a really hard debugging problem.” The next step-function he wants is context — hyperpersonalization — and over time “paying hundred million dollars for getting the people that train the models.”
- His contrarian quickfire belief: “AI is much better than humans and most people don’t agree” — it’s often “very very stupid,” but “if you give it all the context… it’s smarter than humans.” On curves: plateauing “on the things that we care about — nuance, being good at all the different things at once,” but some sigmoids are still exponential — “science and engineering and bioengineering… a lot of new medicines.”
5. The $100M mix: 80% of revenue is people building real businesses
- $100M ARR in 7 months — against which Harry offers scale: “0 to 10 million in two years was the gold standard.” The split: “80% [of revenue] are building real complex applications” — businesses — with roughly 10% enterprise and 10% hobbyist websites, and he calls that “a good split.”
- Enterprise is a fast-growing sleeper: “enterprises are slower to wake up,” but a Google product leader now says “we’re never again writing a document about a product — we have to build a fully working demo.” Harry’s corroboration from 20Product: two Duolingo designers created chess inside Duolingo with a tool he couldn’t recall. Lovable will build an enterprise sales team but “will not become an enterprise company” — no wine-and-dine.
- The CEO question he’d ask instead of “how do we make engineers more productive”: “how can we get the most information about what we should build as quickly as possible” — which requires everyone in the company working in one place. The biggest enterprise bottleneck he sees is change management for humans, not tooling.
- Harry’s investor case, stated to Osika’s face: the reason to own Lovable is that TAM expansion “is actually incomprehensible” — “website builders” is completely the wrong analogy, the way Uber’s market expansion was hard to foresee. Osika’s own end-state: by end-2026 “your perfect co-founder” from idea through growth, email and marketing included — “one opinionated way to do the entire product life cycle” — and in 2035 “the mostly used interface for humans to AI.”
6. AI compresses detailed design
- His product-lifecycle framing: AI has compressed the early work from idea toward “validated, with external users on it” into minutes or hours; the steps after — growth, testing, QA — are what Lovable must build next, so you don’t need a separate product-design-engineering organization.
- On Figma Make entering from the design side: humans are sometimes too obsessed with small details being perfect. The future is talking design philosophy at a high level while AI implements it; his thesis is that detailed design work will slow most people down too much. Figma survives “for some pixel-perfect things.”
- Yet asked which competitor he most respects across Figma, Bolt, Replit: “I respect Figma… they’re good at listening to their users and building a good product. If they can translate that to the full product life cycle, they’re a very formidable competitor.”
- Will we still prompt in five years? “Yes” — but hyperpersonalization absorbs the detail. His analogy: with a great employee “you just have to say, ‘Let’s go to Stockholm and do a hackathon,’ and it just magically becomes what you want it to be.”
7. “All of you guys suck at security” — “Uh, yes”
- The Replit feud, from his side: a competitor announced poorly-built apps as a vulnerability in a way security professionals told him wasn’t proper disclosure — “so then I went in and bashed them. I think that was a very reactive thing,” though he’d happily say it face-to-face.
- Harry raises Jason Lemkin’s Replit disaster (a wiped database, “code red”) and puts it directly: “All of you guys suck at security. Is that true?” The answer, before the reframe: “Uh, yes.”
- The reframe is the self-driving argument: the average developer ships holes; Lovable tells users to go through security reviews and has the AI perform reviews before giving a green light, so “lovable is going to have a lower chance of having a vulnerability” than that average human — “we need to put that at 0% chance.” Security comes up internally “every day.”
8. Slope, founder mode, and impact over hours
- His hiring signal is slope: “if I talk to someone and I learn a lot of things from them and my conversation is very dynamic… their slope will be very high.” The other tool: “if I could be there with a video camera when they worked in the past, that gives me a lot of signals.” Hardest hire: engineering leaders — past performance doesn’t predictably translate.
- On Zuck’s NFL-style contracts: Zuck is paying for the knowledge of “10 people that know everything about how to train foundation models” — “they wouldn’t perform as well as the engineers in my team doing what we’re doing.” Application-layer talent is a very different type of talent, and harder to identify.
- He’ll keep “most of my impact coming from founder mode,” buffered by a “wonderful chaotic protective layer” of ex-founder generalists — “I’m not planning to be that percentile manager myself.” His stated hiring regret: delegating too much and not staying in the details.
- On 996 and Cognition’s six-day ultimatum: over ten years, balance; “over a two-year period, if you really care about something… just work your ass off.” But he manages to impact, not hours — the keeper’s test, plus the Nik line Harry relays approvingly: “I don’t think about culture. I think about winning.” His culture fix: more farmer, less cowboy — “Do move slow so that we can move really really fast” — against Harry’s pro-China “sticky tape” short-termism: with product-market fit and a brand to defend, you can’t. “Can you imagine if Apple: ah, we deleted your cloud, sorry.”
9. Europe on hard mode — but the biggest talent magnet in Stockholm
- The motivation is explicit: “I want to prove that you can build a generational product, a generational company, and a generational team from Europe, and part of it is on hard mode” — the missing piece being the network of people “that have worked on and have context for all the different stages” of scaling. “There are very few people like Elena in Europe.”
- The offsetting edge: “we are the biggest talent magnet in Stockholm right now… it’s much much more difficult to be that in San Francisco” — picking up underutilized talent and 10x-ing it, plus a stronger culture of “humility and low ego” and doing more with less. Harry adds the churn point: Valley employees leave on a bad day for a bigger OpenAI package, killing the compounding of knowledge.
- Capital is “not a bottleneck” for Lovable, and Harry predicts Lovable spinouts getting instant term sheets — “100%. True.” Would Lovable be less successful in the Valley? An honest non-answer: “I honestly don’t know. I think it would be very successful regardless.”
10. Mistakes, changed minds, and what’s underrated
- His regret: not scrapping the GPT Engineer open-source community at the start — “with the perspective of maximal focus, it was just a bad idea to do two things that were a bit too tangentially related.” The operating mantra: “finding the bottleneck for the company and solving for that bottleneck is the best way to move really really fast.” Tomorrow’s board bottleneck: finding the engineers for the product’s next phase while serving “extreme” enterprise pull without losing founder focus.
- What he got wrong: building for agents “before the models were ready for it.” The lesson — get a product as many people as possible use today, “optimize the entire user experience for those users… that’s your data flywheel.” Harry’s own mea culpa: he believed model performance would commoditize and value accrual would fail — “clearly very wrong and very stupid of me.”
- On benchmarks: they “become less useful over time — there’s something called Goodhart’s law. When you start optimizing for a number, that number stops being a good measure for success.” Lovable’s internal example: thumbs-up clicks, hackable with “fun jokes.” The context was Harry’s underrated pick, Surge — the Scale competitor that “never raised a dollar and it’s a billion two in revenue.”
- An instrumental outside adviser is likely Atle Skalleberg; the raw captions render the name as “atlena” and the companies as Muro, Dropbox, and “N segment.” He ran sales and was essentially CEO, and Osika gets a lot of input and help from him.
- Osika’s underrated picks: the browser companies — Strawberry, Dia, Perplexity, plus one name the audio garbles. Harry says Perplexity building a phone is “a good bet”; would Anton invest at $18bn? “It depends on what options I have” — Harry: “your laugh there just kind of said it all.” On jobs, Harry counters that 8 of the top 10 paying jobs today did not exist 15 years ago and that job displacement is usually overestimated; Anton still expects glamorous knowledge work to go the way of the artist — “being an artist was so cool, but clearly you can’t make any money… we’re going to see that again now for a lot of knowledge work, and that’s going to be funny.”