Pioneers Insight Method Research Author
[Special Variety Show · Part 1] “From Hot to Flop” | A Model Ranking Game with 7 Silicon Valley AI Researchers
Back to Episodes

[Special Variety Show · Part 1] “From Hot to Flop” | A Model Ranking Game with 7 Silicon Valley AI Researchers

Summary

  • There was no consensus on model quality: the ranking judge put GPT-5.5 at the top, but when the cards were revealed, its label was “top-tier,” not “hot,” so the round ended in “perfect failure.” The answers included Gemini, GPT-5.2, GPT-5.5 and GPT; one guest ranked Gemini “elite” (“The Big Three is always the third one”), while 周奕超 went straight to “total flop”: “What model in the US is worse than it?” He added: “The biggest flop is the one nobody remembers.”
  • The consensus “hot” pick on pay was “Meta TBD,” which the host called “the universally agreed answer”; the discussion centered on a widening compensation split: AI researchers make far more than before, while pay growth in non-AI roles may be constrained. The panel also noted that talent is already moving from quant into AI; if models can do auto research, “they can do quant research too, and may do it much better than humans.”
  • The job-security round was a clean sweep: Apple was ranked “hot,” while XAI was a “total flop.” The host used Apple’s status as the only company that had never laid people off as the deciding factor, though a guest then pointed out that Apple had in fact done layoffs. XAI was dismissed for “cutting too many people,” alongside the rule of thumb: “When in doubt, ban it first.” OpenAI only made “elite,” on the grounds that “there doesn’t seem to be any large-scale turmoil, but everyone just quietly disappears.”
  • On work-life balance, 王森 said he had originally planned to put today’s OpenAI and Anthropic in the “this generation’s total flop” bucket, but switched to Netflix to avoid a collision. The closing reflection was more telling: automation has spent the past 200 years trying to trade more work for more free time, yet “we are doing more work, but we haven’t gained more time for ourselves.”
  • The AI-era skills ranking was the panel’s most internally coherent: empathy was “hot,” physical health “top-tier,” cooking “elite,” Photoshop “NPC,” and coding “a total flop.” 王森’s test was that coding is “highly verifiable”—write good tests and you can check the result—while Photoshop also involves harder-to-verify aesthetic judgment. He said that since January or February this year, he has not opened a code editor and has used agents to write all his code.
  • The strongest capability signal surfaced in the banter: a guest’s card-playing friend sent GPT-5.5 Pro an open conjecture in probability theory he had considered during his PhD nearly 30 years ago, and the model “then proved it.” Friends cross-checked the result and concluded it was correct; a second version is now being posted. The panel called it “a pretty hot weekend activity.”
  • The AI-bubble verdict is a reflexive argument. “What is a bubble? It is only a bubble when nobody thinks it is one. You now encounter someone talking about an AI bubble every day—so how could it be a bubble?” The corresponding insider view is that frontier-lab employees see the progress of the next generation of models before outsiders do, and may glimpse changes 10 days or 1 month ahead—a kind of privileged experience in this era.

Deep dive

1. Model rankings: no consensus, and GPT-5.5 was not “hot”

  • The 5 answer cards were Gemini, Gemini, GPT-5.2, GPT-5.5 and GPT. The ranking judge, 程明昊, put GPT-5.5 at the top and labeled it “hot,” but the reveal showed that its actual label was “top-tier.” The host called it “wrong on the very first one,” and the round ended in “perfect failure.”
  • The guest who wrote “hot” chose GPT for a straightforward reason: “GPT pays me—it sends me money every month on time,” and “everyone at our company uses GPT.”
  • 董萌 put GPT-5.2 in the NPC slot because it has existed for too short a time: “I don’t even remember what happened with it anymore.”
  • Gemini produced the sharpest split. 吴月忻 ranked it “elite” because “the Big Three is always the third one, so it has to be Gemini.” 周奕超 called it a “total flop”: “What model in the US is worse than it?” He added: “The biggest flop is the one nobody remembers.”

2. Pay: Meta TBD was the unanimous “hot” pick, and quant work may also be exposed

  • The ranking judge put “Meta TBD” at the top, and after the reveal the host called it “the universally agreed answer.” Robinhood was ranked “top-tier,” while Amazon and Google landed lower; the respondent specifically meant “the non-DeepMind part of Google.”
  • The serious takeaway from the round was that compensation is bifurcating. Companies that are not doing AI may see much slower pay growth because AI is replacing parts of their work, while AI researchers are making far more than before.
  • The next step in the argument was that many people are already moving from quant into AI, and much of the remaining quant work may eventually disappear as well. If models can do auto research, “they can do quant research too, and may do it much better than humans.”

3. A perfect job-security round: Apple “hot,” XAI a “total flop,” OpenAI people “quietly disappear”

  • 吴月忻 got the entire ranking right: Apple “hot,” Google “top-tier,” OpenAI “elite,” Meta NPC and XAI “total flop.” In making the call, the host cited Apple as “the only company that has never laid people off,” but another guest immediately countered: “Apple has done layoffs too.”
  • The reasoning on XAI was delivered without hesitation: “They cut too many people.” The shorter summary was: “When in doubt, ban it first.”
  • The explanation for putting OpenAI in the “elite” slot was the panel’s sharpest piece of self-deprecation: “There doesn’t seem to be any large-scale turmoil, but everyone just quietly disappears.”
  • That led to a disagreement over the timeline for AI-driven job replacement. One guest put it at “5-10 years,” describing a slow but “very real” process whose consequence would be that “the balance between labor and capital shifts back toward capital.” Another argued that “human inspiration may never be replaceable,” leaving a meaningful gap on lower-certainty tasks. 王森 used his own experience as the example: his job remains, but “the content and nature of the work have changed dramatically.”

4. Work-life balance and skill rankings: verifiable work gets covered first, while empathy is “hot”

  • XAI was the only correct pick in the work-life-balance round. 王森 said he had initially planned to put the current OpenAI and Anthropic in “this generation’s total flop,” but switched to Netflix to avoid a duplicate; he had heard Netflix was “famous for the benefits” and was very chill.
  • The closing reflection was that for the past 200 years, automation has tried to get people to do more work in exchange for more time of their own. Looking back, however, “we are doing more work, but we haven’t gained more time for ourselves.”
  • 王森 got the entire skills ranking right: empathy “hot,” physical health “top-tier,” cooking “elite,” Photoshop NPC and coding “a total flop.” The core logic was that coding is “highly verifiable”—write good tests and the result can be checked—while Photoshop involves aesthetic judgment that is harder to verify.
  • 程明昊 said human relationships definitely will not be replaced, but communication may not be a real skill. He added that sometimes he would “rather talk more with my cat, 豆,” because the cat might understand faster. When the host asked what empathy meant, 程明昊 answered “empathy,” prompting a joke in the voice of a coding model that he does not even understand things that do not require verification. The footnote to physical health being “top-tier” came from another guest: “I feel like our company mainly looks at whether people are physically fit when hiring.”

5. The Bay Area weekend view: working is “top-tier,” and GPT-5.5 Pro proved a 30-year-old conjecture

  • The weekend rankings were revealed as follows: going to the track was “hot”; hiking split between “elite” and NPC; cherry picking was a “total flop”; and working was “top-tier.” There is a track in Sonoma to the north of the Bay Area, while Laguna Seca is a 1.5-hour drive south.
  • Hiking was called one of the Bay Area’s 3 great clichés. Skiing can be “hot,” cherry picking is a “total flop,” and hiking is NPC. The panel also joked that “not working is the real pleasure”; if you are only eating cherries, cherry picking should count as “hot.”
  • 程明昊 said Monday through Friday is for doing things you have to do, while the weekend lets you “work the job you actually want.” He also said he works overtime every day. Another guest said he never thinks of it as overtime: “A person shouldn’t feel ashamed of liking their work.”
  • One guest wanted to preserve a part of life that always belonged to him. Another said he was willing to spend weekends using AI on side projects, turning things that had previously been difficult or even impossible into possible projects.
  • One guest shared that his card-playing friend had considered an open conjecture in probability theory while pursuing a PhD nearly 30 years ago. He recently sent the problem to GPT-5.5 Pro, and the model “then proved it.” Friends cross-checked the result and concluded it was correct; a second version is now being posted. The panel rated it “a pretty hot weekend activity.”

6. Inside OpenAI: the privilege of seeing the future, and why “AI has no bubble”

  • Multiple guests described high talent density and simultaneous disillusionment. Their colleagues are exceptionally smart and passionate about the work, but one guest said that after joining a frontier lab, many things that seem profound turn out to be straightforward at their core. Often, people decide something is too difficult and then talk themselves out of doing it. Another important lesson was to “do simple things well, correctly and reliably.” The most vivid image was an offsite bus—“the loudest bus I’ve ever been on”—because everyone was discussing what they wanted to build.
  • One guest said there is a sharp divide between the internal and external view of a frontier lab. Insiders can see models one generation ahead and detect their progress before the outside world does, potentially seeing new developments 10 days or 1 month ahead. It was described as “a particularly privileged experience in this era.”
  • Another guest had spent 5 years as a quant trader at a fund, where people hid secrets and were reluctant to share what they were working on. OpenAI, by comparison, is relatively open.
  • The reflexive answer on the bubble question was: “AI definitely has no bubble… What is a bubble? It is only a bubble when nobody thinks it is one. You now encounter someone talking about an AI bubble every day—so how could it be a bubble?”
  • The quick-fire round produced 2 additional lines. The public’s biggest misunderstanding of the company’s AI: “Is the previous model made dumber before the next model is released?” The difference between Chinese and US researchers: “The things they think differently about are much fewer than the things they think alike about. People actually agree on many things.”