Pioneers Insight Method Research Author
【Special Variety Show—Part 2】 Silicon Valley AI Researchers | The Model Guessing Game and Prompt Battle
Back to Episodes

【Special Variety Show—Part 2】 Silicon Valley AI Researchers | The Model Guessing Game and Prompt Battle

Summary

  • A blind test exposed just how homogeneous model outputs have become: even OpenAI researchers repeatedly failed to recognize their own model among the major labs’ offerings. Six guests relied on stylistic tics—parentheses, em dashes, line breaks, and bullet points—to identify the models. Gemini 3.5 Flash, Claude 4.6 Sonnet, and DeepSeek were repeatedly misidentified; Bessie’s unexpected takeaway was that “it’s actually very hard for everyone to reach consensus on which one is their own model… many models are just incredibly similar.” For investors focused on model differentiation, it offered a direct look at how much outputs are converging.
  • The claim that older models are quietly being made dumber before a new model launches received a qualified internal rebuttal. Wang Sen said, “Based on my personal understanding, that should not be the case.” The deeper explanation is that users’ own expectations are unstable: “Even if it gives you an answer, you might not be equally satisfied with it yesterday and today.” Teams do look for clues in social media and user feedback, and after confirming bad behavior, try to eliminate it through algorithmic iteration; users are encouraged to click thumb down.
  • Voice agents are already operating in the real world; AI replacing most of daily life is still a long way off. Zhou Yichao, who recently joined a voice-agent startup, said that when you call Chase to close a credit card, “there’s a very high probability an agent is already taking the call, and it very likely can complete your task quite well.” But AI replacing most things in daily life “is still relatively far off” and “may still depend on the arrival of robotics.”
  • At the frontier, Chinese and US research has more in common than not. “The frontier model companies doing well are basically in China and the US,” and there is broad consensus on how to approach scaling laws and RL. The real differences may lie in resources—“with fewer resources… people work more meticulously”—and in the communications environment: Chinese teams have more Chinese members and more concentrated communication, while US teams are more diverse and require more alignment.
  • The guests’ AI workflows center on sub-agents, documentation, and starting fresh chats frequently. For problems they cannot solve, they first ask the model to lay out a plan, revise it, then “spin up multiple sub-agents to carry out” the work. Long conversations can send a model “suddenly into a dreamlike alternate reality” after compaction, so they keep records in documents. Rather than maxing out the context window in one chat, they start a new conversation—“faster, and sometimes even more accurate.”
  • The panel split openly on whether to sell after an employer goes public. One camp argues that “at current valuations… they can all be held for the long term”; the other says, “Of course you sell… people have waited so long in private companies… you can’t be all in one place,” selling part of the position to buy other companies. FIRE targets ranged from $2M—enough to return to Chengdu or a wife’s hometown, provided one stops comparing oneself with others—to $10M.
  • The Prompt Battle strategy: use verifiable tasks and as many tool calls as possible. The prompt that got ChatGPT 5.5 Thinking High to think as long as possible was: “Search for the 90 most-cited papers for each year since NeurIPS was founded and summarize them in a doc”—it thought for 6 minutes 24 seconds. The competing prompt, “Query at least 100 webpages and produce a research report on how much money it takes to achieve financial freedom,” took just 1 minute 23 seconds. The core strategy is to set a clear, verifiable objective so the model cannot take shortcuts, produce plausible-sounding nonsense, or turn the question back on the user.

Deep dive

1. The Model Guessing Game: Even OpenAI Researchers Couldn’t Recognize Their Own Model

  • The lineup and rules: six guests—程明昊, who works on evals in a post-training team; 余天呈, also in post-training; 董萌, who works on GPU scheduling; 王森, a memory researcher; 吴月忻 Crick, who works on reasoning and test-time scaling; and 周奕超, who had just joined the voice-agent startup Coficia to work on ASR—were split into two teams. They had to identify the model provider from a blind reading of the prompt and answer, with wrong guesses penalized.
  • The guests mainly relied on surface-level stylistic tics. Some said Gemini Flash was shorter and used highlights; GPT was seen as favoring bullet points, lots of parentheses and line breaks, and highly structured answers; Claude was thought to be especially fond of parentheses. One guest described a Gemini answer as having “not much content, but somehow looking like a lot.” Claude 4.6 Sonnet and Gemini 3.5 Flash were both in the lineup, yet the guests still misidentified them repeatedly.
  • The group’s verdict: “It feels like they all have serious problems—lots of parentheses, unnecessary parentheses, em dashes, and line breaks.” Another added, “I thought only we had this problem.” When the style signals are not obvious, “it really isn’t easy to guess, because every model can produce more or less the same answer.”
  • Bessie’s meta-observation was the most revealing. She had expected OpenAI colleagues to know “at a glance” which answer came from 5.5 Instant. “But in reality, it’s actually very hard for everyone to reach consensus on which one is their own model… many models are just incredibly similar.”

2. The Guests’ AI Workflows: Sub-Agents, Documentation, and Fresh Chats

  • The punishment round produced a practical takeaway: “I use code every day.” When faced with a problem they do not know how to solve, they first ask the model to lay out a plan, make revisions, and then “Spin up multiple sub-agents to carry out” the work.
  • The reason for documenting everything was highly specific. 程明昊 said that after compaction, a model can sometimes “suddenly fall into a dream world” and lose its bearings after a conversation runs too long. He added that compaction sometimes fails; 王森 joked that “my compaction has never succeeded.”
  • 王森’s advice was not to remain in the same chat indefinitely: “Every time you open a new chat, it basically still remembers what you were talking about in the previous message.” Chat messages load faster that way. 董萌 added that the result can even be more accurate because the conversation does not run past the context window. He also said it had been “a long time” since he had manually submitted a job or read a job log; AI handles all of it.

3. Public Misconceptions and the “Dumbing Down” Myth Debunked

  • The biggest misconception is that AI can already replace a vast number of jobs. 程明昊 said the technology is “actually still at a very early stage.” Public opinion is binary: “Either people think AI can do everything, or they think it can do nothing. But today we are actually somewhere in the middle.”
  • When models fail at simple tasks, users may even wonder whether “the company is secretly sabotaging things behind the scenes.” 董萌 added that “the distribution of intelligence across models is extremely uneven”; getting simple things right requires substantial work in its own right.
  • On the rumor that the previous model is made dumber before the next one launches, 王森 said, “Based on my personal understanding, that should not be the case.” People’s standards are unstable: “Your thinking yesterday and today… might not leave you equally satisfied.” Teams draw signals from social media and user feedback; once bad behavior is confirmed, they try to remove it through algorithmic iteration and encourage users to click thumb down.

4. AI in Everyday Life: Customer-Service Calls Are Already Being Handled by Agents

  • The guests already have a range of personal use cases, including weight management, diet, and travel planning. 周奕超 also uses AI to handle airline reimbursement or loyalty-points disputes: he asks it to organize the evidence and request into a formal, well-reasoned letter draft, then sends it to the airline’s CEO.
  • 周奕超 used his new employer to illustrate how far deployment has come. When calling Chase to close a credit card, “there’s a very high probability an agent is already taking the call, and it very likely can complete your task quite well.” He also uses coding agents “all the time,” and some simple tasks “have already been replaced.”
  • But he drew a clear line around the current limits: AI replacing most things in daily life “is still relatively far off” and “may still depend on the arrival of robotics.”

5. Chinese and US Researchers: More Alike Than Different

  • One possible difference is resources. “Sometimes there are fewer resources, and people work more meticulously”; when resources are abundant, teams can run more experiments. Each setup has advantages and disadvantages.
  • The more important judgment was that “the frontier model companies doing well are basically in China and the US.” On how to approach scaling laws and how to do RL, “the consensus is actually very similar in many respects.” Chinese teams have more Chinese members and more frequent, concentrated communication; US teams have more diverse backgrounds and need more alignment. “But basically, I think the similarities outweigh the differences.”
  • The panel left one honest question open: if the same group of people moved to a different place, would the result still be the same? “I don’t know.”

6. The Honest Money Talk: Sell After an IPO or Hold, and What Counts as FIRE?

  • The panel split immediately on whether to sell after an IPO. One side said, “Buying and selling is all about price,” and that at current valuations, several of the larger companies could be held for the long term once they list. The other side was blunt: “Of course you sell… everyone has waited so long in private companies… you can’t be all in one place,” selling part of the position to buy other AI companies.
  • FIRE targets varied widely. One estimate was $6M-$8M to endow a professor position, with investment returns covering the professor’s salary. Someone else said $10M. Another said $2M would be enough for a good life in Chengdu or a wife’s hometown—provided one no longer treats wealth or income as validation of personal worth. Otherwise, “you’ll be comparing yourself with others forever… once you cross one socioeconomic class, you start comparing yourself with another group.”
  • The most Silicon Valley answer: “The top weekend entertainment activity is working.” Do the things you have to do Monday through Friday, then “work some of the jobs you want to work” on the weekend—for example, pursue a bold idea of your own.

7. The Prompt Battle: Verifiable Tasks and Search Take More Thinking Time

  • The rules: use a Chinese prompt of no more than 30 characters and ask ChatGPT 5.5 Thinking High to think for as long as possible. Tasks such as calculating pi to (10^{100}) digits were ruled out because the model would “definitely refuse,” so they were not pursued.
  • The winning strategy was worth noting: “Make it do as many tool calls as possible, such as time-consuming tasks like search… and give it a verifiable task.” A clear target provides a way to measure the result, so the model cannot take shortcuts, return something plausible but misleading, or turn around and ask the user.
  • The result: “Search for the 90 most-cited papers for each year since NeurIPS was founded, summarize them in a doc, and think through it in detail on your own” took 6 minutes 24 seconds. “Query at least 100 webpages and produce a research report on how much money it takes to achieve financial freedom” took only 1 minute 23 seconds. The former was clearly more time-consuming.