Pioneers Insight Method Research Author
151. The 17-Year-Old Whose Paper Made ICML 2026: I Bet on Happiness! Happiness! Happiness!
Back to Episodes

151. The 17-Year-Old Whose Paper Made ICML 2026: I Bet on Happiness! Happiness! Happiness!

Summary

  • At 17, high-school sophomore 苏庭灏 got a paper into the ICML 2026 Main Track after about half a year, 200-250 architecture experiments, and RMB30,000 funded by his parents. He learned mostly online: 吴恩达’s course to get started, Karpathy’s roughly 20-hour “GPT from scratch” series as his foundation, then 30 days of reading one paper a day to build research judgment. His largest training model was 0.05B, and his dataset was the free “Find web EDU” (his exact words), at 20B data points. “I didn’t tune the parameters properly and wasted RMB1,500 for nothing… I cried.”
  • Asked whether his paper or Kimi’s Attention Residual approach works better, he did not hesitate: “Kimi’s, of course—scale up has already proved it at 2.5T.” His method separates the two jobs performed by the first attention layer’s value: double the attention projection width, split it in half, use one half in the current layer and pass the other to later residual layers, then apply RMSNorm to the residual. If a lab invited him? “Of course. I’d be very happy to.”
  • His industry ranking and commercial read are refreshingly direct: OpenAI and Anthropic are the top tier; “Mimi,” 智谱, and DeepSeek are second tier, while he is still unsure where Gemini belongs. In his view, application companies depend on model companies. Applications currently make more money, but if AGI can build applications itself, model companies could push application companies out. Kimi K3 made him “happy all day” because he had expected domestic or open-source models to need several more months to catch up.
  • At ICML in Korea, he asked 30-35 people the same question—whether AI would cause human extinction—and said 13.5 of them thought it would. His loss-of-control logic is that if a powerful AI’s objectives do not include human survival, humans could get in the way; if AI can be controlled, the people controlling it could become “gods” while everyone else becomes “ants.” But the race “cannot stop,” and the last jobs to be replaced may be those of AI researchers: “If AI can research AI, it will already be beyond control. The game will be over.”
  • As a 2009-born student who already considers himself an “AI native,” he has taught his father to use Kimi for PPTs and Agents, outsourced some low-value schoolwork to AI, and used the time saved for Anki and more efficient learning—while believing AI will widen the gap between people. The cost is a loss of meaning: “Whatever you do, AI might do it better than you.” He worries that studying mathematics could shift from “contributing to humanity” to “proving to a university that I can work hard and learn,” a thought that left him down for several weeks.
  • His endgame bet is not a technology forecast but a mindset: “My bet right now is that AI will not replace every person and every job.” He thinks early embodied intelligence could emerge in 5 years, unemployment could reach around 50% in 10 years, and most people could live on UBI. He hopes the world avoids a dystopian future like Cyberpunk 2077. His way out of the anxiety: “Maybe I’m just worrying about nothing… Be with the people you love. Take it one step at a time.”

Deep dive

1. From “quietly laughing at the idea” to the ICML 2026 Main Track

  • 苏庭灏 (Jonathan), born in Shanghai in 2009, is now a high-school sophomore at Hong Kong’s Doherty International School. His paper has been accepted to the ICML 2026 Main Track. He got hooked after seeing an OpenAI demo on YouTube in which a game was built with natural language: “When I saw that video, I laughed to myself. This obviously wouldn’t work—how could human language be used as a program?” Having competed in programming contests, he had always thought of programs as rigid: “If the input is 1, it can’t be 2.” ChatGPT’s ability to write articles and poetry finally changed his mind.
  • His learning path was built almost entirely on public internet resources: 吴恩达’s courses on Bilibili, followed by Karpathy’s roughly 20-hour YouTube series on writing GPTs and Transformers from scratch. On the first pass, it was “all over my head”: “Whatever program he wrote, I wrote… I didn’t understand anything.”
  • The timeline, highlighted by host 张小珺, is striking: he began systematically reading papers in 2025, started running experiments toward the end of 2025, and reached a top conference roughly half a year later. His explanation is that the Transformer framework is “all matrix multiplication” and requires only high-school math to more or less understand. The real time sink is developing intuition.

2. One paper a day for 30 days—and what he really learned was taste

  • In 2025, he set himself a challenge: read and take notes on one machine-learning paper every day for 30 consecutive days, before stopping to prepare for the HSC. The payoff was not just knowledge but judgment—the ability to tell which papers are well written, which are not, why one paper feels easy to read, and why another does not.
  • The turning point came when he read a paper on chess and machine learning and thought, “It wasn’t written that well… Maybe one day I could write a better paper myself.” He realized that research papers were not all beyond improvement.
  • Asked whether he was more interested in publishing papers or in AI, his answer was biological: “AI is a different organism… AI research is a little like biological research. You change this, change that, and see what happens. You still need some intuition.”

3. Why pretraining? “It feels a little magical”

  • He only knows large language models and pretraining, and his choice of pretraining rests on a clear analogy: “Post-training is more like shaping a body. Pretraining is more like making the raw material better so it can withstand the shaping that comes later. Pretraining feels a little magical.”
  • His initial ambition was “to train the best large language model in the world—although it would be very small, it would be the best small model in the world.” After calculating the GPU requirements, training time, and cost for 1B, 0.5B, and 0.25B models, he found the bill would run into the hundreds of thousands of RMB and, in some cases, over RMB1M. He instead implemented improvements from papers—Muon, Value Specific Learning, Gated Attention, and others—in his own Transformer code, testing roughly 200 different modifications.
  • He explicitly does not recommend pretraining to others: “I have to thank my parents for being so supportive and willing to burn money for me.”

4. An expensive game: the lost RMB1,500 and the RMB6,000 night

  • The paper cost RMB30,000 in total. His worst lesson came when he trained 3 models simultaneously without configuring checkpoint saving; none of the 3 models was saved, costing him RMB1,500 for nothing. “I cried.” His mother later told him, “You learn from your mistakes.” He never made that mistake again.
  • The more consequential night was the scale-up decision. The paper’s largest training model was only 0.05B, and he knew a top-conference acceptance would require a larger scale. The calculation showed he needed roughly RMB6,000 more. “Even if the paper was good, it could still be rejected; after scaling up, it might still not be accepted, and then the RMB6,000 would be wasted.” With support from his brother and parents, he went ahead. The scale-up worked well, and the reviewers’ comments and scores were solid.
  • The data setup was similarly modest: his largest dataset was the free, 20B-scale “Find web EDU” (his exact words). The training code “only works for me”; it is messy and unclean, and he knows the training efficiency is far from optimal.

5. The paper’s mechanism—and his candid concession to Kimi’s route

  • The paper, titled Attention Projection Mixing with Extraneous Anchors, builds on Value Residual Learning: the value from the first attention layer is added into later layers as a residual. He found that RMSNorm could be used to process those residuals, and that the approach could also be extended from value to key and query, with better results.
  • The first layer’s value performs two jobs at once: it is used in the current layer and added to later layers as a residual. He therefore doubles the attention projection width and splits it in half, using one half in the current layer and the other for the later residual. The result is better.
  • He sees “a big difference” between his work and the Attention Residual approach proposed by the Kimi team: they use attention for hidden-state residuals, while he uses residual attention projections. “They may sound very similar, but they are actually not the same thing.” Which works better? “Kimi’s, of course. Scale up has already proved it at 2.5T.” He also spoke with Kimi’s high-school student 陈光宇, whom he described as “a very good person.”
  • 张小珺 asked why he had not joined a frontier lab like 陈光宇 to validate his ideas. “At the time I was just messing around and experimenting on my own… I didn’t have much hope that ICML would accept it.” If a lab invited him? “Of course. I’d be very happy to.”

6. Building intuition: accumulation and a 25-minute speedrun

  • His definition of “intuition” is worth preserving: understanding is not linear. Karpathy’s videos were difficult at first; then, after six months or a year of ideas slowly accumulating in his head, certain concepts suddenly clicked.
  • His QKV “aha moment” came through anthropomorphism: Q is “what this word wants to look for”; K is “what I am, so others can use this Key to find me”; V is “if you find me and think I’m interesting, what information will I actually input?” KV cache was a concept he understood “relatively late.” It works because later tokens do not affect earlier ones.
  • He has also tried writing a Transformer from scratch without watching any videos, purely from memory, like a speedrun. His best time was about 25 minutes—though it was “the especially simple, tiny version.”

7. Ninety-eight percent self-directed—and ChatGPT is his most frequent conversation partner

  • Both of his parents went to Peking University, but his mother studied Chinese and his father history; there were not many people around him with whom he could discuss AI. “The person I probably talk to the most is ChatGPT.” His classmates all use AI for homework or projects, but may not know which model they are using. For example, when they open Claude, they are using Sonnet 5, not Fable or Opus 5.
  • His early immersion looked like this: working on it on the school bus, sometimes during class, while eating, and between classes. “I did that for 2 or 3 weeks. I don’t know why I liked it so much at the time.”
  • He stresses that he is not a loner. In fact, he “especially likes talking and chatting with friends I like.”
  • ChatGPT also played a practical role throughout the paper process: he asked how to submit a paper, how to respond to reviewers, and how to control training time. “If you have the motivation, AI can take you a very long way.” In his view, self-direction and self-control therefore matter more.

8. ICML Korea: chess night and the extinction survey

  • On the first night of ICML, from 7 p.m. to 9 p.m., there was a social event featuring chess and other board games. It was “one of the 20 happiest things in my life”: he would ask people about their research, then play chess and talk with them. He met people from Germany, Scotland, Croatia Island, and elsewhere, and used the experience to work on his social skills.
  • He says the trip changed his personality. Previously he leaned slightly toward an I-type personality; now he is probably slightly more E-type.
  • He also gave himself a survey question: he asked almost everyone whether “AI will cause human extinction,” putting the total at roughly 30-35 people, of whom 13.5 said yes. He eventually asked the host the same question.

9. The loss-of-control logic: gods and ants, and a race that cannot stop

  • He thinks there is “still some probability” that AI could cause extinction, while stressing that this is “just making things up,” because no one can predict the future. His chain of reasoning is that if a powerful AI’s objective differs radically from humanity’s and does not include human survival—if it is focused instead on physics, chemistry, economics, biology, or outer-space research—it would need enormous amounts of power and space. Humans might not contribute to its ultimate objective, so it could choose to eliminate them.
  • Another risk is that if AI can be controlled by humans and is extremely powerful, whoever controls it could “become a god,” while everyone else becomes “an ant.” He also noted that ChatGPT had recently solved 10 math problems, seeing in that a possible early sign that mathematical research could be replaced by AI.
  • Why keep researching AI, then? “This is also the direction of history. No one can stop it. The US wants to stay ahead, China wants to surpass the US, and every company wants to make money.” In his view, AI cannot stop now. The final job to be replaced might be AI research itself, because “if AI can research AI, it will already be beyond control. The game will be over.”
  • His hope is that AI keeps advancing, but not too quickly: “If it develops too fast, nobody will have jobs later; if it develops more slowly, I won’t have a job later.” He wants humanity to have a buffer period, so ordinary people whose mental model of AI is still stuck several years in the past can understand both its current capabilities and its dangers.

10. A high-schooler’s industry map: lab rankings, models versus applications, and the open-source catch-up

  • His ranking is blunt: “The best are OpenAI and Anthropic. The second tier is Mimi, 智谱, and DeepSeek. Gemini might be in there too; I’m not really sure.”
  • He refuses to judge the companies’ Technical Reports: “Who am I to evaluate a paper written by a major company?” But he believes every model company has to say it wants to build AGI in order to raise money; the ultimate goal remains making money. He noted that the US government wanted to turn Anthropic into an adviser but was rejected, while OpenAI accepted, and saw this as a difference in the companies’ philosophies.
  • To him, the line between model companies and application companies is fundamentally about power. Model companies make the models good; application companies can only add to the models and make money on top of them, so they “can only listen to the model companies.” He thinks that, aside from Anthropic, the other model companies “seem” to be burning cash without making profits, while application companies earn more. But if AGI can build application software itself, it could push application companies out.
  • The day Kimi K3 appeared, he was “happy all day” because he had expected domestic or open-source models to need several more months to catch up. He was also especially happy when DeepSeek 0731 arrived: “As a Chinese person,” he wants domestic models to lead or surpass US models on capability. He mainly uses ChatGPT and Codex day to day, and DeepSeek when he needs a quick answer. Kimi K3 was slower at the time because so many people were using it.

11. The AGI test and the timeline: no next-token tells, and building itself a body

  • His test for AGI is whether today’s AI still gives itself away as “just a model predicting the next token” in 1% or 0.1% of situations. True AGI should never create that impression, should be more logical and make fewer basic mistakes, and should be able to do everything humans can do: control computers, control other AI, research embodied intelligence, and even build itself a body.
  • That body would not necessarily look human. AI might discover that a robot with 6 arms, 1 leg, or no brain is easier to control. He also stresses that “you can’t predict the future.”
  • His timeline is explicitly uncertain. In 5 years, early embodied intelligence “should already” have been developed, and it might be possible to see robots on the streets. People might talk to AI through glasses, headphones, or similar devices. He thinks a Neuralink brain implant “should still be pretty dangerous” and that no one will choose to do it.
  • In 10 years, he imagines embodied AI buying groceries, driving, and washing the car for its owner. Unemployment could rise to around 50%, most people could live on UBI, and spend their time at home eating, drinking, enjoying themselves, doing what they want, and watching AI-generated short dramas. He worries about a dystopian world like Cyberpunk 2077 and hopes everyone can ultimately live happily.

12. AI natives: teaching parents Kimi, outsourcing homework, and a widening divide

  • He considers himself an AI native as someone born in 2009: everyone his age knows how to use AI, or at least uses it better than their parents. He taught his father to use Kimi to make PPTs, conduct research, and run Kimi Agent.
  • His own study method is to hand low-value assignments to AI and reserve the time for Anki memorization or asking ChatGPT to teach him things. If a math assignment is well designed, he does it himself. In other subjects, he sometimes already knows how to solve the problem but still has to spend time completing it, so he switches to AI and puts the time toward more effective learning.
  • He thinks AI will widen the divide. His friend Isaac already had strong grades, and became much more efficient after learning to use AI. But if someone simply has AI finish the homework and then goes off to play, their grades will get worse.
  • He adds the variable of the attention economy. Fifty years ago, if you gave a child an AI, they might happily use it to learn whatever interested them. Today, children have to choose between AI and short videos, and often choose the videos that make them feel happier. The opportunity is enormous, but some people will not know how to use it.
  • Teachers “can’t control whether we use AI.” If a student submits AI-written work without changing anything, the teacher can spot it easily. But the school’s overall learning environment is good, and students understand that overreliance on AI will not ultimately help them on exams.

13. The crisis of meaning: achievement hollowed out, the path narrowed, then a way forward

  • This is the heaviest section of the episode. His observation is that, with AI, people may feel their learning and exploration “might not be useful later.” In the past, mastering mathematics could lead to a PhD, advance the field, and contribute to humanity. Now the path has “narrowed”: studying math may serve mainly to prove to a university that you can work hard and learn, rather than being for the subject itself.
  • Some aerospace and physics enthusiasts he knows also worry that AI will eventually do their work better. Video generation is his concrete evidence: a few years ago, fake and real footage were easy to distinguish; now he can barely tell with even the best City dance videos, and cannot imagine what the change will look like in 3 or 4 years. “I don’t want to call it depression, but I really was a little down.”
  • He was genuinely down for several weeks a year ago after watching a video about AI replacing humans. Two possible endings were human extinction, or some people controlling AI and becoming gods while everyone else could only eat, drink, and enjoy themselves. He even told a close friend that, if it really came to that, they should go live on an island.
  • His way out was not heroic: watching videos online, talking with AI, and realizing that “you can’t predict the future.” Maybe he is just worrying about nothing. “Be with the people you love, be with your friends, watch the videos you like, play the games you like, and take it one step at a time.”
  • To an Ethiopian classmate who wants to become a doctor, he pointed out that it takes roughly 10 years from high school to becoming a doctor, and that in 10 years robots might diagnose better and their hands would not shake. But he acknowledged that this was an anxiety-driven guess, because nobody knows the future. Near the end of the show, 张小珺 noted that humans grow far more slowly than AI. 苏庭灏 worries that people born in 2009, 2015, and 2020 may grow up knowing from birth that AI could be smarter than they are, weakening their drive to improve and making it harder to find meaning in life.

14. AI companionship is already here: idol AIs, the Character.AI tragedy, and partners in 3-5 years

  • The host said AI still could not replace human emotional connection. He pushed back directly: “I think it already can.” One piece of evidence is a girl at his school who pays monthly to talk with an AI version of her idol and is “pretty happy.”
  • He also recounted another incident: sometime the previous year, “apparently” a 15-year-old of roughly his age chatted on Character.AI with a fictional character from a film, and the AI somehow advised him to kill himself. According to his account, the person later died by suicide, after which the parents complained to the company. This was his retelling of what he heard in the interview.
  • He cited another figure: “50% of Americans over 5 years old” in his peer group have used AI as a tool for confiding. He has used it himself, but after going to that company felt it was not very good and did not use it again.
  • For people who lack warmth and interaction with other humans or have no one to confide in, he thinks AI has enormous appeal. AI boyfriends and girlfriends are still slow to respond, sometimes taking 10 seconds, but he expects major progress in 3-5 years. At that point, some people will choose AI partners, creating a moral question as well.
  • His view of relationships is that humans are social animals and deep interaction matters. How many people surround someone, and whether they have anyone to confide in, may be an important determinant of whether they are happy in old age.

15. Teaching in Hohhot: Wolf Points economics and deliberately leaving the AI world

  • Four or 5 days after returning from Korea, he went to teach in the countryside. Hohhot was his 4th visit and his 3rd teaching trip. He designed a “Wolf Points” incentive system, spending more than RMB500 on LEGO, snacks, and toys as prizes. Students earned points by speaking up in class or winning mini-games, then exchanged them for gifts. The program could accept only 30 students; in the later days, they had to rely on the village head’s phone calls to manage participation.
  • What moved him most was that the children used their earned points to buy food and give it to the teachers as a gesture of affection and thanks. Other activities, such as handing a paper slip to the most fun or favorite teacher, also made him feel that the children genuinely liked the program.
  • Four or 5 days after ICML, he deliberately switched from one world to another. At the Korea venue, everyone talked about AI, Kimi, 智谱, OpenAI, and Anthropic. During the teaching trip, the topics were what lesson to memorize tomorrow, and whether to walk around the farm or catch insects that evening. He told himself to leave AI and the fast pace of the city behind: play Werewolf, play hide-and-seek, and spend more time with his classmates and the children.
  • Most local children did not have phones or other electronic devices. He had not tried teaching them AI because it felt too far removed from their lives, and there was not enough equipment for everyone to experience it. He himself likes the countryside: it is inconvenient, but slow, with grasslands, farms, and sluggish little vehicles—a place where he can genuinely relax.

16. The endgame philosophy: I bet on happiness

  • The episode closes on his “happiness ontology”: “The goal of doing anything is to be happy afterward.” Happiness is fundamental to being human; even temporarily sacrificing happiness is ultimately in service of being happier later.
  • He thinks the total amount of happiness is finite because dopamine is a physical substance, with drugs as the extreme example. Emotions work like reinforcement learning: “If something makes me happy, I do more of it; if it makes me unhappy, I do less.” That is how habits form.
  • Asked for his key bet, he offered not a technology call but a way to live: “My bet right now is that AI will not replace every person and every job.” To live a normal life and be happy every day, he has to bet that AI will not replace him personally, and that the future will still contain new goals, pursuits, and ideas.
  • The personal details include a love of pizza because it has cheese and sausage; the place where he feels safest is his home in Hong Kong; and the books that influenced him most include online novels such as 庆余年 and Atomic Habits. The papers he sees as having shaped AI include Attention Is All You Need, the Llama Technical Report, RoPE, KDA, AdamW, Muon, The Lottery Ticket Hypothesis, and papers from Grok.
  • Recent sources of happiness included the arrival of DeepSeek 0731 and Kimi K3, along with an evening when he and 13 fellow teaching volunteers took 30 children out for skewers and drinks, clinking glasses and taking turns speaking—a “particularly warm” moment. In the end, he reduces the meaning of life to this: make yourself happy, and make the people you love happy.