翁家翌: OpenAI, GPT, Reinforcement Learning, Infra, Post-Training, 天授, tuixue, Open Source, CMU, Tsinghua
Summary
翁家翌 sees the make-or-break factor in frontier-model competition as the correctness, throughput and cycle time of post-training RL infra—not flashier algorithms. Frontier ideas can be generated quickly, but flawed experiments amplify misjudgments; in his view, “the one who fixes more bugs trains the better model.” The real scarcity is the ability to complete more correct iterations per unit of time. What truly put OpenAI on alert about DeepSeek was not a one-off leaderboard lead, but its claim of extremely fast iteration.
In his view, AI labs now need strong engineers far more than they need conventional academic credentials, and when the goal is clearly an industry role, a PhD may even be “a waste of life.” His core distinction is that research intuition compounds with time in the field; ideas are cheap to generate and can even be modeled by AI, but “teaching a researcher to do engineering well is far harder than teaching an engineer to do research well.” Talent and organizational investment should therefore be judged by infra context, end-to-end execution and experiment success rates—not paper counts.
ChatGPT was not an inevitable hit that OpenAI had precommitted the entire company to, but an experiment launched by a small team to improve user interaction and collect real-world data. When 翁家翌 joined in 2022, there was not even a formal pre-training/post-training split; the team expected perhaps 10,000-20,000 early users and planned to shut it down after 5 days, only to see demand grow almost exponentially and servers crash repeatedly. “Everyone was drinking from the same tap” was what convinced the company this was worth sustained investment.
翁家翌 chose to “sell shovels,” allowing his career impact to scale horizontally with every model release. He built OpenAI’s internal post-training RL infra, so his name appeared across releases from ChatGPT, GPT-4, GPT-4V, GPT-4o, GPT-4.5 and GPT-5; he even explicitly wrote down his reward function: “maximize the number of times my name appears on the OpenAI blog.” A single research project cannot scale, but a platform used by every researcher can.
OpenAI’s scale advantage comes with organizational drag: stronger user feedback, compute and fundraising capacity, but also a codebase, communication paths and context-sharing burden that grow more bloated with headcount. 翁家翌 joined as employee No. 280; by the time of the interview, the company had more than 3,000 people. Large institutions do not necessarily iterate fastest; they can only slow the decay with small teams, flat structures and “lossless” information flow up and down. “Everyone is bad, and then you see who is less bad” is his unsentimental summary of competition among mature organizations.
He accepts closed source as a real trade-off required for the company to survive and keep scaling, but he has not abandoned his open-source values. From Tsinghua coursework and 天授 to tuixue, he has treated code as “doing charity” and a way to break information asymmetry. But open-sourcing the strongest model could let competitors exploit the work immediately and then close their own systems, weakening OpenAI’s access to funding and compute. His mission is “achieve AGI first, then benefit all humanity”; the latter need not mean releasing weights and could instead mean putting the capability in ordinary users’ hands through free or low-cost products.
He expects researchers may be replaced by AI before infra engineers, but today’s models remain nowhere near capable of independently maintaining the hardest AI infra. To him, agents with RL post-training are not fundamentally different from standard LM RL; they simply add a few rollout steps and possibly more environments. The real obstacle is that AI infra accounts for almost none of the training data, with long feedback loops and expensive validation. After internal discussions around Strawberry, seeing that “the shit mountain is still there” 1-2 years later convinced him replacement will be a slow hill climb, not an overnight discontinuity.
The deeper reward behind his technical and career choices has been to keep “investing in the future” and create impact recognized by real users, but that answer is now becoming unstable. He sees open-source stars, tuixue’s early 1M-plus and possibly later 10M-plus clicks, and daily model inference as unofficial consensus—not GPA, degrees or titles. At the same time, he believes the world may be deterministic and even wants AI to “predict the future.” As the path ahead for RL infra becomes increasingly clear, he has entered another period of uncertainty, ending with only this: “This question is worth thinking about for a lifetime.”
Deep dive
1. “Slow learning” forced him to build a knowledge tree first, then compress conclusions into intuition
翁家翌 started studying mathematical olympiad problems in first grade. He could often finish mental arithmetic before others were halfway through writing, but new knowledge took him 2-3 times longer than it took other people—including reading code today. His explanation was not that he was smarter, but that “I need more time to build my knowledge tree.”
Once the tree was in place, he would “establish a link directly,” with almost no need to derive things layer by layer in application. The same applied to memorizing texts as a child: before bed he could only stumble through them, but after sleeping he could recite them “backward and forward.”
This learning style created an early preference for time arbitrage: he finished high-school math in middle-school year 2 and started calculus in year 3, not for an imminent exam but because “I want to invest in my own future.” Rather than repeatedly farming points on easy problems, he preferred making an early investment in knowledge with potential long-term returns.
2. OI was an admissions tool, but useless constant-factor optimization exposed his real interest
For students outside Beijing, getting into Tsinghua or Peking University was “as hard as reaching the sky,” so competitions were initially an admissions strategy. Moving further in math competitions required too much early accumulation, so he switched to OI, which he had first encountered in middle-school year 1. He could not even pass the Fujian provincial selection in high-school year 1, but in year 2 he used heuristics to score about 70 points on minimum vertex cover and make the provincial team.
Tsinghua’s summer camp offered an unconditional 60-point reduction. If he reached roughly the top-150 silver-medal threshold at NOI, merely clearing the first-tier university line on the gaokao would secure admission; gold meant direct admission. He ended up as Fujian’s only bronze medalist that year—and the last-ranked member of the provincial team—leaving him to weigh the uncertainty between Shanghai Jiao Tong at the first-tier line and Tsinghua with a 60-point reduction. Encouraged by his family, he chose Tsinghua.
After returning his focus to the gaokao, he continued writing code in secret, even coding and submitting raw in iPad Safari without an editor or compiler. The exercise trained complete modeling and rapid error localization. He also optimized both the runtime constant and code length of algorithms with the same complexity: “It’s useless, technically, but it’s interesting.”
3. Open-sourcing Tsinghua coursework was his first systematic fight against information asymmetry
After entering Tsinghua, the first memory that surfaced was not grades or awards, but uploading his coursework and all the historical materials he had collected to GitHub, except for content with copyright issues. Some older students objected, but he insisted: “I think we should break the information asymmetry.”
翁家翌 saw that many capable students were poor at gathering materials and too afraid to ask teaching assistants, sometimes getting stuck on a low-value assignment for 10 or 20 hours. Publishing past assignments and materials teachers had not prohibited from circulating gave them “information equality” and freed their time for what they actually wanted to do.
Half-jokingly, he said students might not know the names of donors engraved on buildings, but they might well know 翁家翌, who helped everyone survive coursework by making assignments available. Impact comes from actual use, not from position—a logic he would repeatedly apply to his own career.
4. Reinforcement learning was a mis-selection; graphics and security were the original romantic pursuits
When choosing a lab direction in his second year, an older student gave him three names: 朱军, 唐杰 and 崔鹏. 朱军 later offered three directions—Bayesian methods, GANs and reinforcement learning. He wanted to work on images, but had no idea which option corresponded to GANs, so he accidentally picked the second option, RL, calling it “quite random.”
What he had really wanted was graphics. The virtual world in Tron: Legacy stunned him; if he could build his own scenes like a creator, “then I’d be fulfilled.” For a graphics final project, he designed a new algorithm to reduce the number of iterations to convergence, then used substantial compute to render a noise-free 16K image, earning one of only 2 A+ grades in the class.
Network security satisfied another hacking impulse. He and an older student found that Tsinghua’s transcript system could be accessed for one cent or even free, downloaded records multiple times, reported the vulnerability to the academic affairs office and helped fix it. To him, security was hacking a system, while graphics was “hacking the real world.”
He ultimately abandoned graphics not because the interest disappeared, but because he believed research required focus and he could not “keep one foot in 2 boats.” The choice kept him in RL, while quickly showing him that what he truly liked was its engineering foundation.
5. Winning ViZDoom confirmed that he hated alchemy but was good at building tools
His first RL project involved killing enemies, picking up health packs, avoiding obstacles and finishing a fixed ViZDoom map. He eventually won the competition, but found the process “not enjoyable at all.” The environment was too narrow, the objective was to overfit aggressively, and large amounts of heuristics were needed to prevent training collapse and handle corner cases.
Compared with Atari and MuJoCo, the task was simple for humans but required the model to absorb a great deal of implicit knowledge about obstacles and other details. 翁家翌 concluded that much RL work at the time was not algorithmic progress, but repeated tuning around a handful of environments—“possibly 10 or 100 times harder than tuning in CV.”
He had an almost “physiological aversion” to hyperparameter tuning, so he shifted toward helping research run smoothly: refactoring code, designing abstractions and improving user experience, leaving people who wanted to grind on algorithms to do so. He later summarized his preference: “I prefer selling shovels.”
6. Mila’s MoE failure proved that the right direction still loses to compute and engineering gaps
To apply to graduate school, he went to Mila for a 2019 summer research internship through his advisor, working with Yoshua Bengio and the postdoc supervising him. The assignment had nothing to do with his RL background: implement an MoE idea in a Transformer language model, with a router selecting among different paths.
翁家翌 spent a long time filling in the Transformer and NLP context before attempting the implementation, but the result did not work well. In retrospect, he believes MoE required compute, strong engineering and scale-up; with only a few GPUs available to one person at the time, he could not validate that kind of direction.
The host tried to connect RL, NLP and his later OpenAI work as a predestined arc. He cut off the hindsight narrative: “If you really want to say that, you can force it to fit.” Even with today’s knowledge, without sufficient computer and engineering resources, “it still wouldn’t have worked.”
7. Getting only a Master’s helped him escape Tsinghua’s single-track evaluation system
His summer research produced no first-author paper, while classmates returned from programs at Stanford, CMU and elsewhere with results; he was not confident about what his recommendation letters would say. He initially applied for PhD programs but ended up with only a Master’s offer. At Tsinghua, he had been affected by the hierarchy that treated PhDs as superior to Master’s degrees, and went through a period of disappointment and adjustment.
He gradually concluded that educational attainment does not automatically determine a career: “It really depends on what you actually did.” If the demand comes from an employer, the question is whether a new hire can work immediately. Highly relevant project experience can offset years of formal experience.
As an undergraduate, he adopted 3 unofficial metrics recommended by his advisor—papers, competitions and 3-digit GitHub stars—to build his own evaluation system. For GPA, he invested only the minimum time needed to clear the threshold. In a course where 87 counted as a B+, once he had reached a usable standard, he was satisfied: “I didn’t want to spend time on even one more point.”
8. 天授 built its first version in 2 weeks, winning through short code and consistent abstractions
The pandemic and geopolitical uncertainty in 2020 prompted him to avoid becoming absorbed in grand narratives and instead write 天授, tuixue and a visa-search tool from home. 天授 began with a simple question: he already had a large amount of RL experiment code, so why not integrate it into a more usable framework?
He studied RLlib for about a month, but its hundreds of thousands of lines and excessive abstractions made it impossible to see where changes belonged, so he “just rewrote it from scratch.” The first version took only 2 weeks; once the abstraction was right, implementing a paper’s algorithm often took fewer than 20 lines.
The goal was not to publish more NeurIPS papers. He already had papers, competitions and course repositories, and his applications were over, so he stated plainly: “I don’t want to publish papers. I think publishing papers is completely meaningless.” What attracted him was open-source code that was usable, easy to modify and capable of advancing other people’s research.
天授 captured researchers’ real needs: short code, a complete set of common algorithms and an almost uniquely determined place to make changes when adding a feature. 翁家翌 believed the project’s greatest value was consistency. When multiple people each write one piece and assumptions cannot be passed between them, copying, bloat and continual decay follow.
9. tuixue turned code into charity and shaped his impact reward
tuixue grew out of his own need to find US visa appointments. Consulates kept closing and no tool existed, so he wrote a crawler and made it freely available. The first version could even be updated manually twice a day; the technology was not important. He says the project had already received more than 1M clicks early on and may now have exceeded 10M.
After the pandemic ended, the consulate website was upgraded and he no longer had time to maintain it, so the project shut down naturally: “It had fulfilled its mission.” He called both 天授 and tuixue “doing charity”: they could be non-profit or even lose money, as long as the tools genuinely improved people’s circumstances.
In high-school year 3, he suddenly came up with a way to score a life: “If life is a game, the final score is the number of people who remember your name.” This was not a desire to chase fame, but a wish to earn spontaneous recognition through something useful. GitHub stars, clicks, citations and model inference could all serve as forms of that consensus.
10. The real choice in 2022 was DeepSeek’s predecessor versus pre-ChatGPT OpenAI
After entering CMU, he spent a year taking classes online from home because of the pandemic. He initially applied to only about 18 companies and received offers from Google and OctoML. He preferred the latter because he did not want to enter a large company and become a “screw” working on front-end and back-end tasks. OctoML was founded by 陈天奇.
Continuing to interview, he received an offer from 幻方’s nascent AI lab—later DeepSeek—for an RL infra role, not quant work. There were also offers from Nvidia and TikTok; FAIR rejected him for process-related reasons. This was therefore not a choice reconstructed after the fact: it genuinely included “DeepSeek versus OpenAI.”
ChatGPT did not yet exist. He chose OpenAI because it and DeepMind were regarded as the 2 strongest RL research labs, and because he wanted to see how frontier industrial research became methodology—not watch a few PhDs hand-build projects at a university.
John Schulman recruited and interviewed him, placing particular weight on his GitHub and engineering ability. The final round was a 3-hour open-ended end-to-end problem; he finished in 2 hours, then fixed a bug live during the demo. According to the interview, only 2 people had been tested on that problem. The other later worked on Codex, and both passed.
11. For industry roles, differentiated experience is more useful than PhD tenure
翁家翌 never considered continuing to a PhD: “If you want to enter industry, then doing a PhD is a waste of life.” A Master’s can serve as a launchpad; the key is to accumulate enough papers and projects during undergraduate or Master’s study to compete directly with PhD candidates and give employers a clear reason to choose you.
He did not think it was possible at the time to simply declare engineering more important, so he tried to satisfy both the research and engineering sides. But in the large-model era, the frontier increasingly depends on whether infra is correct and how many valid ideas can be verified per unit of time.
He quoted a colleague’s judgment: “Teaching a researcher to do engineering well is far harder than teaching an engineer to do research well.” People immersed in a field for years develop research intuition, and can also discuss ideas with experienced researchers such as Alec; reliable implementation and rapid validation are much harder to learn through a crash course.
12. “Selling shovels” scaled his individual contribution with every model release
翁家翌’s empirical view is that every lab’s infra contains bugs to varying degrees. Differences between models may reflect who fixed more bugs, not just the algorithm. “Llama cannot train as well as GPT because Llama has too many bugs” is only his conjecture; he explicitly says, “I don’t know.”
A single research project serves one result, while infra can be reused by every researcher. He therefore wrote a direct reward function for his career: “I want to maximize the number of times my name appears on the OpenAI blog.” If he wanted his name to appear in a scalable way, everyone needed to use his tools.
He helped build OpenAI’s internal post-training RL infra and sat relatively close to its users, so his name appeared on many major model releases. The span cited by the host runs from ChatGPT, GPT-4, GPT-4V, GPT-4o and GPT-4.5 all the way to GPT-5—an exact map of the “selling shovels” strategy.
13. ChatGPT emerged as a fallback from WebGPT, and its breakout far exceeded expectations
When he joined in July 2022, “post-training” was not yet a formal internal category; there was only the RL team under John Schulman. The team had been working on the next version of WebGPT, but GPT-3.5 browsing required tool calls and performed poorly. They therefore took a step back to improve user interaction, narrowing the problem to chat, instruction following and RLHF.
GPT-3.5 already existed, but the PPO pipeline was difficult to use. In practice, most iteration was happening on GPT-3.5 SFT. The new algorithm first got PPO working on GPT-4 around August 2022; GPT-3.5 was still using the old algorithm at the time.
翁家翌 did not immediately see anything game-changing. His first reaction was simply “a model that can talk.” Later he found it could help with some coding problems, though its ability was limited. Internal exposure had been gradual; the real shock came when external users encountered it for the first time.
The sole purpose of the launch was to collect real-world data. The team imagined 10,000-20,000 users at the start, followed by a demand decline and a shutdown after 5 days. Instead, the curve grew almost exponentially, servers crashed repeatedly and users spread it on their own. Only when he saw everyone discussing it at NeurIPS did he realize the work had set off the industry.
14. OpenAI’s advanced productivity was not genius ideas, but higher experiment frequency
His first impression after joining was of “a large lab,” with no complete methodology of the kind he had imagined—just many people with exceptionally strong research intuition. Only after people such as Barrett and Liam joined from Google did the team begin importing “Google’s advanced productivity.”
The core picture was simple: iterations per unit of time and success rate are approximately positively correlated. Increasing experiments from about 30 per week to about 300 raises the probability of finding a productive direction. 翁家翌 saw this as RL projected onto an organization: trial and error repeatedly until the target is hit.
He agrees that high talent density can allow unexpected results to emerge spontaneously, but larger organizations must depend on information flow: decisions at the top must reach the bottom without loss, and progress at the bottom must return upward without loss. Sam once used research assistants to understand internal progress, while Greg went deep into infra. Managing a company, like managing a codebase, requires consistency.
15. The hardest part of RLHF was not raising reward, but confirming the model had actually improved
The key challenge in early PPO was that nobody knew what the performance curve was supposed to look like. A single reward could rise continuously and then saturate, while genuine human reward might rise first and then fall. The divergence between the 2 is reward hacking.
The highest-reward checkpoint was therefore not necessarily the best one, and sampling-based evaluations across benchmarks had high variance. A score crossing a threshold only showed that it had passed; it was not enough to compare the true interactive quality of adjacent models.
The team ultimately had to pull down the models and talk to them directly, then find more people to try them and vote. The host summarized this as “using human feedback to evaluate RLHF.” 翁家翌 acknowledged: “That’s the only way. There’s no alternative.”
This exposed the difficulty of industrial-scale RL: training correctness, sampling noise, checkpoint selection and final human preference form one continuous chain. An infra error anywhere along it can corrupt the research conclusion.
16. Large-model RL moved the bottleneck from the environment to model compute
In toy tasks such as ViZDoom, model training and action sampling are cheap, with difficulty concentrated in the environment. Large-model RL is the reverse: an environment may provide a prompt in only a few microseconds, while model inference and training can consume hundreds or even thousands of seconds.
The first challenges are therefore how to use more GPUs, sample and train more efficiently, and optimize RL, inference and implementation details end to end.
Around August 2022, he realized that academic RL was still overfitting toy benchmarks such as Atari and MuJoCo while industry was using RL to solve real problems. He gradually stopped developing 天授 and shifted his time to OpenAI’s RL infra: “I should spend more time on things that matter more.”
17. The current paradigm has not finished scaling, so the next-generation infra is being rebuilt from scratch
During the most intense period, he woke up and wrote code, debugged and handled problems until sleep, about 6 days a week, eventually going to the hospital because of headaches. After realizing the pace was unsustainable, he began running 3,000 meters twice a week—a distance he could not pass in Tsinghua PE class.
For the next 5-10 years, he refuses to extrapolate linearly from the present. A new RL paradigm or a new pre-training paradigm may emerge. “Anything can happen”; the fact that infra is the main bottleneck today does not prove breakthroughs will never be needed.
The practical sequence for now is to keep hill-climbing large-scale RL, fix more bugs, extract the ceiling from existing methods and compute, and only then decide what comes next. The biggest bottleneck remains throughput: “How many bugs can you fix per unit of time, and how many times can you iterate correctly per unit of time?”
His team is rebuilding OpenAI’s next-generation infra. The previous generation had been in use for more than 3 years, accumulating substantial technical debt and context inconsistency, so they chose to start over. Researchers state requirements, the infra team handles complex implementations such as distributed training, and the end user ultimately only needs to “change one flag.”
18. Agents did not change RL’s essence, but researchers may be automated first
To 翁家翌, agents with RL post-training are not fundamentally different from standard language-model RL. They simply add a few rollout steps in the middle, with potentially more environments. The mechanism remains entering a modelable environment, receiving feedback and improving behavior along that feedback.
He believes researchers may be replaced by AI before infra engineers: ideas are cheap to generate, while experiment configurations, data ablations and simple for loops are easy to automate. Infra engineers need cross-system context, correctness validation and the ability to handle long feedback chains, so replacement will come later.
Sales may be more resistant to replacement because the buyer remains a person and persuasion depends on interpersonal recognition. Roles like Sam’s are also difficult to replace directly: fundraising, compute, commercial judgment and geopolitical resources attach to an identity recognized externally, not merely to a reproducible technical function.
His private definition of AGI is that if a model can complete 80%—90% of the tasks it considers meaningful, it may qualify as AGI. Today it still cannot safely modify AI infra on his behalf, because such code accounts for almost none of the training data and validation is extremely expensive. It has not crossed that line.
19. Strawberry prompted an overreaction, but “the shit mountain is still there”
Before Strawberry was released, OpenAI had already been using it internally for some time. Users briefly felt their jobs would be replaced immediately and joked that they should write code carelessly first, then let the model clean it up later.
1-2 years later, the reality he saw was that “the shit mountain is still there.” Capabilities had improved, but organizations and individuals often overreact, imagining gradual technological progress as an instantaneous discontinuity. Actual replacement is more likely to be slow and continuous.
He also deliberately downplays his own myth. He was certainly lucky to stand in a unique position, but “if you replaced me with anyone, and they had my context, they should be completely capable of doing the job.” What is truly scarce is not a chosen individual, but context accumulated over time and made usable.
20. Closed source is a game-theoretic choice; benefiting humanity does not mean releasing weights
翁家翌 still loves open source and would participate if OpenAI opened a suitable project. John Schulman even once asked whether the internal RL infra should be open-sourced. 翁家翌 himself argued that it was not appropriate because the company had to survive, raise funding and buy compute in order to continue experimenting at scale.
He divides the mission into 2 stages: “The first is to achieve AGI; the second is to benefit all humanity.” The first requires pre-training, RL, compute and scale. The second can be delivered through products, allowing free users to access ChatGPT, voice and other capabilities rather than releasing a bare model ordinary people would not know how to use.
Asked whether opening the technology could absorb community feedback and reach AGI faster, he acknowledged that the theoretical path exists, but execution creates a strategic game. After the leader open-sources its work, competitors can immediately exploit it, become first, continue training and then close their own systems, while the original company may no longer be able to raise funding.
His preferred meaning of “open” therefore leans toward accessibility for ordinary users, not full transparency to other model companies. If resources were infinite and the company never had to worry about survival, he said he would be “very happy” to open-source the relevant infra.
21. The 2023 governance crisis showed that technical leadership still depends on organizational stability
On Sam’s removal in November 2023, 翁家翌 rejected fantastical narratives such as “what exactly did Ilya see?” and said many claims were based on rumors. The internal facts visible to employees were that some directors did not trust Sam and voted to remove him, while rank-and-file employees had almost no prior information and were extremely shocked.
Many employees preferred Sam because achieving AGI requires more than technology: it also requires fundraising, compute, commercial judgment and vision. A purely technical leader may not be able to coordinate that entire chain. What he most wants to avoid is another near-collapse of the organization.
John Schulman’s departure once made him shut down his computer and feel sad for an entire afternoon, but he still believes that “a healthy organization is one in which everyone can be replaced.” The key is not preventing people from leaving, but continually developing new people and maintaining the ability to regenerate, so the organization can recover like a stem-cell system.
22. DeepSeek’s real alarm was iteration speed, not leaderboard rank
Competition from other model companies usually does not directly change his daily work; the DeepSeek episode was an exception. What triggered concern was not one model overtaking GPT on a leaderboard, but the claim on Twitter that its internal iteration speed was extremely fast. OpenAI’s own iteration was relatively slow and therefore had to work to push its speed above 100.
He sees cycle time as the survival line for foundation-model companies. Data ablations and configuration experiments can be accelerated by adding people or automation, but AI infra requires extensive historical context. That is why adding headcount cannot solve every productivity problem linearly.
OpenAI’s hiring number was about 280; by the time of the interview it had more than 3,000 employees. Scale brings user feedback and more use cases, but also expands the codebase, communication costs and assumptions. A small startup focused on one scenario can easily iterate faster, while struggling to compete with a large company on other dimensions.
He summarized competition among mature companies coldly: “Everyone is bad, and then you see who is less bad.” The theoretical solution is an agent with infinite context that handles the organization’s sharing and decisions, perhaps even becoming CEO. AI’s long context may eventually repair the organizational bloat humans repeatedly create.
23. As the technical path became more certain, he lost the answer to what he wanted
If AI could solve one problem facing the world, 翁家翌 would choose “how to predict the future.” He believes the world may be deterministic and free will may not exist, but retains the central contradiction: “I’m very skeptical, but I try to falsify it and find that I can’t. I really want it to be falsified.”
He has considered that time may not flow linearly: from a higher-dimensional perspective, a future self might send information to the past. The impact reward that appeared suddenly in high-school year 3 could have been an idea that simply emerged in his mind—or one that came from the future. It cannot be falsified. His practical response is to “pretend you don’t know about it” and continue experiencing the present.
Entrepreneurship has not been ruled out, but he has not yet seen a good enough idea and prefers products with clear user feedback. tuixue proved that “the technology is not important; what matters is capturing the need.” His ideal self 10 years from now has enough resources and ability to do whatever he truly wants to do at that time.
His concrete investment in the future is to accumulate enough capital to retire early and preserve optionality. But the RL infra he once loved has gradually become something he can “see the end of,” and he once again cannot figure out what he wants. His final honest conclusion is not an AGI prediction, but this: “This question is worth thinking about for a lifetime.”