Pioneers Insight Method Research Author
Quark on Foundation Models + Gaokao: Bringing AI to Tens of Millions
Back to Episodes

Quark on Foundation Models + Gaokao: Bringing AI to Tens of Millions

Summary

  • This year, Quark’s Gaokao admissions-report model generated more than 10M free application reports, with a single-day peak of over 2M on June 25. Its Q&A product and existing application tools together generate roughly 100M page views a day; Jiang Guanjun confirmed that June 25 was the highest peak in Quark’s history. Quark deployed every GPU it could access internally and still fell short, then borrowed additional capacity from Alibaba Group; tens of thousands of GPUs were serving the Gaokao project, making the exam season Quark’s equivalent of “Singles’ Day.” Host Wei Shijie called Quark “the only company in the industry generating application reports,” a characterization Jiang did not directly confirm.
  • Jiang Guanjun attributed this year’s broader focus on the Gaokao to rising social attention, the emergence of college-application advisers as a new profession, and the verticality of the demand, which allows each company to serve the scenario through its own mix of resources. Wei Shijie added that Agents became a major trend this year and foundation-model capabilities became more mature, moving the technology from experimentation toward real applications; last year, combining foundation models with RAG was still largely exploratory, with hallucinations, reasoning and answer quality below par.
  • “Four nines” refers not to model accuracy but to the coverage rate for university, major and score data. Jiang confirmed that the team had done extensive data preparation around more than 3,000 universities and 2,000 majors. Universities, majors and historical admissions scores come from Quark’s existing application tools; only admission probabilities are predicted, and the outputs are checked manually. Historical data can be archived and reused, with each year’s main work focused on incremental updates to admissions plans, new majors and new colleges.
  • Personalization is harder than accuracy. When the application list cannot be filled, the added schools and majors may not fully match the user’s preferences, forcing the team to balance geography, field of study, employment prospects and graduate-school plans, then calibrate the result against each user’s characteristics through multidimensional scoring. The report strategy is “neither aggressive nor conservative”; Jiang stressed that Quark supports decisions rather than making the final choice for users. In one complaint last year, he said the report itself was not at fault—the student missed admission because the application list was too short.
  • The project’s founding goal was to consolidate fragmented information and steadily upgrade the service. Jiang said Quark had wanted to become a personal intelligent assistant as early as 2018 or 2019, when it saw that Gaokao information was scattered and students struggled to choose schools and majors based on their scores and needs. The product evolved from information services to filtering and probability-prediction tools, and this year to free reports generated by a foundation model. More than 50% of the reports came from students in third-tier and below cities; “information parity” was mainly Wei’s framing.
  • The Gaokao project depends on repeated iteration between data, models and experts. Preparation took more than 3 months and involved several hundred experts, outsourced data-cleaning staff, and Quark’s internal algorithm and engineering teams. The model generates an initial result, experts review it and provide feedback, and the team iterates on the data and model to bring the output closer to expert judgment. The reports have long inputs and roughly 10,000 characters of output, with planning, search, reflection, decision-making and execution in between, making them highly GPU-intensive.
  • Quark and Tongyi Qianwen are collaborating as foundation model and application layer. Tongyi Qianwen provides a strong general-purpose base that helps conserve compute; Quark handles continued pretraining and optimization around product capabilities such as Q&A and RAG, then feeds application problems, data and technical solutions back to Tongyi Qianwen. Wei called it “Alibaba brothers fighting as one,” which Jiang said was a fair description. Quark went from anxiety at the end of 2023 to pretraining from scratch for roughly half a year in 2024, after which it had developed a working methodology and continued iterating on its base capabilities.
  • Foundation models have lowered the barrier to building applications while raising the bar for individual capability. Open-source training frameworks, base models and implementation methods make application development easier, but differentiated user experiences require engineers to judge model capabilities more precisely, fill data gaps and tune training methods. AI writing has no single correct answer, so Quark used model self-learning and structural adjustments to run large-scale RLVR, producing a significant improvement; RLHF uses expert labeling to align models with human preferences.
  • As everyone races to build a super-assistant, Quark is choosing its own path. Jiang described the range of approaches as “many roads lead to Rome” and “a hundred Hamlets,” with Quark organizing around work, learning and daily life and using foundation models to connect its existing tools and capabilities. The goal is to turn workflows users once completed manually into processes driven by a single instruction; the Gaokao report is one example of that upgrade. Jiang offered no clear commercialization plan, saying the Gaokao scenario is not well suited to advertising or commerce and that Quark “hasn’t figured out” how to monetize it.

Deep dive

1. Ten Million Reports: Quark’s Annual Compute Peak

  • Quark’s Q&A product and existing application tools generate roughly 100M page views a day, while its newly launched free application reports have just topped 10M this year. More than 2M were generated on June 25 alone, which Jiang Guanjun confirmed was the highest single-day figure in Quark’s history. Wei Shijie summarized the pattern as: “Every year, the Gaokao peak is that year’s highest; this year’s peak is also the highest in history,” and Jiang confirmed the characterization.
  • A sample Wei obtained ran dozens of pages and contained hundreds of application slots, including AI-recommended rankings, admission probabilities and major descriptions. She framed it as a difficult case combining the depth of Deep Research with a highly specific use case, arguing that while Quark cannot fill out applications directly for users, it now provides extensive decision data and judgment support. Her question and characterization of Quark as “the only company in the industry generating application reports” were not directly confirmed by Jiang.
  • The year’s change was not a simple matter of calling a foundation model. Quark used one to help users produce a relatively professional application report and brought in a large number of Gaokao experts to iterate on it.

2. More Than Half of Reports Came from Lower-Tier Cities: From Information Consolidation to Free Reports

  • More than 50% of the 10M reports were generated for students from third-tier and below cities. Wei cited a survey suggesting that many students—especially those in lower-tier cities—lack decision-support information when filling out applications, while parents and teachers may not know the relevant schools and majors either. She also said that during the “Nuanmang Initiative,” some children did not even know what questions to ask.
  • Wei cited a Gaokao expert who said students have only 4 days from the release of scores to the completion of applications. She interpreted the free report as a way to give students access to information, support decision-making and improve information parity; those value judgments came primarily from Wei’s addition.
  • Jiang traced Quark’s origins in the area to roughly 2018 or 2019, when the company wanted to become a personal intelligent assistant and saw that Gaokao information was fragmented. Students often did not know which schools to apply to, lacked an understanding of majors and industries, and might not be able to identify suitable schools based on their scores. Quark first used its search capabilities to consolidate basic information, then built filtering tools and probability prediction, and finally began generating free application reports once foundation models emerged.
  • Wei noted that Gaokao counseling services in lower-tier areas are expensive and that many people do not even know the profession exists. She therefore saw the free report as a way to provide more sources of information; Jiang did not describe it as a customer-acquisition tool.

3. Why This Year: Social Demand, Vertical Scenarios and Maturing Technology

  • Jiang gave 2 reasons for Tencent, Weibo and other companies increasing their investment at the same time. First, the Gaokao has gained social importance and attention, while the emergence of college-application advisers reflects demand from parents and students. Second, the demand is sufficiently vertical for each company to serve it through its own resource mix.
  • Some companies were already working on the Gaokao in previous years, Jiang said, but they committed fewer resources than they have this year. On competition, he argued that broader participation would drive iteration and give users more channels through which to obtain more accurate information.
  • Wei added her own observation: Agents became a major trend this year, foundation-model capabilities entered a more usable phase, and the technology began moving from experimentation to deeper refinement in real applications. The Gaokao is a concentrated, high-frequency and high-intensity social event, analogous to Taobao’s “Singles’ Day” as a training ground for testing product, team and operating capabilities. Jiang agreed that it was a good setting for integrating teams and building capabilities, but did not include “deep thinking” in his 2-part explanation for why companies entered the market.

4. From Specialized Models to Deep Reasoning

  • For the past 6 years, Quark’s Gaokao information Q&A and admission-probability prediction relied on traditional specialized models; the concept of a “Gaokao foundation model” did not yet exist. Last year, companies began experimenting with combining foundation models and RAG materials to answer questions, but hallucinations and reasoning remained below ideal levels, while data and answer quality were still weak.
  • Jiang said the fundamental change at the application layer this year came from deep reasoning. It will be used in more scenarios, giving models stronger planning capabilities and the ability to reflect on and judge their own prior actions. Deep Research and Agent are both developing on top of these capabilities.
  • Wei summarized the capabilities as deep reasoning, planning and reflection, and Jiang confirmed that all 3 are reflected in the Gaokao model.
  • In Jiang’s example, a student enters scores, province, subjects, employment goals, geographic preferences and graduate-school plans. The model first plans which cities fit the request, then calls search tools or search engines for schools, majors and other information, reflects on whether anything was missed, checks the results, and finally assembles the report. Wei identified this as Deep Research.

5. Data Engineering as the Foundation: “Four Nines” Means Coverage

  • Wei noted that Quark had organized score data and admissions policies covering more than 3,000 universities and 2,000 majors across the country. Jiang confirmed that the team had done extensive foundational data work: organizing authoritative web resources, removing low-quality content and cross-checking sources; sourcing major data unavailable online; and consolidating fragmented admissions data from the Ministry of Education, provincial education departments, universities and other sources.
  • Jiang explicitly corrected the claim that accuracy had reached 99.99%: the figure refers to the coverage of national university, major and score data, not model accuracy. The application form comes from Quark’s existing tool; schools, majors and historical admissions scores are accurate, while admission probabilities are predictions.
  • Jiang confirmed that data is checked manually, but did not say that 99.99% coverage was achieved through manual review alone. Historical admission scores, student counts and score-line ranges can be archived and reused; each year’s main task is to organize the current admissions plans, new majors and new colleges, then align them with the existing data.
  • Wei said continued iteration would create a deepening accumulation of capabilities. Jiang confirmed that the accumulation covers both data and an understanding of user needs.

6. Repeated Iteration with Hundreds of Experts

  • On the model side, Quark strengthened its specialized Gaokao knowledge so that more of it was internalized by the model. Experts then reviewed the model’s initial outputs, identified shortcomings, and repeatedly iterated on the data and model to bring the results closer to expert-level judgment.
  • Preparation took more than 3 months, though Jiang said the manpower was difficult to quantify precisely. External contributors included several hundred experts and outsourced staff involved in data cleaning and professional-materials organization; the internal team included algorithm and engineering staff.
  • One recurring source of friction, Jiang said, was that experts found the requirements demanding and that technical staff and experts spoke different professional languages. Wei said an interview with one application expert could last 2-3 hours, making the process a test of patience, attention to detail and offline communication skills. Jiang’s answer was simple: “Patient communication.”

7. Tens of Thousands of GPUs and Online Reliability

  • The reports have long inputs and roughly 10,000 characters of output, with multiple rounds of reflection, decision-making and execution in between, making them extremely GPU-intensive. Quark deployed every GPU it could use internally and still came up short, then borrowed a large number of cards from Alibaba Group; tens of thousands of GPUs were serving the Gaokao project.
  • Jiang said the Gaokao period is Quark’s largest annual draw on resources, effectively its “Singles’ Day,” with the core challenge being how concentrated demand tests peak capacity.
  • The engineering team ran extensive stress tests and carried out engineering and system optimizations across the Gaokao workflow to keep online service relatively stable during the several days of application filing. Wei observed banners in the war room and signs on the desks that colleagues had concentrated their efforts on the project.

8. From Anxiety to “Knowing What We’re Doing”

  • Jiang divided Quark’s foundation-model development into several stages. At the end of 2023, the entire Chinese internet industry was anxious; by mid-2024, Quark had started pretraining from scratch. After roughly half a year, the model’s performance was still fairly ordinary, but the team had learned the methodology and understood that continuous iteration could produce better results.
  • From late 2024 to early 2025, the team focused mainly on improving the model’s base capabilities, while foundation models gradually began going live across Quark. Because Quark wanted to become a personal assistant, it needed to rebuild its existing products and businesses around the foundation model.
  • Wei described the capabilities Quark developed during this process as full-stack foundation-model capabilities. Jiang agreed.

9. “Existing Business Plus Foundation Model” vs. “Foundation Model at the Core”

  • General search ranking still uses Quark’s existing technology stack, but foundation models can generate samples and perform semantic classification, optimizing different parts of the existing business. That is “existing business plus foundation model.”
  • Deep Search outputs answers directly to users, so the retrieval chain and Query CoT must be rebuilt around the foundation model. That is a foundation model at the core.
  • “Super Box” has a historical explanation. Early AI products were mostly specialized models: one tool corresponded to one model, and one search function corresponded to one model, leaving Quark with a large collection of vertical models. Users can now input text, voice, images, video and documents; the central trend is multimodal input and multimodal output.
  • Asked whether vertical models would eventually be consolidated into one foundation model, Jiang first confirmed that there would be “many vertical models,” then emphasized using a foundation model as the core to connect existing tools and capabilities into a Super Box that can cover different needs. The discussion did not explicitly confirm consolidation into a single foundation model.

10. Personalization Is Harder Than Accuracy

  • The application tool mainly presents probabilities for users to screen themselves, while the report may need to provide 80-100 choices. If the list cannot be filled, the system must add schools and majors, but those additions may not fully match the user’s needs, forcing trade-offs among geography, field of study, employment and graduate-school priorities.
  • Jiang said personalization is harder. Accuracy can be improved through data and accumulated legacy tools, but there is no objectively correct path for every student’s preferences. The team combines the user’s characteristics—for example, recommending schools and majors with stronger STEM capabilities to a STEM-oriented student—then scores the report across multiple dimensions and iterates toward the highest-scoring version.
  • Wei recalled that the war room was effectively in debate every day: after producing one version, the team would decide it was “not quite good enough,” send different members to speak with different experts, and then return to make further adjustments. “The experts themselves didn’t agree” was how Wei described the process.
  • Regarding one complaint last year, Jiang said the report itself had no problem; the student was not admitted because too few choices had been submitted. Quark follows a relatively balanced strategy—neither aggressive nor conservative. It does not make the decision for the user, but provides a more professional and accurate reference for the user to judge.

11. Division of Labor with Tongyi Qianwen

  • Wei noted that the project used Tongyi Qianwen’s base model and continued pretraining. Jiang explained that Tongyi Qianwen provides a strong general-purpose foundation that helps Quark save compute, while Quark places product- and business-specific training in the continued-pretraining stage and optimizes capabilities such as Q&A and RAG.
  • Quark also feeds problems, requirements, foundational data and model-optimization proposals from application scenarios back to Tongyi Qianwen, allowing the latter to abstract which improvements are needed in the base model.
  • Wei called the collaboration “a long-awaited, rousing Alibaba brothers’ team brawl” (“久违且振奋人心的阿里兄弟打群架”). Jiang replied: “That’s one way to put it.”

12. Details and Innovation: RLVR, RLHF and the Application Barrier

  • Jiang agreed with the view that “everyone can train models; the difference is in the details.” Building a strong application requires sustained refinement of data, RLVR and RLHF samples, training methods, strategies and the fit with product requirements. But a major breakthrough in base-model capabilities requires more than fine-tuning: it also requires new methods, new architectures, innovative experiments and accumulated experience.
  • AI writing is a large-scale use case. Traditional RLVR works well for questions with standard answers; in writing and other open-ended scenarios, Quark experimented with using model self-learning, training and structural adjustments to conduct RLVR at scale, ultimately producing a significant improvement. Jiang stressed that this is not unique to Quark; other companies will do it as well.
  • Jiang explained that RLVR has the model generate multiple answers for questions with standard answers, compare them with the correct answer, and improve through automated reinforcement learning. RLHF has experts or experienced practitioners label, from several broadly correct answers, the version that best matches human reading preferences—for example, an overview–detail–overview structure—thereby aligning the model with human preferences.

13. Has Algorithm Development Become Easier or Harder?

  • Jiang said foundation models lower the barrier from an application perspective: training frameworks, base models and implementation methods are widely open-sourced, allowing small teams to quickly build creative products. But the bar for individual capability has risen, because otherwise it is difficult to create differentiated products or a superior user experience.
  • The way technical staff work has not changed much. The core task remains translating product-design language into a technically feasible scope, explaining what can and cannot yet be done, and strengthening capabilities by adding data or adjusting training methods. In the end, product and technical capabilities must meet at their feasible intersection.
  • On team management, Jiang emphasized 2 points. First, teams need composure amid fast-moving technology and must distinguish fundamental model capabilities from marketing claims, while aligning on whether the focus is pretraining, PPO or business adaptation. Second, they need a clear experimental path—knowing how to validate a hypothesis and what counts as success—because any single technical investment can expand without limit.
  • On team morale, Jiang said Quark “doesn’t seem to need much” dedicated emotional management. The team is relatively stable, trust is high among members and between operations and technology, and the shared goal of building a personal assistant has been consistent since 2018. The work can be arduous, but there is little need for dedicated morale management.

14. Different Paths to a Super-Assistant and What Comes Next

  • Facing the Agent and super-assistant boom, Jiang said “there are many roads to Rome” and “a hundred Hamlets” to describe the different paths. Quark will focus on work, learning and daily life, using its existing tools and product accumulation to turn workflows users once stitched together manually into intermediate chains completed by a foundation model from a single instruction. The Gaokao report is one example of that demand upgrade.
  • Asked whether the product would be commercialized over the next few years, Jiang offered no clear plan. He said the commercial logic of the foundation-model era should change, while the Gaokao scenario is not suited to advertising or selling goods; Quark “hasn’t figured out” how to monetize it. Wei suggested a model similar to tipping, which Jiang said could be discussed internally.
  • Retention among Gaokao users is fairly strong. Some younger users have already used Quark’s learning products before filling out applications and may use its learning-materials products after entering university. Jiang sees an opportunity to continue serving users across different life-cycle stages through different scenarios.
  • For next year, Jiang said this year was only an initial effort around more complex decisions, and Quark remains far from fully solving users’ application needs. The next steps are to better guide students who do not know how to ask questions and have almost no relevant knowledge or information, improve model interaction, provide more professional experience, and generate genuinely more personalized reports—reducing the team’s repeated internal debates over application trade-offs. Wei summarized the approach as “iterating year by year,” and Jiang agreed.