EN / 中文

The Hidden Battle for High-Quality AI Data: API Relay, Model Routing and Data Distillation in the Industry

by leifengwang·September 29, 2026

Author | Xu Xiaofei, Editor | Liu Wei

Over the past six months or so, a group of domestic AI model vendors have intensively initiated "pre-training rework," with the root cause pointing to the same issue: poor data quality. Among them, some previously spent a year grinding away at trillion-parameter models, only to find their deployment performance outperformed by small models with just 1% of the scale on the market. Ultimately, they discovered that dirty data had held them back.

Some experienced delayed model progress due to engineering flaws such as ambiguous data annotation rules, contaminated evaluation sets, and redundant garbage corpora. Others, constrained by data limitations that led to performance bottlenecks, simply started from scratch and rebuilt their data teams. Meanwhile, intercepting interaction trajectories via API relay stations and performing data distillation have also become a hidden battle beneath the surface of the industry. Data is teaching model vendors a profound and hard-learned lesson.

This lesson is far more hidden than computing power consumption, architecture selection, and talent competition. "The winning hand in the second half of AI is data" sounds like common sense, but it is only today that people truly smell the blood behind it.

01. The "Extensive Detour" in Data

Pre-training rework is, in essence, making up for the missed lessons in data.

The Misconception of "More is Enough"

Zhao Xiang, a data expert at a leading foundation model company, recalled to Leifeng.com that in the early days of pre-training a few years ago, the mainstream view was, "As long as there is enough data, it's fine; a slight drop in quality doesn't matter because the model has its own transformation ability and robustness." Consequently, they began to blindly pile up data extensively, pursuing only scale without regard for quality.

Furthermore, the industry generally lacked a mature large-scale data governance system at that time. In order to assemble a corpus of hundreds of billions or even trillions of tokens, people scraped a massive amount of web garbage and dirty corpora, even mixing in various cleaning packages of unknown origins. The annotation and filtering of much of the data were also poorly handled, resulting in subpar performance of the first trained models.

Only then did people realize that neglecting data quality would leave models stuck at a score of 60 or 70, unable to improve. "Watching overseas models reach a score of 90, we just couldn't catch up," so they had to turn back and redo data optimization.

The Price of Dirty Data

"The harm of underlying dirty data is far greater than imagined. It will cause you to waste the GPUs, money, and manpower you invested, and it will also delay your time window." Wang Bin, a former executive at a model vendor and now an AI entrepreneur, told Leifeng.com. He described a typical case to Leifeng.com. A central foundation model company poured entire batches of scraped tech blogs into the pre-training corpus. It was only halfway through training that they discovered a significant portion of it consisted of machine-generated pages from scraping sites, with the same paragraph repeated tens of thousands of times. The model began inexplicably repeating itself on the evaluation set. By the time the team pinpointed the problem, tens of thousands of GPU hours had already been burned, and that version had to be scrapped and restarted. In fact, such failures are far from isolated in the industry. Many teams only turned back to make up for the missed lessons in data governance after encountering problems with training results.

Inappropriate Data KPIs

Wang Yunhe, founder of Jiyun Lvdong, also mentioned in a media interview that there is nothing new in AI; data is the foundation. If the data is bad, don't blame the computing power; if the data doesn't meet the standards, no algorithm will work, because this violates machine learning theory. He also specifically pointed out that the data team must have a certain say in the pre-training team, and the pre-training team must also make reverse demands on the data. Independent assessments will cause chaos.

In fact, inappropriate assessment of data teams is a pitfall that many model vendors have fallen into. For example, media reports once described that Tencent's Hunyuan team experienced a situation in data governance where "annotation rules were vaguely defined, but the acceptance threshold was set too high. To meet the quota, the team produced a large amount of unusable data." Coupled with the contamination of some benchmark-chasing data and redundant data, they ultimately had to start from scratch, specifically reassembling a team to clean the data from the beginning. "It would be absurd to assess the data team solely based on data volume," added Zhao Xiang, the aforementioned data expert. Because large models compete on the information density and cleanliness of data, not the absolute volume. Assessing only the quantity is equivalent to encouraging teams to use garbage to pad the numbers. Pumping in hundreds of terabytes of dirty data will instead make the trained models dumber, or even cause training collapse.

02. The "Hard Work" and "Clever Tricks" Behind Data

The foundation needs to be relaid, but good data is not something you can just conjure up.

The Historical Accumulation from the Internet Era is No Longer Sufficient

Local life, social networking, gaming, e-commerce, and search—these scenario-specific data were regarded as "barriers" by major tech companies during the Internet era, but they are no longer effective in the AI era. Because foundation models compete on general capabilities, the advantage of data in a specific scenario has limited impact when translated to general capabilities. Although everyone is using their own scenario data to train some vertical small models, the ultimate application scope is not that broad.

"Logically speaking, Google has the most scenario-specific data of various types and should perform the best, but obviously, that is not the case. Vendors like Anthropic and OpenAI, which previously had zero accumulation in data, are actually charging ahead more aggressively, and the same is true domestically," Xu Dong, an algorithm head at a foundation model company, told Leifeng.com. Everyone is standing at the same starting line. At this point, what can widen the gap is: whether you have enough GPUs, how fast you can acquire data, how high the quality is, and whether you have a pipeline that can stably transform data into model capabilities.

No Secret Recipes, Only Hard Work

So, is there really a shortcut for data optimization? Zhao Xiang, the aforementioned data expert, admitted to Leifeng.com that there are actually not many tricks or secret recipes. Data work leans more towards hard and exhausting labor. What matters is to keep at it persistently, investing sufficient manpower and determination to gradually optimize the original 60-point data to 90 points, striving to get a little positive feedback after each optimization. At this time, two accounts need to be calculated: one is the optimization cost, and the other is the training cycle. "There is so much data on the internet; it's impossible to verify every single piece. Training also has its cycle, and we can't wait until all data is optimized before starting training. So we need to pick the important ones first, do it in phases, take small steps and run fast, and iterate continuously," Zhao Xiang added. However, it needs to be clarified that data challenges are not limited to pre-training. The expert data and Agent long-horizon trajectory data required for post-training also involve high difficulty and cost in acquisition.

What Bottlenecks Data Optimization: Money and Capacity

Multiple AI data practitioners told Leifeng.com that currently, the main two factors bottlenecking data work across companies are money and capacity. A counterintuitive point is: big companies don't seem to lack money, but when the budget is broken down, the portion allocated to data is actually not much. "Data does not account for a large proportion of the training budget for model vendors. A rough breakdown circulating in the industry is: for every CNY 100 invested in AI training, CNY 40 is spent on computing power, CNY 30 on talent, CNY 20 on advertising, and the remaining CNY 10 is spent on data," said Zhou Jun, a business head at an AI data supplier.

Taking expert data as an example, Zhou Jun told Leifeng.com that the demand for experts' long-horizon task data is enormous. The budget caliber of several major companies is basically hundreds of millions of CNY, but when it comes to actual payment, they are quite stingy. "For the same expert data, a single piece of research by HLE in the US starts at USD 10,000 to 20,000, while domestically, they basically only offer CNY 1,000 to 2,000."

However, what bottlenecks them more than the budget is capacity. High-quality data has always been scarce. When official channels cannot supply enough, companies start finding their own ways. Zhou Jun told Leifeng.com that some model vendors covertly run relay station businesses, doing Token forwarding. On the one hand, there is revenue, but more importantly, they can collect a large amount of user behavior data in this way. When users invoke various advanced models, the requests and results must pass through this vendor's forwarding layer, so the real task data is left on the platform. This is much closer to the effective samples desired for training than publicly scraped corpora.

Currently, there is also a batch of startups doing Routing (model routing) that are keeping an eye on the asset of "data"; and major companies also have similar routing platforms. However, because they prioritize integrating their own models, it is difficult to accumulate trajectory data from other advanced models. Moreover, due to the incomplete integration of models, the user volume of the platforms themselves is also limited. "In recent years, companies have tried every trick in the book to acquire AI data, but how long these tricks can keep a company ahead is hard to say," Zhou Jun added.

03. The Flywheel Hasn't Started Spinning Yet

There is a voice in the industry that believes large models themselves have no moat; the real moat lies more in the "data flywheel": the stronger the model, the more users are willing to use it to handle difficult tasks, and the more high-quality data flows back. Feeding this data back in makes the model even stronger.

This logic sounds like a closed loop. But as of today, very few have truly got the "data flywheel" spinning. Besides the aforementioned "budget constraints" and "difficulty in acquiring external high-quality data," there is another layer of reasons within the companies themselves. The AI products of major companies generate user behavior every day, but whether this data can return to the training process at scale depends on whether the intermediate pathway has been built. Currently, there are three bottlenecks in this pathway. One is that organizational adjustments are still ongoing, and the dismantling of data walls between departments is still being pushed forward with difficulty; the other is the scarcity of data-related talent.

Multiple AI headhunters told Leifeng.com that the market currently lacks positions such as data cleaning, data pipeline construction, data engineering, and data research. These positions were previously classified as auxiliary roles, and the market demand has only truly picked up recently. Furthermore, more importantly, domestic model vendors also face the problem of a lack of endogenous data. At present, long-horizon trajectory data mainly comes from two scenarios: AI Coding and AI Office. However, a large chunk of the market share in the former has been taken by overseas models like Claude and Codex; the product capabilities in the latter are currently in the initial stage, with limited coverage of work scenarios and complexity, and the accumulated trajectory data is still relatively thin.

04. A Long-term War with No Shortcuts

Returning to the opening statement: "The winning hand in the second half of AI is data." This is broadly correct, but people may have underestimated the cycle and complexity of this battle. Pre-training rework is just the first wave of pain.

Data governance, data acquisition, and the data flywheel—each is a tough bone to crack. Computing power can be bought, and talent can be poached, but there is no shortcut to accumulating high-quality data. It requires time, determination, and systematic investment at the organizational level. Multiple industry insiders admitted to Leifeng.com that in the next year or two, the industry is very likely to see more cases of "starting from scratch." This is not a problem for just one company; the entire industry is making up for lost lessons.

The true divergence will probably become clearer only after this round of data infrastructure is basically completed. Whoever can build the data pathway from user behavior to model training first will have the opportunity to stand out in the next stage. But at least from the current perspective, the end of this road is still far away. Leifeng.com will continue to release in-depth observations on AI data in the future. Friends who are interested in this topic or want to discuss the industry reality are welcome to add the author below for communication and discussion.

* The interviewees in the article, Zhao Xiang, Wang Bin, Xu Dong, and Zhou Jun, are all pseudonyms.