What is AI model distillation and why is it becoming a US-China flashpoint?

What is AI model distillation and why is it becoming a US-China flashpoint?


For years, the US-China AI race has been fought over semiconductors, computing power and the ability to build increasingly capable AI systems. Now, the competition is moving towards access to the capabilities inside those systems, as Washington grows concerned that rivals could learn from the outputs of leading US models and use that knowledge to build their own AI.At the centre of this debate is a technique known as AI model distillation. It allows a smaller model to learn from the outputs of a more powerful model, giving it selected capabilities while requiring less computing power. It is a standard technique used in AI development.The controversy arises when proprietary models are systematically queried and their outputs are used to train competing systems without authorisation.The issue has now entered the defence domain. Reuters reported that Chinese military-linked researchers used outputs from US AI models in research involving defence applications.The International Institute for Strategic Studies (IISS) has also identified illicit model distillation as a growing issue in the US-China competition, particularly because it could allow compute-constrained actors to acquire useful AI capabilities without developing comparable systems entirely from scratch.

What is AI model distillation

What is AI model distillation

What is AI model distillation?AI model distillation is a method of transferring capabilities from a larger, more powerful AI model to a smaller one. The larger model is commonly described as the teacher, while the smaller model is the student. Instead of training the student to learn everything independently, developers use the teacher’s outputs as training data.In a typical distillation process, the student model is given large numbers of questions or tasks. The teacher model generates responses, which are then used to train the student. Over time, the student learns to produce similar outputs for particular tasks without having access to the teacher’s underlying model weights.The result is not necessarily a copy of the original model. A distilled model can retain selected capabilities of the teacher while being smaller and requiring fewer computing resources. This makes distillation useful when developers want to deploy AI models at lower cost or on systems with limited computing capacity.The concept is not new. The modern knowledge-distillation approach was described by Geoffrey Hinton, Oriol Vinyals and Jeff Dean in a 2015 research paper, which showed how knowledge from large neural networks could be transferred to smaller models.The main advantage of distillation is efficiency. Training and running a large AI model can require substantial computing power. A smaller distilled model can perform selected tasks with fewer computational resources, making it cheaper and easier to deploy.This is particularly useful when an AI system has to operate under hardware or power constraints. Research on knowledge distillation has examined its use on resource-limited devices, including mobile and embedded systems, where large models may be difficult to run.Why does model distillation matter for defence?The defence relevance of model distillation comes from the growing use of AI in military operations. The US Department of Defense has identified AI as a tool for improving battlefield decision-making, including battlespace awareness, force planning and kill chains. It has also emphasised deploying AI capabilities closer to the tactical edge.For such applications, smaller and more efficient AI models can be useful because military systems may have limits on computing power and connectivity. Distillation can reduce the resources needed to run a model while retaining selected capabilities of a larger system. This creates a potential route for deploying capable AI closer to where military operations take place.The importance of AI in defence is increasing as militaries move it beyond analysis and into operational systems. The US, for example, is testing AI-controlled F-16s under DARPA’s VENOM programme, while the Pentagon is developing AI-enabled applications for battle management, decision support and intelligence.

How China is using distilled AI in defence

How China is using distilled AI in defence

How China enters the defence debateThe defence implications are already visible in Chinese military research. A Reuters review of more than 80 Chinese academic papers and patents found that researchers linked to the People’s Liberation Army and other military institutions had used outputs from US AI models to develop specialised systems for defence-related applications.One 2024 paper from the PLA‘s National University of Defense Technology described using model distillation to shrink an image-processing model for unmanned aerial vehicles. The system was designed to analyse live video and support navigation and targeting decisions in real time, including when communications were disrupted.Reuters also found a study by China’s Academy of Military Sciences that used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines.The reported research focused on using distillation to transfer selected capabilities into smaller systems rather than reproducing a complete frontier model.How is AI being used in defence today?AI is already being used across several military functions, particularly where large volumes of data have to be processed faster than humans can manage. The US Department of Defense, for example, has used AI through Project Maven to analyse imagery and video for intelligence, surveillance and reconnaissance, helping identify objects of military interest.AI is also being developed for autonomous military systems. DARPA’s OFFSET programme has focused on autonomous swarms of unmanned aerial and ground systems, while its RACER programme is developing autonomous ground vehicles capable of operating in complex environments.Beyond autonomous platforms, the Pentagon is using AI for decision support, intelligence processing and battlefield awareness. The Department of Defense has said AI can help commanders process information, understand the battlespace and make decisions faster.This makes AI models increasingly relevant to military capability. The issue is therefore not simply whether a country has access to AI, but whether it can develop and deploy models that can process intelligence, support decisions and operate on autonomous systems.

How distilled AI helps on the battlefield

How distilled AI helps on the battlefield

What capabilities can be extracted through distillation?Distillation does not have to transfer an entire AI model’s capabilities. The Center for a New American Security (CNAS) identifies several ways in which outputs from a more capable model can be used to improve another system, including generating synthetic training data, extracting reasoning information, cleaning training data and providing reward signals for reinforcement learning.This means an actor can target particular capabilities rather than attempting to reproduce a frontier model in its entirety. For example, outputs can be used to generate large amounts of specialised training data or improve the reasoning and evaluation capabilities of another model.For defence applications, this distinction is important. A military developer does not necessarily need a replica of a frontier model; it can use selected capabilities to build a smaller system designed for a specific task, such as image processing or target recognition. Reuters’ findings on Chinese UAV and maritime research provide examples of this approach.How can distilled AI help military systems?The main advantage is that distillation can transfer selected capabilities of a larger AI model into a smaller model that needs less computing power. For defence systems, this can be useful where space, power and onboard computing capacity are limited.The military value is clearest when the model is designed for a specific task. Reuters reported that Chinese military researchers used a distilled target-recognition model on tactical hardware during simulated operations involving drones, ships and unmanned submarines.This can be particularly useful for platforms that need to process data locally rather than rely on a constant connection to remote computing infrastructure. In a battlefield environment, reducing the computing requirements of an AI model can make it easier to deploy AI capabilities directly on tactical platforms where bandwidth, power and processing capacity are limited.

Why AI distillation is a US-China flashpoint

Why AI distillation is a US-China flashpoint

Why is Washington concerned?The concern is tied to the US effort to maintain its lead in advanced AI while restricting China’s access to the computing resources needed to develop frontier models.IISS notes that the performance gap between US and Chinese frontier models has narrowed, while the gap in available computing capacity remains significant. China holds an estimated 14% of global AI compute, compared with 74% for the US, according to figures cited by IISS.This makes model distillation relevant to the wider technology competition. According to IISS, training a leading frontier model requires significantly more computing than systematically querying an existing model and collecting its outputs. For a compute-constrained developer, using those outputs to train another model can therefore provide a way to acquire selected capabilities without bearing the full computing cost of developing them independently.CNAS similarly identifies adversarial distillation as one of the ways Chinese developers could seek to overcome the computing deficit created in part by US restrictions on advanced semiconductors and manufacturing equipment.For Washington, the concern is therefore not that distillation can simply reproduce an entire US frontier model. The issue is whether capabilities developed through large investments in computing, data and research can be transferred into another AI system without repeating the same development process.What is the US doing to protect its AI advantage?Washington’s response is moving beyond controlling the hardware used to build advanced AI. In April 2026, the White House issued NSTM-4, directing federal agencies to treat illicit AI model distillation as a national-security concern. The memorandum also called for accountability measures and greater diplomatic efforts to raise awareness of Chinese distillation activity.The US is also strengthening security around the models themselves. A June 2026 National Security Presidential Memorandum directed the national-security establishment to work with private AI companies to protect advanced AI technologies from malicious distillation attacks. The measures include sharing threat intelligence, joint red-team exercises, security research and stronger cyber and physical protection for AI data centres.At the same time, Washington continues to use semiconductor controls to limit China’s access to advanced computing. The Commerce Department has maintained restrictions on advanced computing chips and semiconductor manufacturing equipment, although the policy has also evolved.In January 2026, the US began reviewing licences for Nvidia H200 and AMD MI325X chips for approved Chinese customers on a case-by-case basis, subject to security conditions.As the US and China compete over AI, the ability to protect, transfer and deploy those capabilities could become as important as the computing power used to develop them. For militaries, that could make model distillation an issue not only of AI efficiency, but also of who can turn advanced AI capabilities into usable systems on the battlefield.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *