The global race for supremacy in artificial intelligence has expanded beyond commercial tech firms directly into national security and defense infrastructure. Despite stringent export controls imposed by the United States to restrict China's access to advanced AI chips and cutting-edge technologies, Chinese military research institutions are successfully enhancing their defense capabilities by leveraging the outputs of leading American AI models. An examination of more than 80 Chinese patents and research papers reveals that the People's Liberation Army is actively building specialized, small-scale defense AI systems derived from top-tier US models like those from OpenAI and Anthropic.
Understanding Model Distillation and How It Functions
Central to this strategy is a machine learning technique known as model distillation. In basic terms, researchers query a massive and highly sophisticated frontier AI model, such as ChatGPT or Claude, and capture its responses alongside its underlying reasoning patterns. This output is then used as a training dataset to construct a much smaller, highly efficient AI model tailored for specific tasks. The chief advantage of model distillation lies in its computational efficiency. Developing the smaller model does not require multi-billion-dollar supercomputing clusters or thousands of high-end graphics processing units. The resulting compact model can run efficiently on standard hardware with limited computing power, effectively neutralizing the impact of hardware export bans.
Bypassing US Hardware Export Restrictions
Over recent years, the United States introduced aggressive trade restrictions targeting high-performance AI processors to impede China's military modernization. However, evidence indicates that Chinese defense institutions treat model distillation as an effective workaround. Rather than attempting to train proprietary frontier models from scratch, Chinese scientists are taking a strategic shortcut by harvesting the reasoning capabilities of leading Western systems to power domestic hardware.
Specific Military Applications Across PLA Units
The practical application of distilled AI models spans multiple defense sectors and specialized military branches. PLA Unit 96941, an entity focused on military intelligence and cyber warfare operations, utilized OpenAI's GPT-3.5 to process and summarize complex, sensitive military source code. The insights gained were subsequently used to build a localized model operating strictly within isolated military networks.
Similarly, the National University of Defense Technology developed a lightweight AI model optimized for deployment directly on autonomous drones. This system enables drones to process real-time video feeds, recognize targets, and maintain navigation even when external communications or satellite links are disrupted.
The Academy of Military Sciences applied distillation methods to construct target-recognition models for uncrewed submarines, naval vessels, and aerial drone fleets during combat simulations. Furthermore, North University of China, which maintains close ties to the national defense manufacturing sector, generated synthetic datasets using Anthropic's Claude 3 Haiku to facilitate automated social media monitoring and content moderation.
Extracting Reasoning Chains and Internal Security Concerns
Sunny Cheung, a researcher at the Jamestown Foundation, observed that Chinese defense scientists are not merely gathering raw factual answers from Western AI models, but are specifically extracting their decision-making logic and reasoning pathways. While teaching an AI static facts is straightforward, instilling complex reasoning mechanisms is vastly more difficult, making Western frontier models highly valuable targets for distillation.
Paradoxically, the widespread adoption of model distillation has raised security concerns within China's own military establishment. Researchers from the Army Engineering University published a study focused on data-free distillation, highlighting how third parties can replicate proprietary model capabilities purely by observing model outputs without accessing original training data. The study outlined potential vulnerabilities associated with this technique alongside recommended defensive measures to prevent unauthorized model cloning.



















