If you think you’ve excluded all Chinese AI from your tech stack, you might want to think again.
The US government’s attitude towards Chinese AI is that it is a threat to national security. That alone could convince patriotic Americans to shun AI labeled “Chinese” in favor of AI without that label. Cisco has released details of new research by itself and VAIL demonstrating that national labels don’t give an accurate picture of a model’s pedigree and features.
In a blog post titled “The US vs. China AI Trap: An Incomplete Proxy for AI Security,” Cisco argued that country labels do not provide a complete picture of the components of an AI model due to a phenomenon it calls origin entanglement.
The researchers used two AI model fingerprinting methods to analyze model weights and behavioral patterns and found that they did not accurately or necessarily reflect the model’s publisher name and country of origin.
They suggest that AI models require their own form of SBOM, but note that there are difficulties: “dependencies are not listed in a manifest file, they are built into the learned weights themselves.” This is because the model maker is likely to refine its AI with existing control points rather than training it from a blank page. In short, it can inherit weights, biases, and behavior patterns from a different model originating from a country other than the one listed on the label. An American model may contain behavior inherited from a Chinese model, while a Chinese model may inherit American characteristics.
The AI tag’s developer and country of origin retain value, Cisco says, providing insight into the responsible developer, applicable jurisdiction and “authorized” procurement process — but not providing an accurate assessment of what might exist inside the model.
The Nemotoron and Qwen models were selected for study because some Nemotoron models are known to use Qwen base weights.
Researchers have searched for surviving links between the two families using Cisco’s Model origin kit and on VAIL Behavioral fingerprinting. The former examines the artifact from the inside, while the later examines the behavior of the inference from the outside.
“Both methods found that the Nemotron models built from Qwen base weights were significantly more similar to the Qwen models than chance would predict.” Their final conclusion is that post-training and a new publisher name do not necessarily erase detectable links with upstream model families.
This is important because the origin of the model can create problems similar to the software supply chain threat for which SBOMs were created. For example, say the researchers, “If an upstream model is later found to contain a backdoor, systematic bias, or exploitable behavior, organizations will need to know which downstream models may require review.”
He suggests three areas that need improvement to effectively use AI:
enterprises consideration of the use of a particular model must “treat the publisher’s identity as one piece of the puzzle.” Due diligence should include provenance, learning dependencies, behavioral analysis and operational controls: label name is not a substitute for potential risk.
Regulators need a better understanding of the model’s upstream dependencies to build a true picture of vulnerabilities, biases, and limitations stemming from the model’s origins.
AI developers should treat genealogy disclosure as routine, not optional. As in all things, transparency is the best disinfectant and would allow users to understand upstream dependencies before integrating a model into their technology stack.
In a geopolitical context, the full model is often known simply by the labeled country of origin. But this can lead to blind spots and false equivalencies.
“A Bill of Materials model can expand the responsible adoption of AI by recording baseline control points, extraction methods, master datasets, synthetic data generators, teacher and reward models, licenses, and entities with post-implementation access. Technical fingerprinting can confirm these disclosures or identify relationships that may require further review. Industry does not have to wait for regulation to do this routinely.”
So, the researchers conclude, while country-of-origin labels matter, they don’t define a model’s technical heritage. “Models don’t have passports. They have supply chains.”
Connected: Cisco releases open-source tool for AI model provenance
Connected: Weights of AI: Securing the heart and soft underbelly of artificial intelligence
Connected: Addictions in Artificial Intelligence: Can AI be trusted?
Connected: AI and Cybersecurity – Everything You Wanted to Know But Were Afraid to Ask