Articles

ACT 1: Mistral and the Chinese Paradox: What Really Happened?

THE MISTRAL PARADOX: 3 STEPS TOWARD RETHINKING SOVEREIGNTY IN THE AGE OF AI

By Dr. Tawhid CHTIOUI, Founding Presidentof aivancity School of AI & Data for Business & Society; selected by Keyrus as one of the 25 most influential global figures in the field of AI and data (January 2025).

Discover the introduction to the series “The Mistral Paradox” and what it’s all about: Read the introduction


For three years, France’s narrative on sovereignty in artificial intelligence has been based on a fairly simple intuition: to avoid relying exclusively on major American models, we needed to be able to develop our own. Mistral AI has rightly become the most visible embodiment of this ambition. France finally had a player capable of defending a certain degree of technological autonomy against OpenAI, Google, and Anthropic.

On August 11, 2026, however, Mistral introduced a shift that deserves more attention than a mere commercial announcement. The company unveiled a new phase of its “sovereign AI” strategy, based in particular on regional control of inference, control over computing capabilities, and access to open models developed by third parties. The first model offered in this way is GLM-5.2, developed by the Chinese company Z.ai. Mistral’s technical documentation is particularly clear: the model is hosted by Mistral and “served without modification” on its part (Mistral AI, 2026a, 2026b).

We need to be clear about what this means. Mistral is not abandoning its own models, is not becoming a mere distributor of Chinese technology, and is not necessarily giving up its ambition for sovereignty. Its reasoning is actually consistent: a country or a company is considered sovereign as long as it can choose the models it uses, decide where they run, control the computing infrastructure, and retain control over the value produced. According to this view, using a model designed elsewhere does not automatically mean becoming dependent on its designer (Mistral AI, 2026a).

But it is precisely this consistency that makes Mistral's choice so interesting.

For sovereignty is then redefined. It no longer necessarily lies in producing the artificial intelligence that one uses, but in the ability to control the conditions under which it is used. This shift might seem purely technical were it not occurring at the very moment when the global landscape of artificial intelligence is undergoing a major transformation. The 2026 Stanford AI Index notes that the gap between the top American and Chinese models has virtually closed: since early 2025, models from both countries have taken turns at the top of several rankings, while China now holds a dominant position in several areas of AI research (Stanford HAI, 2026).

We are therefore no longer in a world where Europe is simply seeking to reduce its dependence on primarily American technology. A second AI powerhouse is now capable of producing models that are performant enough for a French market leader to deem it rational to offer them to its own customers.

That is where the debate over sovereignty takes on much greater depth.

As long as we thought of digital sovereignty primarily in terms of data, the cloud, and infrastructure, the central question was where our information was located and who could access it. Generative AI raises another question: Who built the intelligence that will interpret it?

Because a model doesn't just provide computing power. It was trained on specific corpora, in specific languages, and according to specific technical choices; it was then adjusted and aligned according to standards, constraints, and trade-offs that are never entirely detached from the context in which it was developed.

The Mistral case, therefore, does not merely raise the issue of our technological dependence on China. It forces us to ask a question we should have been asking long ago about American models themselves: when we entrust an increasing portion of our access to knowledge, our analyses, and sometimes our decisions to a machine, over what do we truly wish to retain sovereignty?

Perhaps we are discovering that digital sovereignty is not just about control over our data and infrastructure. It is also, now, about the representations of the world through which machines learn to respond to us.

To understand what is really changing at Mistral, one should focus less on the nationality of GLM-5.2 and more on the concept of sovereignty that the company is currently developing. In its August 11 announcement, Mistral defined this concept around three capabilities: controlling the models used, choosing where the AI runs, and having the computing power necessary to run it at the desired scale. Its ambition, therefore, is no longer merely to offer French models capable of competing with those from major foreign laboratories; it also consists of providing the environment in which different forms of intelligence can be selected, executed, secured, and tailored to the needs of organizations (Mistral AI, 2026a).

It is in this context that the release of GLM-5.2 should be understood. Developed by Z.ai, an international brand of the Chinese company historically known as Zhipu AI, this model is presented by Mistral as a third-party model intended specifically for code generation and long-context agentic applications. The documentation specifies that it is hosted by Mistral and served “without any modifications by Mistral” (Mistral AI, 2026b).

This clarification is important because it helps avoid two equally misleading interpretations. The first would be to say that Mistral is “becoming Chinese,” which makes little sense: the company continues to develop its own models, its own multimodal technologies, its specialized models, and its computing infrastructure. The second would be to view the addition of GLM-5.2 as a mere catalog detail with no strategic significance. In reality, Mistral is not changing its nationality nor is it giving up on producing artificial intelligence; rather, it is significantly expanding its role. It now aims to be a model designer, a provider of computing power, a deployment platform, and an intermediary enabling its customers to access artificial intelligence developed elsewhere.

This shift is far from trivial. A model lab derives part of its power from its ability to produce intelligence that others use. An orchestration platform derives more of its power from its ability to allow users to choose among multiple forms of intelligence, run them in a controlled environment, and replace them when better ones emerge. Value does not disappear; rather, it shifts along the chain—from the model itself to infrastructure, integration, security, customization, and operational control.

In fact, this trend did not begin with GLM-5.2. A much more technical—and therefore less noticeable—signal had appeared eight months earlier with Mistral Large 3.

Unveiled in December 2025 as Mistral’s most powerful model to date, Large 3 is a Mixture of Experts model with 675 billion parameters, 41 billion of which are activated to process each token. Mistral reports that it trained the model from scratch on 3,000 NVIDIA H200 GPUs (Mistral AI, 2025). It is therefore not a DeepSeek model that has been repurposed, adapted, or retrained.

Its architecture, however, bears striking similarities to that of DeepSeek-V3, released a year earlier by the Chinese lab DeepSeek. Both models consist of 61 layers, use a hidden dimension of 7,168, employ the MLA (Multi-head Latent Attention) mechanism, and retain the first three dense layers before transitioning to a mixture-of-experts architecture. Several attention compression parameters are also identical. The differences, however, are substantial: Notably, Large 3 uses fewer but wider experts, activates fewer of them per token, employs a different positioning system, and adds a visual encoder that is absent from DeepSeek-V3 (DeepSeek-AI et al., 2024; Mistral AI, 2025).

The point is clear enough that Patrick von Platen, a member of the Mistral team, articulated it himself during a technical discussion on Hugging Face: Mistral Large 3 was not built or trained “on” DeepSeek-V3, but its architecture is, in his own words, “heavily inspired” by it. At the same time, he highlights the differences in the structure of the experts, the positioning, and the integration of vision (von Platen, 2025).

This distinction is fundamental because the word “model” often encompasses very different technical realities in public debate.

Architecture describes the overall organization of the system: how its various components are arranged and interact. Weights are the billions of numerical parameters acquired during training, in which much of what the model has learned is embodied. The training data constitutes the informational material from which this learning takes place. The tokenizer, for its part, determines how the text is broken down into units that the machine can process. Added to this are the various post-training and alignment steps that shape the behavior of the final model.

Using a similar architecture does not, therefore, mean replicating the same model, any more than using a building plan leads to the construction of two identical buildings. The materials may differ, as may the equipment and the building’s intended uses. Two models can share much of their architecture yet exhibit very different behaviors because their weights, data, tokenizers, or post-training processes are not the same.

This distinction helps dispel the simplistic notion that Mistral “copied DeepSeek.” The publicly available information does not support such a claim. Instead, it reveals something more interesting: some of the architectural solutions that now underpin a major model developed in France had been experimented with and taken to great lengths by a Chinese laboratory.

In fact, this evolution of architectural approaches must be viewed within the broader history of artificial intelligence. Scientific research has always been built on borrowing, improving, and recombining ideas. The Transformer architecture, which still forms the foundation of a vast majority of contemporary models, itself stems from a paper published in 2017 by Google researchers (Vaswani et al., 2017). It would therefore be absurd to ask each country to reinvent on its own all the technical building blocks it uses.

What deserves our attention is not the existence of this traffic, but the change in the direction of traffic.

For much of AI’s recent history, the leading architectures, training methods, and benchmark models against which the rest of the world sought to measure itself came primarily from American laboratories. DeepSeek has shown that a Chinese laboratory can now, in turn, produce innovations effective enough to become international technical benchmarks. Its MLA architecture, its hybrid expert organization, and its innovations designed to reduce training and inference costs were specifically conceived in the context of severe constraints on access to the most advanced GPUs. DeepSeek-V3 thus sought to turn a hardware constraint into an efficiency advantage, with a model featuring 671 billion parameters, of which only 37 billion are activated for each token (DeepSeek-AI et al., 2024).

The story unfolding between DeepSeek, Mistral Large 3, and then GLM-5.2 is therefore more nuanced than a simple tale of dependence. In December 2025, a leading French company built its own model, drawing heavily on an architecture developed in China. In August 2026, he takes it a step further by offering a Chinese model—unmodified—directly within his infrastructure.

In both cases, Mistral retains control over its own technology choices. But in both cases as well, China is now seen not merely as a vast market, a competitor, or an imitator, but as a source from which technologies are disseminated—technologies that European players deem relevant enough to incorporate into their own strategies.

This is likely the most significant change. The global center of gravity for innovation in artificial intelligence is no longer exclusively American, and the issue of French sovereignty can no longer be viewed simply as a head-to-head competition with Silicon Valley. It must now contend with a China that has also become a producer of standards, architectures, and models capable of circulating far beyond its borders.

Executive MBA in AI & Business Transformation, aivancity
aivancity

Executive MBA in AI & Business Transformation

The MBA Redesigned for the Age of AI. For experienced executives who want to lead the transformation of their organizations. Paris, Nice, and Dubai.

12 months — Part-time At least 10 years of experience Early bird: 20,000 € Paris · Nice · Dubai

The DeepSeek case could still be interpreted as the spectacular emergence of an exceptional player. But that would be to underestimate what is happening in China today. A much broader ecosystem has formed around DeepSeek: Alibaba is developing the Qwen family, Z.ai is working on GLM models, Moonshot AI is developing the Kimi family, while MiniMax is pursuing a particularly aggressive strategy focused on models designed for agent-based applications. These companies differ in their history, resources, and technological choices, but their proliferation shows that China’s rise is no longer based on a single, isolated laboratory.

The figures from the 2026 Stanford AI Index illustrate this shift. While in 2023 the top U.S. models still held considerable advantages on several benchmarks, U.S. and Chinese models have repeatedly traded the top spots since early 2025. In March 2026, the measured gap between the top U.S. model and the top Chinese model was just 2.7%. The United States continues to produce more models considered “notable,” but China now leads in the volume of scientific publications, citations, and the number of AI patents (Stanford HAI, 2026).

Of course, these rankings should be interpreted with caution. Benchmarks measure only a portion of a model’s actual capabilities, and the Stanford AI Index itself highlights their growing reliability issues. But the trend is hard to dispute: the technological frontier of artificial intelligence is no longer exclusively American.

Above all, the nature of competition is changing. A model’s maximum performance still matters, but it is no longer the sole determining factor. The cost of training and inference, the number of parameters actually activated, the ability to access weights, model size, ease of customization, and the ability to run models on user-selected infrastructure are all becoming strategic considerations in their own right.

In this field, several Chinese labs have made resource constraints a guiding principle for innovation. DeepSeek-V3, in particular, demonstrated that a model with 671 billion parameters could activate as few as 37 billion for each token, with a particular focus on training efficiency (DeepSeek-AI et al., 2024). Qwen has adopted a family-based approach, with models of vastly different sizes designed to facilitate their adaptation to varying computational capabilities and use cases (Yang et al., 2025). GLM-4.5, from Z.ai, also combines a Mixture of Experts architecture with a number of active parameters far below its total size, while Moonshot AI’s Kimi K2 and the MiniMax M2 family pursue the same goal of efficiency in coding and agent applications (GLM-4.5 Team et al., 2025; Kimi Team, 2025; MiniMax et al., 2026).

The term “open” must also be used precisely. Not all of these models are open source in the strict sense, particularly because the training data, complete production methods, or certain components often remain unavailable. On the other hand, a significant portion of the Chinese ecosystem relies on open-weights models, whose weights can be downloaded and which, depending on their licenses and technical specifications, can be hosted, adapted, or integrated outside of their developers’ platforms.

This characteristic profoundly alters their geopolitical potential. A company, government agency, or nation that is hesitant to entrust its data to a closed U.S. API can now consider running a high-performance Chinese model locally, adapting it to its specific business needs, or integrating it into its own infrastructure. A study published in late 2025 by the Stanford Institute for Human-Centered AI specifically highlights that Chinese developers are increasingly favoring models that are both efficient and flexible in their deployment, and that their international dissemination could permanently alter the balance of technological dependence (Meinhardt et al., 2025). Qwen, in particular, is already described by the authors as widely used by developers around the world.

This is a significant strategic shift. For a long time, the dominant model for generative AI was to build the most powerful system possible and then attract users to the platform that controlled it. Part of the Chinese ecosystem is exploring a different path: making its models powerful enough, cost-effective enough, and accessible enough to serve as building blocks for systems built by others.

So perhaps China isn't just trying to produce the best model in the world. It's also trying to produce models that the world can easily adopt.

If this strategy succeeds, a country’s strength in artificial intelligence will no longer be measured solely by the number of users connected to a national platform. It will also be measured by the number of companies, developers, and countries that build their own systems using models designed elsewhere. This is what gives the choice of Mistral a significance that far exceeds Mistral itself.

The choice of GLM-5.2 takes on an additional geopolitical dimension when one considers the status of its developer in the United States.

In January 2025, the Bureau of Industry and Security (BIS), which is part of the U.S. Department of Commerce, added Zhipu AI—operating under the name Beijing Zhipu Huazhang Technology—along with several of its affiliated entities, tothe Entity List. This list includes organizations that the U.S. government considers to be involved, or likely to be involved, in activities contrary to its national security or foreign policy interests. In the case of the entities added in January 2025, the BIS explicitly cited their contribution to China’s military modernization through the development and integration of advanced research in artificial intelligence (U.S. Department of Commerce, BIS, 2025).

In practical terms, this designation subjects exports, reexports, or transfers to Zhipu of goods and technologies subject to U.S. regulations to particularly restrictive licensing requirements. For Zhipu AI, the regime outlined by the BIS covers all products subject to the Export Administration Regulations, with a presumption of denial for license applications (U.S. Department of Commerce, BIS, 2025).

The scope of this decision must, however, be interpreted carefully.The Entity List is an instrument of U.S. export control policy. It therefore primarily reflects the U.S. government’s strategic assessment of the risks that an organization poses to its own national security and foreign policy interests. It does not constitute a European ban on the use of Z.ai’s technologies, nor does it demonstrate that the responses generated by GLM would be systematically censored, manipulated, or politically biased. Nothing in the BIS’s decision supports such a conclusion.

This nuance is essential, but it in no way diminishes the power of the symbol.

We find ourselves in a rather unusual situation: a company that Washington considers strategic enough to restrict its access to U.S. technologies is, at the same time, a supplier of a model proposed as part of a European infrastructure that is claimed to be sovereign.

The United States therefore views certain Chinese AI players as a national security threat; Mistral sees one of their models as a technology valuable enough to incorporate into its product lineup.

But this divergence reveals something essential: sovereignty in AI is no longer a purely technological concept. It is becoming a source of geopolitical conflict, in which the same models can simultaneously be viewed as strategic assets to be contained by one power and as technological resources to be integrated by another.

The paradox is therefore no longer just a Franco-Chinese one. It has become a three-way one: a Europe seeking to reduce its dependence on the United States may now be forced to rely on Chinese technologies that the United States is specifically trying to contain.

And this is where the issue becomes even more delicate. For once we accept that a Chinese model can be technically advanced, economically attractive, and hosted on European infrastructure, we must still examine what the infrastructure does not change: the way in which this model has learned to represent the world and respond to it.

The European hosting of GLM-5.2 addresses certain sovereignty issues: data can remain within an infrastructure controlled in Europe, and the organization can choose its runtime environment and reduce its dependence on a foreign platform. But it leaves a more difficult question unresolved: what does a model carry with it when its weights are moved from one country to another?

A large language model is not merely a neutral computer program that is simply run on a server. Its behavior results from several successive layers: the data to which it has been exposed, the languages and cultures represented in that data, the learning that took place based on that data, and then post-training, alignment, and the rules designed to determine desirable, sensitive, or prohibited responses. The infrastructure therefore controls where the model operates; it does not automatically rewrite what the model has learned.

The case of China brings this issue into particularly sharp focus because the regulatory environment there is explicit. Since 2023, Chinese regulations governing generative artificial intelligence services intended for the public have, among other things, required compliance with “fundamental socialist values” and prohibited the generation of content that could undermine national security, the country’s unity, or the image of the state. At the same time, these rules impose obligations to combat discrimination, protect individual rights, ensure transparency, and improve the reliability of content (Cyberspace Administration of China, 2023).

However, it would be wrong to conclude from this that a Chinese model hosted in Europe would necessarily be a propaganda tool or that all of its responses would be politically biased. Regulations applicable to a service deployed in China, the choices made during a model’s training, and the behavior of an “open weight” version running abroad are three distinct matters.

Research is nevertheless beginning to show that these environments leave measurable traces. A study published in 2025 in *Proceedings on Privacy Enhancing Technologies* shows that exposing models to corpora derived from a censored information environment can influence their responses on certain politically sensitive topics (Amed et al., 2025). Other comparative studies have observed, on issues related to China and its relations with Taiwan, that several Chinese models—including GLM and Qwen—align more closely with the dominant positions in their country of development (Lin & Fan, 2024).

These findings deserve to be taken seriously, but they should certainly not lead us to imagine that there is, on the other hand, a culturally neutral Western artificial intelligence.

Numerous studies show precisely the opposite. A study published at ACL 2024 highlights the predominance of English-language cultural references in several major models, even when users query them in other languages (Wang et al., 2024). Other research identifies a bias toward Western references in Arabic-language use, while a study presented at COLING 2025 shows that model behavior varies depending on the cultures against which they are compared and the languages used for their training or adaptation (Naous et al., 2024; Masoud et al., 2025).

Even the political values of American models are not set in stone. A study examining several versions of ChatGPT observed measurable political biases and how they evolved over time, serving as a reminder that alignment is indeed a series of human and technical choices—not a matter of achieving some form of algorithmic neutrality (Liu et al., 2025).

The real question, then, is not whether Chinese models have values while American models do not. It is a matter of recognizing that all major models are culturally situated, even when their global dissemination ultimately obscures that origin.

Perhaps this is where the French debate on sovereignty must take a new step forward. We have spent a great deal of time wondering where our data is stored, under which laws it is protected, and on what infrastructure it is processed. We must now also ask ourselves which systems interpret this data, what representations of the world they have learned, and to what extent we are able to identify, audit, and, when necessary, correct them.

This question becomes all the more important given that these models are no longer used merely to complete a sentence or translate a document. They are being applied in education, research, recruitment, law, healthcare, journalism, strategic consulting, and—even more so in the future—in agents capable of carrying out entire chains of reasoning and action.

As we entrust them with part of our access to the world, cognitive sovereignty does not, therefore, consist in demanding that machines be devoid of culture—which would be impossible. It consists in knowing which cultures, norms, and choices they embody, so that we remain capable of never confusing them with the world itself.

One could wrap up this story with a simple conclusion: Mistral would have abandoned its original ambition, China would have won yet another battle, and French sovereignty would have been undermined. However, nothing we have observed supports such a simplistic interpretation.

Mistral continues to develop its own models while pursuing a different strategic approach: recognizing that its value may also lie in its ability to select, host, secure, and orchestrate the best available AI systems, regardless of their origin. While this strategy is open to debate, it reveals above all that sovereignty in artificial intelligence has become far more complex than the narrative that characterized it just a few years ago.

We used to think we had to choose between relying on American models and building our own. The rapid emergence of the Chinese ecosystem now introduces a third player into this equation. Moreover, the spread of open-source models blurs the line between producing foreign technology and depending on it: a model may have been designed in Beijing or Hangzhou, downloaded and run in Paris, adapted by a French company, and operated without its designer’s involvement in its day-to-day operation. But this operational autonomy does not eliminate all dependencies; it merely shifts them.

Behind the question of where artificial intelligence operates now lies the question of where it was designed, what it has learned, and the choices that shaped it. That is why the sovereignty of the future can no longer be measured by a single criterion of nationality, hosting, or ownership. It must also take into account our ability to understand the models we use, to recognize their limitations, to audit them, to modify them, to replace them, and to preserve sufficient expertise to continue producing them ourselves.

The Mistral paradox thus leads us to a less comfortable but far more fruitful conclusion: technological dependence does not begin simply when we do not possess a technology; it begins when we are no longer able to understand, choose, or replace the technology we need.

That leaves one question that the Mistral case makes it impossible to avoid: How did we get here?

France is not lacking in talent in the field of artificial intelligence. It has internationally recognized mathematicians, engineers, researchers, and laboratories. With Mistral, it has even managed to produce one of the very few European players capable of competing on the global stage in the realm of large language models. Yet, as this competition intensifies, groundbreaking innovations seem to be increasingly concentrated in the United States and China.

This situation may force us to look beyond the quality of our research alone. For there is a vast gap between producing scientific knowledge and transforming that knowledge into industrial power.

It is this distance that we must now understand. And it has less to do with Mistral than with the French and European ecosystem surrounding it.


Don't miss our upcoming articles!

Get the latest articles written by aivancity experts and professors delivered straight to your inbox.

We don't send spam! Please see our privacy policy for more information.

Don't miss our upcoming articles!

Get the latest articles written by aivancity experts and professors delivered straight to your inbox.

We don't send spam! Please see our privacy policy for more information.

Related posts
Articles

The Blind Spot in AI Training: Data Management

By Dr. Guendalina CALDARINI, Program Director, MSc in Data Management | Professor at aivancity | AI education today faces a paradox. Schools and boot camps are training graduates who are capable of refining…
Articles

ACT 3: After Mistral: Who Will Rule in the New AI World Order?

THE MISTRAL PARADOX: 3 STEPS TOWARD RETHINKING SOVEREIGNTY IN THE AGE OF AI By Dr. Tawhid CHTIOUI, Founding President of aivancity School of AI & Data for Business & Society; selected as one of the 25 leading figures…
Articles

ACT 2: France knows how to develop talent. Why does it struggle to build power?

THE MISTRAL PARADOX: 3 STEPS TOWARD RETHINKING SOVEREIGNTY IN THE AGE OF AI By Dr. Tawhid CHTIOUI, Founding President of aivancity School of AI & Data for Business & Society; selected as one of the 25 leading figures…