The dispute over Claude distillation looks like another technological fight between the United States and China. But underneath it sits a more interesting question: what happens when artificial intelligences begin to modify the very information environment from which the next generation will learn?
Claude answers a question. On the other side, there is not necessarily a person trying to solve a problem. There may be another artificial intelligence taking notes.
There is something strange about the scene because it looks less like a conversation than one of those situations in which someone walks into a classroom, sits in the back row, and copies diligently enough to reconstruct much of what the professor taught. The difference is that this student can ask millions of questions, does not get tired, does not need to sleep, and can repeat the operation until it discovers not only what the professor answers, but something about how the professor tends to answer.
At a certain scale, it even becomes difficult to keep calling it a student. It is more like a silent crowd entering through one door, asking questions for weeks, and leaving through another carrying tiny pieces of someone else’s intelligence. Nobody sees the pieces. They weigh nothing. They set off no alarm at the airport. Yet on the other side they may eventually become another machine.
That, simplifying considerably, is what sits behind one of the most interesting disputes now moving through the AI industry.
Anthropic said on September 10 that it had detected operations by several Chinese companies aimed at extracting capabilities from Claude. The largest case it attributed to Alibaba involved more than 151 million fraudulent interactions, allegedly intended to improve Alibaba’s Qwen models. Reuters reported the allegations while making clear that these are Anthropic’s claims, not judicially established facts.
China rejects the way the United States is describing the phenomenon. Its Commerce Ministry argues that distillation is a standard technique in artificial intelligence and accuses Washington of turning a technical and commercial issue into a tool for restricting competitors.
And this is where the argument becomes interesting, because both statements can contain part of the truth.
The student who learns from the professor
Distillation did not appear with DeepSeek, nor is it a clandestine invention from some Chinese laboratory.
The basic idea is to use the outputs of a powerful model to help train another model. A large system can act as a kind of teacher, producing examples, evaluations, or responses that a smaller system can then use.
The conflict begins when a competitor obtains another model’s outputs at massive scale, through access the owner considers fraudulent, in an attempt to transfer its capabilities.
At that point, the professor analogy starts to break down.
Copying a classmate’s notes does not automatically turn a student into the professor. The student may memorize answers while still failing to understand other things. Something similar is true in AI: obtaining millions of Claude outputs does not mean copying its weights, knowing its original training data, or reproducing exactly how it was trained.
But the outputs are not merely notes either.
They are manifestations of capabilities produced through vast amounts of compute, data, research, and engineering. If another model can absorb part of those capabilities at a fraction of the cost, a difficult economic and legal question appears:
At what point does learning from another model stop being learning and become technology extraction?
The answer is not as simple as the political fight may suggest.
The strange part comes later
Suppose Qwen learns partially from answers produced by Claude.
Then someone uses Qwen to write a technical document.
Someone else uses DeepSeek to program a library.
A company generates millions of product descriptions with GPT.
A student revises a paper with Gemini.
A journalist uses an AI system to reorganize notes and eventually publishes the article.
All of this slowly mixes back into the internet.
It does not happen at once. Nobody pulls a lever and converts the web into synthetic content. It is more like pouring ink into a river. At first you can still distinguish the water. A few miles downstream it becomes difficult to say which part came from the mountain and which part came from the bottle.
It is worth taking a small detour here.
A traditional library lets us imagine a certain genealogy. One book cites another; that book cites a paper; the paper points to a researcher. Follow the path patiently enough and it is still possible to arrive at a person who once sat at a desk, in a laboratory, or inside an archive.
The internet was never that orderly, but it retained something of the same logic.
Now imagine a library in which some books were written after reading other books that had themselves been summarized by machines trained partly on earlier versions of the same library. The books remain on the shelves. They still have authors, titles, and dates. Only at night, some sentences seem to have changed families without anyone knowing exactly when it happened.
Intellectual genealogy begins to form circles.
There is no need to imagine this as a catastrophe. It is enough to recognize that the circuit has changed.
Back to Claude.
If some Claude outputs help improve Qwen, and content later produced with Qwen enters the web, a future generation of models trained on web content may end up learning indirectly from Claude through Qwen.
And Claude itself could one day learn from information that counts an earlier Claude among its ancestors.
It is a strange genealogy: a machine finding an earlier version of itself among its grandparents.
That is where reflexivity enters.
When observing the world also changes it
George Soros used the concept of reflexivity to explain a phenomenon in financial markets.
Market participants try to understand economic reality in order to make decisions. But their decisions then modify the very reality they were trying to understand.
If enough investors believe a company has an extraordinary future, they buy its shares. The price rises. That higher price may make it easier for the company to raise capital, hire better people, or acquire competitors. The original perception ends up altering some of the fundamentals on which the company will later be judged.
Reality changes beliefs, and beliefs change reality.
Something similar, though not identical, is beginning to appear in artificial intelligence.
A model learns from the information environment.
Then it produces information.
That information changes the environment.
The next model learns from the changed environment.
And produces information again.
AI → information → AI → information.
It looks like a simple arrow. But repeat it enough times and it stops looking like an arrow and starts to resemble one of those country roads that, after several hours, returns the traveler to the same tree they were certain they had never passed before.
The object being studied no longer stays still while we build the instrument meant to understand it.
This does not mean AI will inevitably become stupid
There is an easy temptation here: conclude that if machines begin learning from other machines, they will inevitably deteriorate until they produce nonsense.
There is a real problem related to that idea, but it needs to be handled carefully.
In 2024, researchers published a paper in Nature examining what happens when successive generations of models are trained indiscriminately on information generated by previous models. They described a phenomenon called model collapse: less frequent parts of the original distribution begin to disappear, and the model can progressively drift away from the real data it was supposed to represent.
An intuitive way to imagine it is to photocopy a photocopy and then photocopy that one again. The major lines survive for quite a while. The small details disappear first.
Until one day the grandmother in the photograph is still there, but she has lost a wrinkle, then a mole, and finally the expression that made it possible to recognize her as herself rather than as just another woman.
That is precisely the worrying part: the most obvious things are not necessarily the first to disappear. The rare cases, the tails, the unusual details, the small irregularities that are also part of reality can go first.
But it would be a mistake to conclude that all synthetic data are bad.
AI-generated data can be deliberately used to train specific capabilities, create scarce examples, generate problems with verifiable solutions, or improve smaller models. Later research has explicitly explored how synthetic data can be used while avoiding or reducing these degradation processes.
The distillation Anthropic is alleging and model collapse are not the same thing either.
A distillation process can carefully select high-quality outputs from a stronger model. The model collapse studied by researchers appears in recursive training cycles where synthetic information increasingly replaces or distorts the original distribution.
The connection between the two phenomena is elsewhere.
Both show that the output of one AI can become raw material for another AI.
And that raw material is growing.
Human data may become more valuable, not less
Something counterintuitive happens here.
If generating text, code, images, and synthetic data becomes almost free, it may seem as though data themselves will become less valuable.
The opposite may happen with certain kinds of data.
An authentic conversation between people.
A sensor measurement.
A laboratory experiment.
A properly anonymized clinical record.
A survey.
A business record.
A photograph whose provenance can be demonstrated.
A document written before the explosion of generative models.
Anything that can prove it came directly from the world begins to acquire a new property: it is outside the synthetic loop.
Perhaps we will start treating information the way we treat certain old objects. Not because they are necessarily better, but because we can prove they were already there before the machine arrived.
A newspaper printed in 1998 may contain no better ideas than one written in 2030. But nobody will need to wonder whether its paragraphs were produced by a model that learned from another model that had summarized a page written by a third system.
Provenance starts becoming a quality of its own.
The authors of the Nature paper make a related point explicitly: as model-generated material becomes more common online, data from genuine human interactions may become increasingly valuable for training future generations.
It is a small paradox.
The cheaper information becomes to manufacture, the more expensive information with demonstrable provenance may become.
So who had the idea?
There is another, even more uncomfortable consequence.
Suppose a model discovers a particularly elegant way to solve a mathematical problem.
The solution appears in thousands of responses.
Some of those responses are used to train other models.
Those models discover variations.
The variations appear in documents.
The documents feed new systems.
Three generations later, a model produces a similar but significantly better idea.
Whose idea is it?
The question sounded philosophical until models began to participate in real scientific research.
This week OpenAI published a solution generated by one of its internal systems to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. The publication includes a formalized proof in Lean and a section on concurrent work and scientific responsibility.
We do not need to settle the question of scientific authorship here to notice what has changed.
Our system of intellectual attribution is built around a relatively simple genealogy:
someone learned something → thought something → produced something.
AI introduces new branches:
someone produced something → a model learned from it → another model learned from the first → someone used that second model → something new appeared.
Finding the ancestor of an idea may become considerably more difficult than finding the author of a document.
Perhaps an idea of the future will arrive at an office without knowing who its father was. It will have documents, dates, versions, and hundreds of digital traces. It will be missing only what once seemed simplest: a family history.
The dispute with China is only the first layer
That is why reducing the Anthropic episode to “China is copying American AI” wastes the most interesting part of the story.
There is a commercial dispute.
There is a geopolitical dispute.
There is a legitimate argument about intellectual property, terms of service, and access to models.
There is also a race in which the United States is trying to preserve a technological advantage while China tries to narrow it, and in which model efficiency has become as important a variable as raw compute.
But underneath all of that, something else is happening.
Models are no longer merely consumers of the knowledge accumulated by humanity.
They are also becoming producers of the epistemic environment in which their successors will be trained.
Somewhere inside a data center, an answer produced today may sleep for a few months, reappear transformed inside another model, and return years later as if it were a new idea. Nobody will hear the journey. Servers do not keep ghosts, at least not in the traditional sense. But they are beginning to keep genealogies.
That changes the question.
Until now, we have worried about what artificial intelligence would learn from us.
The next question may be considerably stranger: what will artificial intelligence learn from a world that has already been partially written by itself?
References
- Reuters — “Anthropic disrupts Russian, Chinese AI campaigns targeting its Claude models,” September 10, 2026.
- Associated Press — China’s response to U.S. accusations over AI distillation, September 9, 2026.
- Shumailov et al. — AI models collapse when trained on recursively generated data, Nature, 2024.
- Zhu et al. — How to Synthesize Text Data without Model Collapse?, ICML 2025.
- George Soros — writings on fallibility and reflexivity.
- OpenAI — On the Navier–Stokes Millennium Prize Problem, September 8, 2026.
