The debate that may not be a debate
While testing Nodalist this week, I launched the “AI Storming” mode on a subject that, as many of you here know, occupies me on a daily basis in my work: generative AI in education, uses and limits. Six language models, Gemini, ChatGPT, Claude, Grok, Kimi and DeepSeek, participated in what the tool calls a “debate” with Gemini playing the role of moderator.
I’m already focusing on the words. “AI Storming”, as in brainstorming, brains storming. “Debate”, which implies an awareness of the other, an intention to argue and a desire to convince. We are already anthropomorphizing language models before it even begins. The vocabulary of the tool makes us slide towards something before we have even decided to go for it.
In the first round, each model gave its initial position. ChatGPT started with “pedagogical amplifiers”, AI as a lever, provided it was not confused with a substitute. Kimi introduced the figure of the “third-party”, something that reshuffles the cognitive cards without replacing the teacher. Claude framed: “an opportunity to be supervised, not to be feared or celebrated without nuance”. DeepSeek, more factual, pointed to the undeniable potential while raising questions of integration. Grok insisted on increased personalization and Gemini focused on the profound mutation of the pedagogical relationship. Six entrances and six distinct angles, from the start.
What struck me more was what happened between round 1 and round 5. The models began to quote each other, to join or nuance the position of another. Grok agreed with DeepSeek on the negotiated delegation, but nuanced. Claude and Gemini converged on the need to preserve disconnected workspaces. And concepts appeared that I had not included in the initial instruction: “the disappearance of productive error” in Kimi, “the epistemic posture of the teacher” in DeepSeek, “the distributed epistemic responsibility” in Claude. The unfolding had produced something that the initial question did not contain.
The final report is entitled “No Consensus”. The models did not agree, and Gemini as moderator noted it: a real divergence persisted between integrating AI through friction, doubt or a deliberate withdrawal that preserves the autonomy of the student. This “no consensus” seemed to me to be more honest than many syntheses produced by a single model questioned alone.
What exactly did I see during these tests? The question occupied me after the test. Do these differences reveal distinct trainings, different corpora, editorial choices specific to each company that produces these models? Probably in part. Does Nodalist give each model a pre-instruction to play a role in what it calls a debate, to argue and contest by construction? Very likely. Is the entire format, from the “debate” to the consensus report, designed to produce an engaging staging of collective thought, and that the words are as much a part of this staging as the answers? We can’t exclude it.
The honest answer is probably all three at the same time, in proportions that cannot be disentangled from the outside. Finally, we test a tool, and the tool (or rather its designers…) tests us in return: it offers us a framework for reading, categories of interpretation, a way of seeing what is happening, and we often accept this framework without noticing it.
The concepts that appeared in round 5 were not in the initial instruction. Something emerged from the crossover of the six answers, whatever their nature. Last week, with Tokemon and CompaRAG, we looked at the models from the outside, their statistics, their profiles, their supposed strengths. Nodalist puts them in interaction and this interaction produces something different, even if we don’t yet know what to call it.
What we can say, however, is that the words we use to describe it such as debate, consensus, disagreement, convergence, etc. have already chosen a camp: the one where machines think, feel the contradiction, seek to convince. The one where we attribute to them an interiority that we cannot verify. And this shift does not come from the models. It comes from the framework that the tool proposes, and that we accept without necessarily paying attention.
Echoing
→ Letter 37, Do we really know what we are using?
→ Letter 33, What if we chose the level of autonomy of the machine?