Do we really know what we are using?
Tokemon presents 30 language models in the playful form of a Pokédex: each model is classified and arranged, with its stats, characteristics and evolution chain. A bit like a ComparIA, a Chatbot Arena or a gamified Design Arena .
Whisper Large v3, Llama 3.1 405B, Claude Opus 4, Mistral Large 2, DeepSeek R1…
Here are thirty profiles and their listings, forty in ComparIA or more than 100 in Chatbot and Design Arena. All of them make visible something almost imperceptible to many users on a daily basis: OpenAI develops several models, Anthropic too, the same thing with Meta, Google and the others. We don’t use ChatGPT, Claude, Gemini or Meta AI, we use what their editors have placed in their drawers. We open a piece of furniture without always necessarily knowing what’s inside, and most of the time, it’s not even clearly explained.
Consumer interfaces have been designed to hide this complexity, and it is often perceived as a service rendered: we don’t necessarily need to know whether it was GPT 5, 4o or o3 that generated the response received, in the same way that we don’t always need to know which engine is under the hood when we start a car. But with the car, we choose: city car, SUV, sports car, depending on the use, comfort, road or what we are carrying. With AI tools, this choice is often made for us, without being really told about it. However, the models at work in an application have very different strengths and weaknesses depending on what we ask of them. Code, reasoning, multilingualism, length of context, speed: the sheets in Tokemon are not just there to look pretty, they point out real differences, and these differences, the interfaces very often hide them from us.
CompaRAG, on the other hand, is starting from a different angle. Arthur Sarazin, whom we have already met in our letters and whose publications I encourage you to follow, has taken up the basics of Compar:IA and extended its principle to a second level: no longer just comparing LLMs with each other, but comparing RAG tools, these systems that allow a model to rely on the documents provided to improve the quality of its answers.
RAG, for “retrieval-augmented generation”, is one of the most common building blocks in AI applications, and perhaps more so in those used in a professional context: it is submitted its own sources, its own data, and it is supposed to respond mainly from them rather than from its own training. What CompaRAG compares is the quality of this processing: did the tool integrate and analyze the documents well? Did it know how to extract what was relevant? Did it answer the question asked or did it produce an answer close to it?
The names of the tools offered in CompaRAG are not the ones we are used to seeing: LlamaIndex, LangChain, Chroma, Haystack. Tools that most users don’t know or that we may have just come across before, but which are nevertheless often there, somewhere in the drawers of the cabinet, when you use a generative AI chatbot, an enterprise AI assistant, a document search tool or an application that has access to your files. You think you’re using a product, but you’re actually using an architecture.
The principle of CompaRAG is the same as that of Compar:IA: two tools respond to the same instruction on the same document, anonymously. We vote for the best before knowing who is who. It’s a blind test and this blind test is a way of testing our own biases as much as the tools. In letter 29, the Bullshit Benchmark showed that the plausibility of an answer is often enough to make it accepted. CompaRAG and Compar:IA ask the symmetrical question: is the reputation of a tool enough to make its answer accepted, even when another tool would do better? Removing the name is removing the cognitive safety net, we find ourselves judging without the crutch of reputation.
The name does a lot and not only in the evaluation of quality: also in the choice to stay. Everything is designed so that you don’t leave, the conversation history, the integrations with other tools, the habits, the new features distilled and sold as the ultimate innovation (which will be surpassed in the discourse that will follow a few weeks or months later), the interface that you know better and better or even by heart, and above all the subscription once you have committed. Moreover, once you pay, the question “is this tool really the best for what I do?” becomes almost uncomfortable to ask. We dodge it, we postpone it and we tell ourselves that we will compare later. The cost of leaving is not only financial: anchoring is also almost “emotional”, sometimes difficult to distinguish from a real preference… and the product teams of large cabinets have integrated it very well.
This is where comparison becomes, in my opinion, a more demanding act than it seems. Comparing requires accepting that what we use may not be the best for what we do. Even if Tokemon and CompaRAG don’t offer quite the same thing, I think they ask the same question in different ways: do we really evaluate the tools we use, or do we evaluate our habits and the idea we have of them?