Away From Reductionism: “Artificial Intelligence” in All Its Shining Beauty

October 07, 2026 • 00:18:25
Away From Reductionism: “Artificial Intelligence” in All Its Shining Beauty
Localization Today
Away From Reductionism: “Artificial Intelligence” in All Its Shining Beauty

Oct 07 2026 | 00:18:25

/

Hosted By

Eddie Arrieta

Show Notes

By Sasha Fedorova, Florian Jaton, and Dmitrii Shevchuk 

AI is hyped as a force that will revolutionize the future. But the future isn’t here yet, and in the meantime, many argue the conversation surrounding the technology has been defined by reductionism and thought-terminating cliches.

View Full Transcript

Episode Transcript

[00:00:00] Away from reductionism Artificial intelligence in all its shining beauty By Sasha Fedorova, Florine Giton and Dmitry I. Shevchak Time and again, scrolling social media feeds and skimming news websites leaves us reeling from an overdose of reductionism. [00:00:19] Experts and all sorts of countless commentators keep repeating the techno deterministic mantra. Artificial intelligence is revolutionizing startup founders and tech developers in love with their digital tools reduce our life to customer needs and technological solutions. [00:00:36] Innovators dreaming of crossing the chasm reduce society to segments from early adopters to late majority and laggards. Geoffrey Hinton reduces the hospital to a computer vision tool playground and advises stop training radiologists. Now venture capitalists think about tenfold returns from a batch of ChatGPT wrappers and raising the next artificial intelligence funding round. Accounting clerks, back end coders, and translators are scared and paralyzed because someone said that artificial intelligence has already replaced them. The chief financial officer of a multinational company is adding savings from personnel cuts at headquarters and in the subsidiaries and subtracting license fees to frontier model vendors to estimate the bottom line of digital transformation gains. Everyone is fueling the fear of missing out. How has mainstream discourse become so reductionist and impoverished? Attempts to thoroughly assess neural network based technologies in real life settings are still rare compared with the amount of artificial intelligence promotional slop. Among the rare examples is an assessment of automatic speech to speech translation in a conference setting by the World Health Organization interpretation team. From a technical point of view, automatic STS translation is a stack of speech transcription, machine translation, and text to speech technologies which are fundamentally based on the neural network paradigm. The WHO assessment adapted criteria developed for human interpreters. [00:02:12] In summary, the results of this exercise showed that grades for automatic interpretation ranged from 5% to 83% across all United nations language combinations of Arabic, Chinese, English, French, Russian, and Spanish. [00:02:28] Only 1 out of 90 renditions English into French received a passing grade, which was set at 75%. [00:02:36] None of the automatic STS translations were free of reputational risks, which ranged from 1 to 9 per rendition. A single reputational risk was treated as eliminatory due to its potential for miscommunication. [00:02:49] According to the head of the WHO interpretation team, automatic interpretation produced surprising errors that differed from the mistakes evaluators would expect from a human interpreter. Automatic STS translation tackles the practical task of multilingual communication in a conference setting differently from human interpreters. [00:03:08] The interpreting community's folklore circulates an example of how automatic STS translation, while good at numbers that are traditionally challenging for interpreters, can mix up avocado and advocate. A report by the Council of Europe Interpretation Service documents an example of a numbered list item 6 judges that in which the numeral and the verb came out as six judges who completely changing the meaning. When it is hard to make sense of the input speech interpreters use the technique of word by word translation to maintain the speech delivery flow. Interpreters use this time gaining technique in exceptional cases to grasp key points and continue delivering the message while maintaining multilingual communication. Automatic STS translation parrots output speech, merely mimicking understanding. [00:03:59] To put it in terms coined by linguist Emily Bender from the University of Washington, what resists reduction is what reality is made of. Human interpretation cannot be reduced to word by word. Language processing and automatic STS translation is fundamentally different from human interpretation. The two are not reducible to one another, and more broadly it requires an informed view on steering the development and deployment of machine systems, starting from the decision of their necessity in the first place. [00:04:30] Setting aside the metaphysical question of whether a machine can think or have intentions, any case needs to be understood in historical perspective and through its technical details in an effort to make sense of what it is not reduced to a myth of linear technological progress or anthropomorphizing intelligence metaphors, but in all its shining beauty, as Bruno Latour puts it, in irreductions. [00:04:54] This is a twinning story of attempts to find a solution of the worldwide translation problem through the use of electronic computers of great capacity, flexibility and speed. This quote takes us Back to the 1949 translation memorandum by Warren Weaver, director of the Rockefeller Foundation's Natural Sciences Division. [00:05:14] This initial inspiration was followed by a moment when, in the wake of the Sputnik launch in 1957, machine translation efforts were generously funded by the US National Research Council to speed up the translation of Soviet scientific papers. Disillusionment followed when systems failed to meet expectations. [00:05:34] By 1966 it was concluded that there has been no machine translation of general scientific text and none is an immediate prospect and funding was canceled. Recent progress in language processing Another turn in this 70 year long story, can be attributed to technological developments in neural network software architectures and graphics processing unit hardware originally developed for video gaming. One might think of GeForce, Nvidia or Radeon graphics cards in their first desktop computer running the Quake 3 shooter and later widely used for cryptocurrency mining. GPUs proved to be well suited hardware for vector calculations and probabilistic estimation. [00:06:19] In 2012, researchers from the University of Toronto, Jeffrey Hinton and his doctoral students Alex Krzyevsky and Ilya Setzkever, convincingly demonstrated that a previously marginalized neural network software architecture, when run on GPU hardware, could perform computer vision tasks more accurately than established approaches like scale invariant, feature transform and histogram of oriented gradients. [00:06:45] In the late 2000 and tens, neural networks using the transformer architecture outperformed established approaches that relied on complex syntactic or semantic rules and drew conceptually from the Chomskyan transformational generative grammar paradigm. For details, see the recent book by computational linguist Thierry Poirot from the Paris Artificial Intelligence Research Institute. Neural networks further demonstrated efficiency in natural language processing. [00:07:13] More precisely, they used large training datasets to find statistical patterns and powerful graphical computational hardware to perform vector calculations and probabilistic estimations sufficient to produce outputs that sound meaningful. The practical efficiency of neural networks in language processing was demonstrated to a Wider audience in 2016, when Google Translate transitioned to neural network architecture. [00:07:38] A further Milestone came in 2019 when the California based company OpenAI launched the Generative Pre Trained Transformer model. This GPT2 model was trained on statistical patterns of word combinations from a corpus of 8 million web pages equivalent to about 6 billion words. The model was capable of generating text by putting words one after another into a plausible textual construction. [00:08:05] Developers prompted GPT2 with three lines of a fictional discovery. In a shocking finding, scientists discovered a herd of unicorns living in a remote, previously unexplored valley in the Andes mountains. Even more surprising to the researchers was the fact that the unicorns spoke perfect English. The GPT2 model's output was about 2,000 characters long and, in a surprisingly plausible manner, narrated that a fictional Dr. Jorge Perez from the University of La Paz reported that the population was named Ovid's Unicorn after its distinctive horn. The fictional scientist concluded, they seem to be able to communicate in English quite well, which I believe is a sign of evolution or at least a change in social organization. [00:08:52] In late 2022, this time publicly, OpenAI launched its flagship generative AI chatbot, ChatGPT, based on the GPT 3.5 model that triggered unhealthy hype dynamics around the topic of artificial intelligence. ChatGPT became a global phenomenon, reaching 1 million users in just five days. The newer GPT 5.5 is claimed by OpenAI to excel at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. Tech companies do not disclose the exact data set sizes for newer frontier models, but estimates suggest that generative models are trained on data equivalent to hundreds of terabytes of text, a growing proportion of which is synthetic or generated by other models. Conceptually, today's frontier models operate on the same principle as in the case of the fictional Dr. Jorge Perez and Ovid's unicorns. Trained on patterns from datasets, a model generates statistically plausible outputs which the following example illustrates. A 20 line Python code and an open access GPT type model QWEM3.1.7B were used to dive deeper into the normally obscured technicalities of the underlying mechanism. [00:10:19] Once the prompt is provided, the model tokenizes this input, splitting it into chunks of symbols, vectorizes it into mathematical representations for processing, and estimates probabilities of the next token in a sequence. To put it differently, a model guesses one word after another, referring to statistical patterns in the datasets it processed during the training stage. Let's return to the phrase above In a shocking finding, scientists discovered a herd of unicorns living in a remote, previously unexplored valley in the Andes mountains. Even more surprising to the researchers was the fact that the unicorns spoke perfect. The final word is missing, leaving the model to complete the phrase first, the model chunks the phrase into more than 40 tokens, such as shocking finding and scientists on YouTube one might find insights from OpenAI's Andrej Karpathy on GPT2 tokenization. [00:11:17] Notably, toponyms and the main subject of the sentence are lost at this step as they become tokens such as n, es, unic, and orns. [00:11:29] Next, the model translates every token into a mathematical vector representation and processes it to estimate probabilities based on patterns in training data sets of the next token see Figure 2. The most plausible next token options statistically estimated would be English Bravo Qwen 3.1.7 B, French and German, among others. The probability of Latin is half that of Chinese. Trained on vast amounts of text from books and websites, models generate outputs by identifying statistical patterns in language. [00:12:06] Although outputs may seem coherent, it remains debated whether producing text by predicting plausible word sequences is the same as understanding. The critique emphasizes that meaning requires grounding in the real world, that is Connecting words to actual objects, events, and contexts. Learning statistical relationships between words alone does not provide this connection. With a non 0 probability, unicorns might be statistically estimated to speak Hebrew or Japanese. [00:12:37] The fictional Dr. Jorge Perez might appear as Advocate Perez or even Avocado Perez, able to communicate in Latin as a sign of evolution or a change in social organization in the neural network paradigm. Further developments are inherently dependent on the scaling of training datasets and data centers as their materialization. [00:12:58] Scaling is supposed to increase the amount of statistical multi layered contextual data considered when estimating the probability of the next tokens. [00:13:08] Scaling from hundreds of terabytes used to train GPT 5.5 to petabytes or even zettabytes of text for training envisioned in advanced models is expected to be supported by tens of data centers the size of Manhattan, such as Meta's Hyperion facility in Louisiana. Already today, training data of sufficient quality is becoming scarce. [00:13:29] Facing constraints on available training data, developers increasingly use synthetic data neural network generated data for training. The specificity of synthetic data will unavoidably steer model outputs as in early models like GPT1 that were trained on corpora of vampire sagas, which gave their outputs a flavor of dark romanticism. [00:13:52] More consequentially, the use of synthetic data and potentially second order synthetic data, meaning data generated by models trained on synthetic data, will further sharpen the question of generative models lacking grounding in reality. Mainstream discourse on artificial intelligence has grown reductionist through misleading anthropomorphizing metaphors, technodeterminist assumptions, and the near complete absence of rigorous assessment of technology in the wild. [00:14:21] 70 years of machine translation research have shown neither the linear progress the current hype implies nor any inevitability of continued advances. [00:14:29] A closer look at the underlying technical mechanism of neural networks, tokenization and statistical next token estimation reveals a process that should be thoroughly understood before a deployment decision is made. The architecture of this statistical paradigm sets the ceiling on a potential breakthrough that is needed to bring it to the next level. Data scarcity is only one of the constraints Ecological costs and infrastructure have not even entered the conversation. [00:14:57] Examining these technicalities also exposes the gap between outputs that sound meaningful and anchoring in reality. [00:15:05] Human and machine processes are fundamentally different, and deploying machine systems in professional settings requires careful adaptation and configuration of existing practices. [00:15:16] Who steers this development can affected communities, professions and broader society find a place in an arena almost completely crowded out by big tech and their investors? [00:15:27] And will they enter it fed on technodeterminist slogans or armed with fully fledged argumentation? [00:15:33] Sasha Fedorova is a conference interpreter and a PhD candidate at the University of Geneva, where her thesis examines the practice of simultaneous interpreting as lived experience. [00:15:45] Drawing on an inactive and semiotic research framework. She works between French, English and Russian for Geneva based international organizations. [00:15:55] Florian Gatone is a researcher and lecturer at the University of Lausanne Faculty of Social and Political Sciences and at EPFL College of Humanities. He worked at the Donald Brent School of Information and Computer Science at the University of California at the Centre d Sociologie d' Innovation at Mainz Paris PSL University and at the Geneva Graduate Institute of International and Development Studies. [00:16:22] His research interests are the sociology of algorithms, the philosophy of mathematics and the history of computing. He is the author of the book the Constitution of Ground truthing programming formulating. MIT Press, 2021 open access Dmitry Shevchak is a master's student in digital society. He is interested in socio technical visions of machine learning technology, the processes through which they are shaped and shared, and the role they play in technology development and adoption. This article was written by Sasha Fedorova, a conference interpreter and a PhD candidate at the University of Geneva where her thesis examines the practice of simultaneous interpreting as lived experience, drawing on an inactive and semiotic research framework. She works between French, English and Russian for Geneva based international organizations. Florian Giton A researcher and lecturer at the University of Lausanne Faculty of Social and Political Sciences and at EPFL College of Humanities. He worked at the Donald Bren School of Information and Computer Science at the University of California, at the Centre de Sociologie de l' Innovation at Mainz, Paris PSL University and at the Geneva Graduate Institute of International and Development Studies. His research interests are the sociology of algorithms, the philosophy of mathematics and the history of computing. He is the author of the book the Constitution of Ground truthing, programming formulating MIT Press, 2021 open access and Dmitry Izhevchuk a master's student in digital Society. [00:18:07] He is interested in sociotechnical visions of machine learning technology, the processes through which they are shaped and shared, and the role they play in technology development and adoption. Originally published in Multilingual Magazine issue 256October 2026.

Other Episodes

Episode 212

September 12, 2022 • 00:07:29
Episode Cover

GALA urges continued support for Ukraine from localization community

The Globalization and Localization Association (GALA) has released the results of a survey on support to Ukraine, conducted with 40 respondents of its community.

Listen

Episode 20

January 31, 2022 • 00:06:11
Episode Cover

Largest LSP breaks billion-dollar barrier

Long-time industry leader TransPerfect just revealed that the company reached USD 1.1 billion in revenue in 2021. This is up from USD 852.4 million...

Listen

Episode 179

June 05, 2024 • 00:09:39
Episode Cover

Reimagining the Content Lifecycle With Strategic Use of Generative AI

By Kajetan Malinowski Content marketers are asking themselves: should they continue to create content in one language and then localize it for many markets,...

Listen