Cross-Linguistic Epistemic Variance in Large Language Model Cognition

Authors

  • Muhammad Saad Nizami Association for Computing Machinery (ACM) at ISL, Pakistan
  • Muhammad Ramzan Shahid Khan Department of Computer Science, Namal University, Pakistan
  • Muhammad Bilal Department of Computer Science, Namal University, Pakistan

Keywords:

Large Language Models, Multilingual NLP, Cross-lingual Evaluation, Linguistic Bias, AI Fairness

Abstract

LLMs are used worldwide, but whether they perform consistently across different languages is a question that has not been rigorously tested. This paper examines cross-lingual prompt sensitivity by running four models (Kimi K2 Instruct, Llama 3.3 70B, Qwen3 32B, and GPT-OSS 120B) on 48 semantically equivalent prompts across English, Spanish, Urdu, Japanese, and Arabic, covering six prompt categories and producing 960 total responses. The observed disparities were larger than expected. Japanese prompts received responses averaging 187.1 words versus 646.9 for English, a statistically significant 71% gap (Cohen’s d = 1.14, 95% CI [0.91, 1.37], p < 0.0001), and 26.6% of Japanese responses were classified as minimal compliance compared to 4–6% in the other languages. One-way ANOVA confirmed that response length varies significantly by language (F (4, 955) = 34.93, p < 0.0001, η² = 0.128, medium-to-large effect), and chi-square analysis confirmed the same for refusal patterns (χ² (12) = 100.54, p < 0.0001, Cramér’s V = 0.187). Model effects on response length were even larger (F (3, 956) = 128.57, p < 0.0001, η² = 0.287). Japanese responses showed the highest Type-Token Ratios (mean = 0.823, 95% CI [0.797, 0.849]), reflecting a measurement artifact from brevity rather than genuine vocabulary richness. All four models showed this pattern across all six prompt categories, making it difficult to attribute to any single architectural choice. Training data imbalances and tokenization inefficiencies are the most plausible explanations. For a technology marketed as globally accessible, these results raise questions worth taking seriously.

References

Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, “PaLM: Scaling Language Modeling with Pathways,” arXiv:2204.02311, 2022, [Online]. Available: https://arxiv.org/abs/2204.02311

Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, “LoRA: Low-Rank Adaptation of Large Language Models,” arXiv:2106.09685, 2021, [Online]. Available: https://arxiv.org/abs/2106.09685

Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, “MEGA: Multilingual Evaluation of Generative AI,” arXiv:2303.12528, 2023, [Online]. Available: https://arxiv.org/abs/2303.12528

Ian Magnusson, Noah A. Smith, Jesse Dodge, “Reproducibility in NLP: What Have We Learned from the Checklist?,” arXiv:2306.09562, 2023, [Online]. Available: https://arxiv.org/abs/2306.09562

Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, “Unsupervised Cross-lingual Representation Learning at Scale,” arXiv:1911.02116, 2020, [Online]. Available: https://arxiv.org/abs/1911.02116

Su Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna Wallach, “Language (Technology) is Power: A Critical Survey of ‘Bias’ in NLP,” arXiv:2005.14050, 2020, [Online]. Available: https://arxiv.org/abs/2005.14050

I. Solaiman et al., “Evaluating the Social Impact of Generative AI Systems in Systems and Society,” Jun. 2024, Accessed: Jun. 20, 2026. [Online]. Available: http://arxiv.org/abs/2306.05949

Orevaoghene Ahia, Sachin Kumar, Hila Gonen, “Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models,” arXiv:2305.13707, 2023, [Online]. Available: https://arxiv.org/abs/2305.13707

Aleksandar Petrov, Emanuele La Malfa, Philip H.S. Torr, Adel Bibi, “Language Model Tokenizers Introduce Unfairness Between Languages,” arXiv:2305.15425, 2023, [Online]. Available: https://arxiv.org/abs/2305.15425

Yanzhu Guo, Guokan Shang, Chloé Clavel, “Benchmarking Linguistic Diversity of Large Language Models,” arXiv:2412.10271, 2025, [Online]. Available: https://arxiv.org/abs/2412.10271

Xu Huang, Wenhao Zhu, Hanxu Hu, Conghui He, Lei Li, Shujian Huang, Fei Yuan, “BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models,” arXiv:2502.07346, 2025, [Online]. Available: https://arxiv.org/abs/2502.07346

Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, “Language Models are Multilingual Chain-of-Thought Reasoners,” arXiv:2210.03057, 2022, [Online]. Available: https://arxiv.org/abs/2210.03057

Viet Dac Lai, Nghia Trung Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, Thien Huu Nguyen, “ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning,” arXiv:2304.05613, 2023, [Online]. Available: https://arxiv.org/abs/2304.05613

Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, “The Flan Collection: Designing Data and Methods for Effective Instruction Tuning,” arXiv:2301.13688, 2023, [Online]. Available: https://arxiv.org/abs/2301.13688

Niklas Muennighoff, Thomas Wang, Lintang Sutawika, “Crosslingual Generalization through Multitask Finetuning,” arXiv:2211.01786, 2023, [Online]. Available: https://arxiv.org/abs/2211.01786

B. H. Pinzhen Chen, Simon Yu, Zhicheng Guo, “Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?,” arXiv:2406.12822, 2024, [Online]. Available: https://arxiv.org/abs/2406.12822

Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, “Bias and Fairness in Large Language Models: A Survey,” arXiv:2309.00770, 2024, [Online]. Available: https://arxiv.org/abs/2309.00770

Krithika Ramesh, Sunayana Sitaram, Monojit Choudhury, “Fairness in Language Models Beyond English: Gaps and Challenges,” arXiv:2302.12578, 2023, [Online]. Available: https://arxiv.org/abs/2302.12578

Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, Sebastian Ruder, “UNKs Everywhere: Adapting Multilingual Language Models to New Scripts,” arXiv:2012.15562, 2021, [Online]. Available: https://arxiv.org/abs/2012.15562

L. W. Julia Kreutzer, Isaac Caswell, “Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets,” arXiv:2103.12028, 2022, [Online]. Available: https://arxiv.org/abs/2103.12028

Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, Quoc V. Le, “Finetuned Language Models Are Zero-Shot Learners,” arXiv:2109.01652, 2022, [Online]. Available: https://arxiv.org/abs/2109.01652

Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, “Scaling Instruction-Finetuned Language Models,” arXiv:2210.11416, 2022, [Online]. Available: https://arxiv.org/abs/2210.11416

P. M. McCarthy and S. Jarvis, “MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment,” Behav. Res. Methods 2010 422, vol. 42, no. 2, pp. 381–392, May 2010, doi: 10.3758/BRM.42.2.381.

Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, Chenguang Zhu, “G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment,” arXiv:2303.16634, 2023, [Online]. Available: https://arxiv.org/abs/2303.16634

Mingqi Gao, Xinyu Hu, Jie Ruan, Xiao Pu, Xiaojun Wan, “LLM-based NLG Evaluation: Current Status and Challenges,” arXiv:2402.01383, 2025, [Online]. Available: https://arxiv.org/abs/2402.01383

Rotem Dror, Gili Baumer, Segev Shlomov, Roi Reichart, “The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing,” ACL 2018 - 56th Annu. Meet. Assoc. Comput. Linguist. Proc. Conf. (Long Pap., 2018, [Online]. Available: https://aclanthology.org/P18-1128/

P. Naresh, S. V. N. Pavan, A. R. Mohammed, N. Chanti, and M. Tharun, “Comparative Study of Machine Learning Algorithms for Fake Review Detection with Emphasis on SVM,” Int. Conf. Sustain. Comput. Smart Syst. ICSCSS 2023 - Proc., pp. 170–176, 2023, doi: 10.1109/ICSCSS57650.2023.10169190.

Uri Shaham, Jonathan Herzig, Roee Aharoni, Idan Szpektor, Reut Tsarfaty, Matan Eyal, “Multilingual Instruction Tuning With Just a Pinch of Multilinguality,” arXiv:2401.01854, 2024, [Online]. Available: https://arxiv.org/abs/2401.01854

E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?,” FAccT 2021 - Proc. 2021 ACM Conf. Fairness, Accountability, Transpar., pp. 610–623, Mar. 2021, doi: 10.1145/3442188.3445922.

M. R. Hayat, A. Bakar, M. Bilal∗, and M. R. S. Khan, “Data-Driven Analysis and Predictive Modeling of Urban Air Quality for Environmental Management,” Int. J. Innov. Sci. Technol., vol. 8, no. 3, pp. 508–527, 2026, Accessed: Jun. 30, 2026. [Online]. Available: https://ideas.repec.org/a/abq/ijist1/v8y2026i3p508-527.html

Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer, “QLoRA: Efficient Finetuning of Quantized LLMs,” arXiv:2305.14314, 2023, [Online]. Available: https://arxiv.org/abs/2305.14314

Cheng-Han Chiang, Hung-yi Lee, “Can Large Language Models Be an Alternative to Human Evaluations?,” arXiv:2305.01937, 2023, [Online]. Available: https://arxiv.org/abs/2305.01937

Downloads

Published

2026-06-17

How to Cite

Nizami, M. S., Muhammad Ramzan Shahid Khan, & Muhammad Bilal. (2026). Cross-Linguistic Epistemic Variance in Large Language Model Cognition. International Journal of Innovations in Science & Technology, 8(3), 1299–1314. Retrieved from https://journal.50sea.com/index.php/IJIST/article/view/1932