The impact of generative models on legal translation: a multidimensional analysis

Authors

DOI:

https://doi.org/10.4151/S0718-09342026012201394

Keywords:

corpus linguistics, generative artificial intelligence, multidimensional analysis, legal discourse.

Abstract

This exploratory study applies Bibers (1988) multidimensional analysis (MDA) to examine the linguistic features of human translation (HT), neural machine translation (NMT), and translation produced by a generative AI model (GPT-4) within the legal domain, specifically in judgments issued by the Court of Justice of the European Union (CJEU). While MDA has been widely adopted in corpus linguistics, its application to translation studies remains incipient. Drawing on a corpus of over 2.6 million words, the study conducts a novel factorial analysis using MFTE-tagged texts to identify co-occurring linguistic patterns across the three translation methods. The results reveal significant differences, particularly in two main dimensions: one related to syntactic complexity and formal exposition, and another to procedural and event-focused language. GPT-4 translations exhibit higher syntactic density and procedural expression, closely aligning with the communicative functions of judicial discourse. In contrast, NMT outputs tend to be more conceptually abstract and less structurally varied, while HT demonstrates a balance between formality and communicative adaptation. These findings underscore the value of MDA in translation studies and highlight the stylistic and functional variation introduced by different translation technologies. The study offers both methodological and pedagogical implications, paving the way for future research into translation quality and variation through a multidimensional lens.

Author Biographies

María del Mar Sánchez Ramos, Universidad de Alcalá

María del Mar Sánchez Ramos holds a PhD in Translation and Interpreting from the Jaume I University of Castelló (Spain). She is currently Full Professor in Translation Studies at the Department of Modern Philology (Universidad de Alcalá, Spain), where she combines teaching with her research work. Dr Sánchez Ramos’s main research interests are translation technologies (i.e., machine translation and post-editing), corpus-based translation studies, and discourse analysis. She has extensively published about these topics in international and national top-ranked indexed journals. She has been alto a Visiting Professor at several  universities, the University of Limerick (Ireland), the University of Geneva (Switzerland) and Charles University in the Czech Republic. She has been a member of Spain’s FITISPos-UAH Research Group (Training and Research in Translation and Interpreting in Public Services) since 2012.

Muhammad Shakir, Universidad de Münster

Muhammad Shakir is a post-doctoral researcher and completing his Habilitation at the University of Muenster, Germany since 2021. His research interests include language variation, corpus linguistics, register studies, MD analysis, and computer-mediated communication (CMC). He completed his PhD on MD analysis of Pakistani English CMC. In the Habilitation he focuses on variation in the English language on a regional level in South Asia and the Caribbean. For this purpose, he has compiled the South Asian Online Englishes (SAOnE) corpus. His recent research has been on the use of indigenous discourses markers in the English CMC of South Asians, register variation in South Asian CMC, the use of indigenous memes in English tweets for political satire, and variation in the use of the modal must versus semi-modals of obligation have (got) to, have to, and need to. For the second phase of his Habil, he is currently compiling the Caribbean Online Englishes (CAOnE) corpus, the Caribbean version of SAOnE.

References

Albors-Llorens, A. (2020). Judicial protection before the Court of Justice of the European Union. In C. Barnard & S. Peers (Eds.), European Union law (3rd ed., pp. 283–333). Oxford University Press.

Baker, M. (1993). Corpus linguistics and translation studies: Implications and applications. In G. Francis & E. Tognini-Bonelli (Eds.), Text and technology: In honour of John Sinclair (pp. 233–252). John Benjamins.

Berber-Sardinha, T. (2024). AI-generated vs human-authored texts: A multidimensional comparison. Applied Corpus Linguistics, 4(1), Article 100083.

Biber, D. (1988). Variation across speech and writing. Cambridge University Press.

Biber, D. (1995a). Dimensions of register variation: A cross-linguistic comparison. Cambridge University Press.

Biber, D. (1995b). On the role of computational, statistical, and interpretive techniques in multi-dimensional analyses of register variation: A reply to Watson. Text: Interdisciplinary Journal for the Study of Discourse, 15(3), 341–370.

Biber, D. (2006). University language: A corpus-based study of spoken and written registers. Benjamins.

Biber, D., Johansson, S., Leech, G., Conrad, S., & Finegan, E. (1999). Longman grammar of spoken and written English. Longman.

Briva-Iglesias, V., Dogru, G., & Cavalheiro Camargo, J. L. (2024). Large language models "ad referendum": How good are they at machine translation in the legal domain?. MonTi Monografías De Traducción E Interpretación, (16), 75–107.

AUTHOR

Castilho, S., & Resende, N. (2022). Post-editese in literary translations. Information, 13(2), Article 66. https://doi.org/10.3390/info13020066

Chou, I., & Liu, K. (2024). Style in speech and narration of two English translations of Hongloumeng: A corpus-based multidimensional study. Target, 36(1), 77–111.

De Sutter, G., & Lefer, M. A. (2020). On the need for a new research agenda for corpus-based translation studies: A multi-methodological, multifactorial and interdisciplinary approach. Perspectives, 28(1), 1–23.

Egbert, J., & Staples, S. (2019). Doing multidimensional analysis in SPSS, SAS and R. In T. Berber-Sardinha & M. V. Pinto (Eds.), Multidimensional analysis research methods and current issues (pp. 125–144). Bloomsbury.

Frankenberg-Garcia, A. (2022). Can a corpus-driven lexical analysis of human and machine translation unveil discourse features that set them apart? Target, 34(2), 278–308.

Giampierie, P. (2025). AI-Powered contracts: A critical analysis. International Journal of Semiotics Law, 38, 403-420.

Heiss, C., & Soffritti, M. (2018). DeepL traduttore e didattica della traduzione dall’italiano in tedesco. alcune valutazioni preliminari. In L. Anderson, L. Gavioli & F. Zanettin (Eds.), InTRAlinea. Special Issue: ‘Translation and Interpreting for Language Learners’ (TAIL). https://www.intralinea.org/specials/article/2294

Ilisei, I., & Inkpen, D. (2011). Translationese traits in Romanian newspapers: A machine learning approach. International Journal of Computational Linguistics and Applications, 2(2), 319–332.

Jiao, et al. (2023), Is ChatGPT A Good Translator? Yes with GPT-4

as the Engine. arXiv. https://doi.org/10.48550/arXiv.2301.08745

Killman, J. (2023). Rendering morphosyntactic features of legal Spanish judgments using NMT and SMT. In J. Zhao, D. Li, & V. L. C. Lei (Eds.), New advances in legal translation and interpreting (pp. 221–242). Springer.

Krüger, R. (2020). Explicitation in neural machine translation. Across Languages and Cultures, 21(2), 195-216.

Kruger,H., & Van Rooy, B. (2016). Constrained language. A multidimensional analysis of translated English and a non-native indigenised variety of English. English World-Wide, 37(1), 26–57.

Lapshinova-Koltunski, E. (2015). Variation in translation: evidencie from corpora. En C. Fantinuoli, F. Zanettin (Eds.), New directions in corpus-based translation studies. Language Science Press.

Le Foll, E. (2022). Textbook English: A corpus-based analysis of the language of EFL textbooks used in secondary schools in France, Germany and Spain [Doctoral dissertation, University of Osnabrück]. osnaDocs. https://doi.org/10.48693/278

Le Foll, E., & Shakir, M. (2023). MFTE Python [Computer software]. https://github.com/mshakirDr/MFTE

Le Foll, E., & Shakir, M. (2024). The Multi-Feature Tagger of English (MFTE): Rationale, description and evaluation. Research in Corpus Linguistics, 13(2), 63–93.

Michał Ziemski, Marcin Junczys-Dowmunt, & Bruno Pouliquen. (2016). The United Nations Parallel Corpus v1.0. In N. Calzolari, K. Choukri, T. Declerck, S. Goggi, M. Grobelnik, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, & S. Piperidis (Eds.), Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16) (pp. 3530–3534). European Language Resources Association.

Neumann, S., & Evert, S. (2021). A register variation perspective on varieties of English. In E. Seoane & D. Biber (Eds.), Corpus-based approaches to register variation (pp. 143–178). John Benjamins.

Niu, J. & Jiang, Y. (2024). Does simplification hold true for machine translations? A corpus-based analysis of lexical diversity in text varieties accross genres. Humanities and Social Sciences Commnications, 11, 1-10.

Oliver, A. (2021). MTUCOC-Translator [Computer software]. https://github.com/aoliverg/MTUOC-translator

R Core Team. (2024). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.R-project.org/

Revelle, W. (2024). psych: Procedures for psychological, psychometric, and personality research (Version 2.4.6.26) [Computer software]. Northwestern University. https://cran.r-project.org/package=psych

Rico Pérez, C. (2025). Estudio del efecto de imprimación de la traducción automátcia sobre un corpus de textos del español institucional. Revista de Humanidades Digitales,10: 48-72.

Rico Pérez, C., Sánchez Ramos, M. M., & Oliver, A. (2020). INMIGRA3: Building a case for NGOs and NMT. In A. Martins, H. Moniz, S. Fumega, B. Martins, F. Batista, L. Coheur, C. Parra, I. Trancoso, M. Turchi, A. Bisazza, J. Moorkens, A. Guerberof, M. Nurminen, L. Marg, & M. Forcada (Eds.), Proceedings of the 22nd Annual Conference of the European Association for Machine Translation (pp. 469–471). European Association for Machine Translation.

Şahin, M., & Dungan, N. (2014). Translation testing and evaluation: A study on methods and needs. Translation & Interpreting, 6(2), 67–90.

Toral, A. (2019). Post-editese: An exacerbated translationese. In M. Forcada, A. Way, B. Haddow, & R. Sennrich (Eds.), Proceedings of Machine Translation Summit XVII: Research track (pp. 273–281). European Association for Machine Translation.

Trklja, A., & McAuliffe, K. (2019). Formulaic metadiscursive signalling devices in judgments of the Court of Justice of the European Union: A new corpus-based model for studying discourse relations of texts. International Journal of Speech, Language and the Law, 26(1), 21–55.

Vanmassenhove, E., Shterionov, D., & Way, A. (2019, August 19–23). Lost in translation: Loss and decay of linguistic richness in machine translation [Paper presentation]. Machine Translation Summit XVII, Dublin, Ireland.

Vigier Moreno, F., & Sánchez Ramos, M. M. (2017). Using parallel corpora to study the translation of legal-system bound terms: The case of names of English and Spanish Courts. In R. Mitkov (Ed.), Computational and corpus-based phraseology: Second International Conference, Europhras 2017, London, UK, November 13–14, 2017 proceedings (pp. 260–273). Springer.

Volansky, V., Ordan, N., & Wintner, S. (2015). On the features of translationese. Digital Scholarship in the Humanities, 30(1), 98–118.

Wiesmann, E. (2019). Machine translation in the field of law: A study of the translation of Italian legal texts into German. Comparative Legilinguistics, 37, 117–153.

Published

2026-10-01

How to Cite

Sánchez Ramos, M. del M., & Shakir, M. (2026). The impact of generative models on legal translation: a multidimensional analysis. Revista Signos. Estudios De Lingüística, 59(122). https://doi.org/10.4151/S0718-09342026012201394