The impact of generative models on legal translation: a multidimensional analysis
DOI:
https://doi.org/10.4151/S0718-09342026012201394Keywords:
corpus linguistics, generative artificial intelligence, multidimensional analysis, legal discourse.Abstract
This exploratory study applies Bibers (1988) multidimensional analysis (MDA) to examine the linguistic features of human translation (HT), neural machine translation (NMT), and translation produced by a generative AI model (GPT-4) within the legal domain, specifically in judgments issued by the Court of Justice of the European Union (CJEU). While MDA has been widely adopted in corpus linguistics, its application to translation studies remains incipient. Drawing on a corpus of over 2.6 million words, the study conducts a novel factorial analysis using MFTE-tagged texts to identify co-occurring linguistic patterns across the three translation methods. The results reveal significant differences, particularly in two main dimensions: one related to syntactic complexity and formal exposition, and another to procedural and event-focused language. GPT-4 translations exhibit higher syntactic density and procedural expression, closely aligning with the communicative functions of judicial discourse. In contrast, NMT outputs tend to be more conceptually abstract and less structurally varied, while HT demonstrates a balance between formality and communicative adaptation. These findings underscore the value of MDA in translation studies and highlight the stylistic and functional variation introduced by different translation technologies. The study offers both methodological and pedagogical implications, paving the way for future research into translation quality and variation through a multidimensional lens.
References
Albors-Llorens, A. (2020). Judicial protection before the Court of Justice of the European Union. In C. Barnard & S. Peers (Eds.), European Union law (3rd ed., pp. 283–333). Oxford University Press.
Baker, M. (1993). Corpus linguistics and translation studies: Implications and applications. In G. Francis & E. Tognini-Bonelli (Eds.), Text and technology: In honour of John Sinclair (pp. 233–252). John Benjamins.
Berber-Sardinha, T. (2024). AI-generated vs human-authored texts: A multidimensional comparison. Applied Corpus Linguistics, 4(1), Article 100083.
Biber, D. (1988). Variation across speech and writing. Cambridge University Press.
Biber, D. (1995a). Dimensions of register variation: A cross-linguistic comparison. Cambridge University Press.
Biber, D. (1995b). On the role of computational, statistical, and interpretive techniques in multi-dimensional analyses of register variation: A reply to Watson. Text: Interdisciplinary Journal for the Study of Discourse, 15(3), 341–370.
Biber, D. (2006). University language: A corpus-based study of spoken and written registers. Benjamins.
Biber, D., Johansson, S., Leech, G., Conrad, S., & Finegan, E. (1999). Longman grammar of spoken and written English. Longman.
Briva-Iglesias, V., Dogru, G., & Cavalheiro Camargo, J. L. (2024). Large language models "ad referendum": How good are they at machine translation in the legal domain?. MonTi Monografías De Traducción E Interpretación, (16), 75–107.
AUTHOR
Castilho, S., & Resende, N. (2022). Post-editese in literary translations. Information, 13(2), Article 66. https://doi.org/10.3390/info13020066
Chou, I., & Liu, K. (2024). Style in speech and narration of two English translations of Hongloumeng: A corpus-based multidimensional study. Target, 36(1), 77–111.
De Sutter, G., & Lefer, M. A. (2020). On the need for a new research agenda for corpus-based translation studies: A multi-methodological, multifactorial and interdisciplinary approach. Perspectives, 28(1), 1–23.
Egbert, J., & Staples, S. (2019). Doing multidimensional analysis in SPSS, SAS and R. In T. Berber-Sardinha & M. V. Pinto (Eds.), Multidimensional analysis research methods and current issues (pp. 125–144). Bloomsbury.
Frankenberg-Garcia, A. (2022). Can a corpus-driven lexical analysis of human and machine translation unveil discourse features that set them apart? Target, 34(2), 278–308.
Giampierie, P. (2025). AI-Powered contracts: A critical analysis. International Journal of Semiotics Law, 38, 403-420.
Heiss, C., & Soffritti, M. (2018). DeepL traduttore e didattica della traduzione dall’italiano in tedesco. alcune valutazioni preliminari. In L. Anderson, L. Gavioli & F. Zanettin (Eds.), InTRAlinea. Special Issue: ‘Translation and Interpreting for Language Learners’ (TAIL). https://www.intralinea.org/specials/article/2294
Ilisei, I., & Inkpen, D. (2011). Translationese traits in Romanian newspapers: A machine learning approach. International Journal of Computational Linguistics and Applications, 2(2), 319–332.
Jiao, et al. (2023), Is ChatGPT A Good Translator? Yes with GPT-4
as the Engine. arXiv. https://doi.org/10.48550/arXiv.2301.08745
Killman, J. (2023). Rendering morphosyntactic features of legal Spanish judgments using NMT and SMT. In J. Zhao, D. Li, & V. L. C. Lei (Eds.), New advances in legal translation and interpreting (pp. 221–242). Springer.
Krüger, R. (2020). Explicitation in neural machine translation. Across Languages and Cultures, 21(2), 195-216.
Kruger,H., & Van Rooy, B. (2016). Constrained language. A multidimensional analysis of translated English and a non-native indigenised variety of English. English World-Wide, 37(1), 26–57.
Lapshinova-Koltunski, E. (2015). Variation in translation: evidencie from corpora. En C. Fantinuoli, F. Zanettin (Eds.), New directions in corpus-based translation studies. Language Science Press.
Le Foll, E. (2022). Textbook English: A corpus-based analysis of the language of EFL textbooks used in secondary schools in France, Germany and Spain [Doctoral dissertation, University of Osnabrück]. osnaDocs. https://doi.org/10.48693/278
Le Foll, E., & Shakir, M. (2023). MFTE Python [Computer software]. https://github.com/mshakirDr/MFTE
Le Foll, E., & Shakir, M. (2024). The Multi-Feature Tagger of English (MFTE): Rationale, description and evaluation. Research in Corpus Linguistics, 13(2), 63–93.
Michał Ziemski, Marcin Junczys-Dowmunt, & Bruno Pouliquen. (2016). The United Nations Parallel Corpus v1.0. In N. Calzolari, K. Choukri, T. Declerck, S. Goggi, M. Grobelnik, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, & S. Piperidis (Eds.), Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16) (pp. 3530–3534). European Language Resources Association.
Neumann, S., & Evert, S. (2021). A register variation perspective on varieties of English. In E. Seoane & D. Biber (Eds.), Corpus-based approaches to register variation (pp. 143–178). John Benjamins.
Niu, J. & Jiang, Y. (2024). Does simplification hold true for machine translations? A corpus-based analysis of lexical diversity in text varieties accross genres. Humanities and Social Sciences Commnications, 11, 1-10.
Oliver, A. (2021). MTUCOC-Translator [Computer software]. https://github.com/aoliverg/MTUOC-translator
R Core Team. (2024). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.R-project.org/
Revelle, W. (2024). psych: Procedures for psychological, psychometric, and personality research (Version 2.4.6.26) [Computer software]. Northwestern University. https://cran.r-project.org/package=psych
Rico Pérez, C. (2025). Estudio del efecto de imprimación de la traducción automátcia sobre un corpus de textos del español institucional. Revista de Humanidades Digitales,10: 48-72.
Rico Pérez, C., Sánchez Ramos, M. M., & Oliver, A. (2020). INMIGRA3: Building a case for NGOs and NMT. In A. Martins, H. Moniz, S. Fumega, B. Martins, F. Batista, L. Coheur, C. Parra, I. Trancoso, M. Turchi, A. Bisazza, J. Moorkens, A. Guerberof, M. Nurminen, L. Marg, & M. Forcada (Eds.), Proceedings of the 22nd Annual Conference of the European Association for Machine Translation (pp. 469–471). European Association for Machine Translation.
Şahin, M., & Dungan, N. (2014). Translation testing and evaluation: A study on methods and needs. Translation & Interpreting, 6(2), 67–90.
Toral, A. (2019). Post-editese: An exacerbated translationese. In M. Forcada, A. Way, B. Haddow, & R. Sennrich (Eds.), Proceedings of Machine Translation Summit XVII: Research track (pp. 273–281). European Association for Machine Translation.
Trklja, A., & McAuliffe, K. (2019). Formulaic metadiscursive signalling devices in judgments of the Court of Justice of the European Union: A new corpus-based model for studying discourse relations of texts. International Journal of Speech, Language and the Law, 26(1), 21–55.
Vanmassenhove, E., Shterionov, D., & Way, A. (2019, August 19–23). Lost in translation: Loss and decay of linguistic richness in machine translation [Paper presentation]. Machine Translation Summit XVII, Dublin, Ireland.
Vigier Moreno, F., & Sánchez Ramos, M. M. (2017). Using parallel corpora to study the translation of legal-system bound terms: The case of names of English and Spanish Courts. In R. Mitkov (Ed.), Computational and corpus-based phraseology: Second International Conference, Europhras 2017, London, UK, November 13–14, 2017 proceedings (pp. 260–273). Springer.
Volansky, V., Ordan, N., & Wintner, S. (2015). On the features of translationese. Digital Scholarship in the Humanities, 30(1), 98–118.
Wiesmann, E. (2019). Machine translation in the field of law: A study of the translation of Italian legal texts into German. Comparative Legilinguistics, 37, 117–153.
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Revista Signos. Estudios de Lingüística

This work is licensed under a Creative Commons Attribution 4.0 International License.
Copyright agreement:
Authors who have a manuscript accepted for publication in this journal agree to the following terms:
Authors will retain their copyright and grant the journal the right of first publication of their work by means of this copyright agreement document, which is subject to the Creative Commons Acknowledgment License that allows third parties to share the work provided that its author and first publication in this journal are indicated.
Authors may adopt other non-exclusive license agreements for distribution of the published version of the work (e.g., depositing it in an institutional repository or publishing it in a monographic volume) as long as the initial publication in this journal is indicated.
Authors are allowed and encouraged to disseminate their work via the internet (e.g., in institutional publications or on their website) before and during the submission process, which can lead to interesting exchanges and increase citations of the published work (read more here).