Deep Learning Framework Analyzes Grammatical Gender Shift from Latin to Occitan
Researchers have introduced a new interpretable deep learning framework to investigate the diachronic evolution of grammatical gender systems during the transition from Latin to Romance languages, specifically focusing on Occitan. This linguistic shift involved restructuring from a tripartite system (masculine, feminine, neuter) to a bipartite one (masculine, feminine). The study addresses challenges in low-resource historical settings by demonstrating that conventional tokenization strategies are insufficiently robust. The authors propose a specialized tokenizer that significantly improves performance over baseline methods. The framework analyzes gender prediction at both lexical and contextual levels. At the lexical level, it evaluates the contribution of morphological features, while at the contextual level, it quantifies how different part-of-speech categories influence gender prediction. These analyses collectively characterize the distribution of gender information between the lemma and its sentential context. To foster further research and reproducibility, the team has made their codebase, datasets, and results publicly available. This work bridges computational linguistics and historical philology, offering new tools for understanding language evolution through advanced artificial intelligence techniques.
Wire timeline
Deep Learning Framework Analyzes Grammatical Gender Shift from Latin to Occitan
Researchers have introduced a new interpretable deep learning framework to investigate the diachronic evolution of grammatical gender systems during the transition from Latin to Romance languages, specifically focusing on Occitan. This linguistic shift involved restructuring from a tripartite system (masculine, feminine, neuter) to a bipartite one (masculine, feminine). The study addresses challenges in low-resource historical settings by demonstrating that conventional tokenization strategies are insufficiently robust. The authors propose a specialized tokenizer that significantly improves performance over baseline methods. The framework analyzes gender prediction at both lexical and contextual levels. At the lexical level, it evaluates the contribution of morphological features, while at the contextual level, it quantifies how different part-of-speech categories influence gender prediction. These analyses collectively characterize the distribution of gender information between the lemma and its sentential context. To foster further research and reproducibility, the team has made their codebase, datasets, and results publicly available. This work bridges computational linguistics and historical philology, offering new tools for understanding language evolution through advanced artificial intelligence techniques.
cs.AI updates on arXiv.org