In Data or Invisible: Toward a Better Digital Representation of Low-Resource Languages with Knowledge Graphs
This PhD proposal addresses the widening digital divide in Open Access Data between high-resource and low-resource languages, which excludes many communities from global digital transformation. The research focuses on improving language coverage within Linked Open Data knowledge graphs (LOD KGs). Initially, the study identifies key variables characterizing language distribution, such as Wikipedia article counts and language-tagged entities, analyzing them across three major multilingual LOD KGs: DBpedia, BabelNet, and Wikidata. Building on this analysis, the author intends to investigate the impact of cross-lingual transfer candidate selection on multilingual KG completion. Specific strategies under examination include those based on linguistic proximity and curated annotated alignments. Furthermore, the proposal explores the potential of analogical reasoning, leveraging similarities and differences to identify cross-lingual correspondences. This approach aims to enhance KG completion performance and significantly improve language representation in LOD systems. The work was presented at the 23rd European Semantic Web Conference in May 2026, highlighting efforts to make low-resource languages more visible and inclusive in emerging digital technologies and artificial intelligence frameworks.
Wire timeline
In Data or Invisible: Toward a Better Digital Representation of Low-Resource Languages with Knowledge Graphs
This PhD proposal addresses the widening digital divide in Open Access Data between high-resource and low-resource languages, which excludes many communities from global digital transformation. The research focuses on improving language coverage within Linked Open Data knowledge graphs (LOD KGs). Initially, the study identifies key variables characterizing language distribution, such as Wikipedia article counts and language-tagged entities, analyzing them across three major multilingual LOD KGs: DBpedia, BabelNet, and Wikidata. Building on this analysis, the author intends to investigate the impact of cross-lingual transfer candidate selection on multilingual KG completion. Specific strategies under examination include those based on linguistic proximity and curated annotated alignments. Furthermore, the proposal explores the potential of analogical reasoning, leveraging similarities and differences to identify cross-lingual correspondences. This approach aims to enhance KG completion performance and significantly improve language representation in LOD systems. The work was presented at the 23rd European Semantic Web Conference in May 2026, highlighting efforts to make low-resource languages more visible and inclusive in emerging digital technologies and artificial intelligence frameworks.
cs.AI updates on arXiv.org