Structural similarity improves cross language scientific sentence classification
Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity
Computation and LanguageDigital Libraries
Summary
Scientific papers often organize sentences into sections like introduction, methods, and results, which follow similar patterns across languages. This paper shows that using the structural similarities of these patterns helps computers better understand and classify sentences in scientific papers written in different languages. The authors created a multilingual dataset and found that language closeness matters less than how sentences are organized. They developed methods that use this structure information and improved classification accuracy when transferring models to languages without training data.
What this means in practice
- •For digital library developers: Improve automatic structuring and classification of scientific papers in multiple languages to enhance search and accessibility.
- •For multilingual nlp engineers: Build cross-lingual sentence classifiers that perform better by incorporating the structural patterns of scientific writing.
Authors
Kazuhiro Yamauchi, Marie Katsurai
Abstract
Sequential sentence classification (SSC) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries. Cross-lingual transfer is a promising approach to address the scarcity of training data in non-English languages. Prior work on other natural language processing tasks has shown the benefits of capturing linguistic similarity between source and target languages. However, SSC inherently depends on patterns at the discourse level, such as label sequences and positional regularities, which appear consistently across languages regardless of linguistic differences. To examine the factors that determine transfer success in SSC, we constructed a multilingual SSC dataset covering 13 non-English languages collected from five academic databases. Our cross-lingual transfer experiments, using both encoder-based and generative models, show that linguistic proximity has no consistent predictive power for transfer performance, whereas structural similarity in rhetorical organization shows a weak but consistent positive correlation across models. After controlling for source-language performance, the similarity of label distributions is the most consistent predictor. Building on this finding, we propose a set of three methods that explicitly leverage structural information using generative models. In the in-domain evaluation, the best combination reaches parity with the strongest encoder baselines, and in transfer to languages unseen during training, it outperforms the strongest encoder baseline.