The Multilingual FrameNet Corpus
2026-08-24 • Computation and Language
Computation and LanguageArtificial Intelligence
AI summaryⓘ
The authors created a new collection of language data called the Multilingual FrameNet Corpus (mFNC) by combining information from nine different languages with an existing English dataset. They trained computers to understand meaning in sentences using this new multilingual data. Their models worked better than previous ones in recognizing sentence meaning across multiple languages. The authors have made their dataset and models freely available for others to use.
FrameNetFrame Semantic ParsingMultilingual CorpusCross-lingualNatural Language ProcessingSemantic RolesLanguage ResourcesMachine Learning ModelsMultilingual Training Data
Authors
Beatrice Fiumanò, Nicolas Lazzari, Simone Paolo Ponzetto, Valentina Presutti
Abstract
This paper introduces the Multilingual FrameNet Corpus (mFNC), a novel resource that extends the English Berkeley FrameNet corpus by collecting and harmonizing existing language-specific corpora across nine additional languages: Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian and Swedish. By training models that rely on different architectures on the mFNC, we consistently outperform existing state-of-the-art Frame Semantic Parsers in both multilingual and cross-lingual settings, underscoring the importance of multilingual training data. The mFNC and our trained FSP models are openly available at https://github.com/beatrice-f/mFNC.