Computational method finds biblical references in karen blixen stories

Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales

Computation and Language

Summary

Finding hidden references to the Bible in literature can be hard, especially when the text doesn't quote directly but hints or paraphrases. The authors worked on spotting these references in Karen Blixen's Seven Gothic Tales by comparing story passages to old Danish Bible translations. They tested different methods and improved one by training it with examples, which helped spot many more subtle references. Their approach serves as a helpful tool to suggest possible connections for scholars to explore further, rather than replacing expert reading.

What this means in practice

  • For digital humanities teams: Assist scholars by automatically retrieving likely biblical references in historical literary texts for further expert analysis.
  • For language technology developers: Develop improved natural language search tools that detect paraphrased and allusive text references in multiple languages.

Tested on one dataset.

Authors

András Kovács, Alexander Conroy, Daniel Hershcovich, Jens Bjerring-Hansen

Abstract

Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Tales. Drawing on the commentary to a critical edition, we construct a benchmark of 189 annotated references and evaluate retrieval against all 31,170 verses of historically plausible Danish Old and New Testament translations. We compare TF-IDF and BM25 with multilingual and Danish sentence encoders, examine the effect of linguistic normalization, and fine-tune a Danish encoder using hard negatives and five-fold cross-validation. We analyze performance across automatically derived lexical-overlap strata representing quotations, paraphrases, and allusions. Linguistically normalized BM25 provides a strong zero-shot baseline, attaining an overall R@10 of 0.365 and retrieving every quotation within its ten highest-ranked verses. The best zero-shot dense model achieves a comparable overall score of 0.360 while performing better on allusions. Fine-tuning DFM-large raises its overall R@10 from 0.265 to 0.508 and more than doubles its performance on allusions, from 0.138 to 0.339. However, evaluation against editorial annotations alone understates the model's scholarly usefulness: a literary scholar judged seven of 30 selected rank-one predictions counted as false positives to be meaningful additional references. These findings show both the potential and the epistemic limits of computational intertextual retrieval. Rather than treating scholarly annotations as exhaustive or model outputs as discoveries, we propose retrieval models as heuristic co-readers that recover documented references and generate candidates for expert-led close reading.