Vocabulary Growth Fundamentals: Bernstein Functions and Hausdorff Sequences

Computation and Language

Summary

The authors review how vocabulary growth can be understood using certain mathematical functions called Bernstein functions and Hausdorff sequences, which connect to specific random processes. They link past studies with newer models that focus on the rate of unique words appearing once (hapax legomena). Notably, they prove that a logistic model for hapax rates fits nicely within the Bernstein function framework, addressing a previous question. The authors also explore how these theories might change or break down when more complex random processes are considered.

Authors

Łukasz Dębowski

Abstract

We survey the theory of vocabulary growth founded in the setting of stochastic processes. In particular, we model the expected number of types through Bernstein functions and Hausdorff sequences. These classes of mathematical objects, defined by alternating signs of their derivatives or differences, can be related to continuous-time Poisson point processes and discrete-time IID processes, respectively. Building on previous accounts of the vocabulary growth, we integrate the broader theories of Bernstein functions and Hausdorff sequences and connect them with recently developed hapax rate models. In particular, we prove that the logistic hapax rate model has a non-negative spectrum and hence it defines a Bernstein function, thereby solving an earlier posed problem. We also analyze the limitations of the Bernstein--Hausdorff theory of the vocabulary growth by considering its generalizations under stationary and Weibull renewal processes.