Factorengram improves language models by sharing and gating memory components
FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
Computation and LanguageArtificial Intelligence
Summary
Language models often remember phrases and word patterns by storing them as single chunks, which limits their ability to pick out the right meaning depending on the sentence. The authors propose a new method called FactorEngram that breaks these chunks into smaller shared parts and lets the context decide which parts to use. This makes the model better at understanding and predicting language by reusing components across patterns and adjusting their importance depending on the surrounding words. Their experiments show improvements in language tasks and help find the best way to add this memory system into existing models.
What this means in practice
- •For natural language processing engineers: Enhance language understanding systems by integrating factorized memory components that selectively retrieve relevant word patterns.
- •For conversational ai developers: Improve response relevance in chatbots by using context-driven gating to better represent polysemous phrases and multi-word patterns.
Authors
Bowen Yang, Jingbo Zhou, Qinghong Miao, Hua Wu
Abstract
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.