LLM serving decisions shaped by multiple sustainability impacts
Beyond Energy: When Sustainability Dimensions Reshape LLM Serving Decisions
Computers and SocietyDistributed, Parallel, and Cluster ComputingMachine LearningPerformance
Summary
Running large language models (LLMs) impacts the environment in many ways like energy use, carbon emissions, water use, and harming biodiversity. The authors show that how much energy a computer setup uses is separate from where and when the LLM is run, which affects carbon and water impacts differently. They created PRISM, a framework that considers all these factors together to help pick better ways to serve LLMs that balance these impacts. Their tests show PRISM can significantly reduce mistakes people make when choosing setups based on just one impact type.
What this means in practice
- •For cloud infrastructure teams: Improve LLM deployment choices by balancing energy, carbon, water, and biodiversity impacts to reduce environmental trade-offs in data centers.
- •For environmental policy makers: Assess and guide LLM practices by considering regional and temporal effects on sustainability dimensions beyond just energy use.
Authors
Tianyao Shi, Xipeng Shen, Yi Ding
Abstract
Large language model (LLM) serving has environmental impacts across energy consumption, carbon emission, water consumption, and biodiversity loss. Yet these dimensions are largely evaluated in isolation, leaving it unclear when and how they lead to different optimization decisions. We present PRISM, a unified framework for characterizing and optimizing LLM serving across energy, carbon, water, and biodiversity impacts. Our analysis reveals a fundamental distinction: computing configurations determine energy consumption, whereas where and when LLM serving is deployed determine its carbon, water, and biodiversity impacts. Under a fixed deployment choice and operational-only accounting, all dimensions preserve the same energy-based configuration ranking. Deployment rankings can diverge across dimensions, while embodied impacts can break configuration invariance when they exceed a lifecycle crossover boundary. PRISM identifies these conditions, quantifies cross-dimensional regrets, and balances the four dimensions. In regional-routing experiments, PRISM reduces median worst-case regret by 50.2% relative to the strongest baseline.