Greenhouse climate control improves yields and safety with adaptive learning

Safe Greenhouse Climate Control Using Lagrangian-Constrained PPO with Kolmogorov-Arnold Networks

Artificial Intelligence

Summary

Greenhouses need to keep temperature, humidity, and CO2 just right for plants to grow well. The paper finds that using a special kind of artificial intelligence model that follows strict safety rules helps keep these factors within safe ranges more reliably than older methods. This approach adjusts its rules automatically to avoid hurting plant growth, while also improving profits. The method also uses a new type of neural network that understands complicated patterns better. Tests on lettuce-growing simulations show it reduces climate problems and boosts economic returns.

What this means in practice

  • For greenhouse managers: Use adaptive reinforcement learning to maintain crop-friendly climate conditions and improve yield profitability in greenhouse operations.
  • For precision agriculture system developers: Implement safer RL controllers with nonlinear network models to reduce operational risks and improve crop performance under changing environmental conditions.

Authors

Hangzun Liu, Yuling Fan, Fang Tian, Zhilong Bie, Zaiwen Feng, Yongliang Qiao

Abstract

Greenhouse climate control balances economic return with maintaining temperature, humidity and CO2 within crop-adapted growth ranges. Conventional reinforcement learning (RL) greenhouse controllers use fixed reward penalties to limit climate constraint violations, yet such heuristic penalties cannot explicitly constrain long-term cumulative violations. Poorly tuned weights either lead to overly conservative policies and lower yields, or fail to suppress persistent climate deviations that harm photosynthesis and induce crop diseases. To address this issue, we formulate greenhouse climate regulation as a Constrained Markov Decision Process (CMDP) and use a Lagrangian safe RL framework RCPO-PPO to separate economic optimization and cumulative safety constraints, enabling adaptive penalty adjustment without manual tuning. To handle strong nonlinear, time-varying coupling between greenhouse microclimate and crop growth, Kolmogorov-Arnold Networks (KANs) replace Multi-Layer Perceptrons (MLPs) as policy and value approximators for improved nonlinear representation. Sinusoidal cyclic time features are embedded in observations to capture diurnal environmental periodicity. Simulations use a classic winter lettuce greenhouse model driven by 40-day real weather disturbances. Compared with vanilla penalty-based PPO, our method cuts cumulative climate violations by 18.65% and raises lettuce economic profit by 2.91%, keeping violations stable near the safety threshold. This decoupled CMDP optimization with KAN-based policy representation mitigates long-term climate risks and boosts planting profits, offering a constraint-aware control strategy for precision greenhouse cultivation.