Representation alignment improves chemical synthesis planning with large language models
Representation Alignment as a Bottleneck in LLM-Based Retrosynthesis Planning
Artificial IntelligenceComputation and Language
Summary
Planning how to make chemical compounds is hard for large language models, especially when they try to do everything at once. The authors found that breaking this task into smaller steps—matching molecules, mapping reactions, and then generating plans—helps the models succeed. This shows that the main challenge is how information is represented and passed between steps, not how powerful the models are. Using these careful intermediate steps could make future chemical planning tools work better.
What this means in practice
- •For chemical software developers: Create retrosynthesis planning tools that improve success by breaking down tasks into intermediate molecule and reaction mappings before plan generation.
- •For ai system architects: Design AI workflows for complex planning problems that use intermediate representations to improve alignment between reasoning stages.
Authors
Hyunwoo Yoo, Cassie Huang, Haebin Shin, Li Zhang, Gail L. Rosen
Abstract
While LLMs show promise in general reasoning, symbolic planning in chemistry remains a bottleneck. Direct ''SMILES-to-PDDL'' attempts fail because they force models to juggle chemical analysis and planning-language structuring simultaneously. We hypothesize that this failure stems from a lack of intermediate abstractions rather than insufficient model capacity. By decomposing retrosynthesis into molecule mapping, reaction mapping, and PDDL generation, we achieve high success rates where end-to-end approaches fail. This provides evidence that a primary bottleneck lies in representation alignment rather than raw model capacity. Our structural analysis demonstrates that intermediate representations are essential in retrosynthesis planning, highlighting the importance of representation-centric design in future systems.