Apolo improves ontology learning from text using optimized prompts
APOLO: Automatic Prompt Optimization for Ontology Learning
Artificial Intelligence
Summary
Building structured knowledge from text, known as ontology learning, is hard because there isn’t much labeled data to train on. The authors created APOLO, a way to automatically improve prompts given to large language models to better extract this knowledge. They generate training examples using multiple agents that pair text with expert-made ontologies. Then, they try two methods to learn ontologies and optimize their prompts with a special evolutionary approach. Their experiments show that prompt optimization helps, especially with one method that builds ontologies step-by-step.
What this means in practice
- •For biomedical data teams: Automatically generate or improve biomedical ontologies by optimizing prompts for large language models without retraining them.
- •For agricultural knowledge engineers: Improve plant ontology construction by applying prompt-optimized systems that better capture complex ontology structures.
Authors
Huu Tan Mai, Roman Kochnev, Cuong Xuan Chu, Lukas Lange, Heiko Paulheim, Daria Stepanova
Abstract
Ontology Learning (OL) from text has advanced with the emergence of Large Language Models (LLMs), but it remains challenging due to the limited availability of annotated training data and the difficulty of adapting LLMs to perform OL effectively. We address this via APOLO - Automatic Prompt Optimization for Ontology Learning, by casting OL as an explicit prompt optimization problem over LLM modules. To obtain training data, we employ a multi-agent system that generates text-ontology pairs from existing expert-curated ontologies. We then propose two ontology learner architectures: a greedy and an autoregressive learner, and optimize both using GEPA, a greedy evolutionary prompt optimizer built on DSPy. Experiments on two ontologies - a biomedical (DOID) and a plant ontology (PO) show consistent improvements after optimization across nearly all model and mode combinations, with autoregressive learners achieving the largest gains. Our results demonstrate that prompt optimization is a viable and lightweight alternative to fine-tuning for OL, and that the autoregressive formulation better captures ontological structure than the greedy approach.