Jev improves semantic choices for scientific workflow decisions
Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences
Computation and LanguageArtificial Intelligence
Summary
Scientific workflows often need to pick the right meaning for relationships like culture or treatment before calculations can be done. The authors tested a tool called Jev to make these semantic choices and link them to arithmetic steps. They found Jev performed well in correctness and speed compared to alternatives. Sometimes incorrect semantic choices still gave the right final answer, showing why checking all relations and quantities is important. This work highlights Jev's useful role in scientific decision-making tasks.
What this means in practice
- •For scientific workflow engineers: Use Jev to enhance semantic correctness and reduce delays when assigning relations in automated scientific workflows.
- •For clinical data integration teams: Improve accuracy of treatment and culture classifications in clinical datasets by incorporating Jev as a semantic decision component.
Authors
Boyuan Deng, Shuyi Fan, Hongyang Zhang, Xinhong Xie
Abstract
Scientific workflows often require choosing among known relations before a deterministic calculation can proceed. Whether observations share a culture, treatment or reference standard can change the scientific meaning of the resulting count or comparison. We evaluate Jev as a semantic decision component using a harness that follows its documented guidance and assigns arithmetic to code. The study compares twelve model configurations on twenty source-grounded Choices across ten scientific cases, each repeated five times. We measure semantic selections, downstream outputs and final claim labels separately. Jev matched five other configurations at complete semantic correctness and achieved the lowest observed median latency among successful responses. Across three comparison models, seven wrong selections on one culture-history question changed downstream counts while preserving the correct final label. These results identify a useful role for Jev in prepared scientific decision tasks and show why evaluating that role requires checking the relations and quantities that a workflow will reuse.