Ai enabled tools improve scalable measurement of durable skills
Towards Scalable Measurement of Durable Skills
Human-Computer Interaction
Summary
Measuring important skills like creativity and teamwork is hard because tests need to feel like real life but also be fair and repeatable. The authors created a system where people chat with AI teammates that act realistically and help guide conversations to show these skills clearly. They also use AI to score how well someone demonstrates these skills. Their tests on human-AI chats show the method works well and matches expert grading, making skill measurement easier and more reliable.
What this means in practice
- •For corporate training teams: Use AI teammate systems to assess employee soft skills in scalable, reproducible ways during training sessions.
- •For online education platforms: Integrate automated AI evaluators to grade student creativity and collaboration during interactive learning tasks.
Authors
Amir Globerson, Amy Keeling, Anisha Choudhury, Anna Iurchenko, Aviad Segal, Avinatan Hassidim, Ayça Çakmakli, Ben Gomes, Benn Witt, Cathy Cheunga, Cristine Legare, Diana Akrong, Eliad Carmi, Elisabeth Bauer, Gal Elidan, Hadas Gelbart, Hairong Mu, Katherine Chou, Lev Borovoi, Nir Kerem, Niv Efron, Noa Kerrem Gilo, Preeti Singh, Rajvi Kapadia, Rena Levitt, Roni Rabin, Ronit Levavi Morad, Rotem Yulzary, Shashank Agarwal, Sophie Allweis, Tracey Lee-Joe, Tzvika Stein, Yael Bar Moshe, Yael Haramaty, Yaniv Carmel, Yishay Mor, Yoav Bar Sinai, Yoav Bergner, Yossi Matias, Yuri Lev
Abstract
Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural real-world human interaction between humans. On the other hand, it should be scalable, controllable and reproducible. Here we argue that LLMs can be used to better capture both of these aims. Concretely, we develop a framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an "Executive LLM" setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency. We complement this with an AI evaluator that can be used to measure skill proficiency in such interactions. We evaluate our assessment protocol based on transcripts of interactions of human participants with our AI framework, for multiple durable skills. For the skill of creativity, we further demonstrate the efficacy of an autorater for evaluating complex tasks performed by real students. Our analysis shows that the use of the Executive LLM significantly increases elicited evidence and that LLM-automated scoring of conversations largely agrees with that of expert annotators. This research demonstrates the utility of orchestrated LLMs approaches for measuring complex social and cognitive constructs in a scalable and controllable manner.