vix.ing · top · new · best · stats · spec

Assessing the feasibility of Large Language Models for detecting micro-behaviors in team interactions during space missions

2025/06/27 by Anurag Raut, Raut, Ankush, Projna Paromita +7
Medicine · Neuroscience · Psychology · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Language Development and Disorders #Neurobiology of Language and Bilingualism #Spaceflight effects on biology

paper · pdf · doi:10.48550/arxiv.2506.22679

openalex publication_date 2025/06/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We explore the feasibility of large language models (LLMs) in detecting subtle expressions of micro-behaviors in team conversations using transcripts collected during simulated space missions. Specifically, we examine zero-shot classification, fine-tuning, and paraphrase-augmented fine-tuning with encoder-only sequence classification LLMs, as well as few-shot text generation with decoder-only causal language modeling LLMs, to predict the micro-behavior associated with each conversational turn (i.e., dialogue). Our findings indicate that encoder-only LLMs, such as RoBERTa and DistilBERT, struggled to detect underrepresented micro-behaviors, particularly discouraging speech, even with weighted fine-tuning. In contrast, the instruction fine-tuned version of Llama-3.1, a decoder-only LLM, demonstrated superior performance, with the best models achieving macro F1-scores of 44% for 3-way classification and 68% for binary classification. These results have implications for the development of speech technologies aimed at analyzing team communication dynamics and enhancing training interventions in high-stakes environments such as space missions, particularly in scenarios where text is the only accessible data.

Citations

Related