vix.ing · top · new · best · stats · spec

View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior

2024/07/01 by Tanush Chopra, Chopra, Tanush, Michael Li +2 · 2 citations
Business, Management and Accounting · Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Ethics in Business and Education #FOS: Computer and information sciences #Information and Cyber Security #Machine Learning (cs.LG) #Securities Regulation and Market Practices

paper · pdf · doi:10.48550/arxiv.2407.00948

openalex publication_date 2024/07/01 · openalex created_date 2024/07/06 · openalex updated_date 2026/07/28

Abstract

When large language models (LLMs) are asked to perform certain tasks, how can we be sure that their learned representations align with reality? We propose a domain-agnostic framework for systematically evaluating distribution shifts in LLMs decision-making processes, where they are given control of mechanisms governed by pre-defined rules. While individual LLM actions may appear consistent with expected behavior, across a large number of trials, statistically significant distribution shifts can emerge. To test this, we construct a well-defined environment with known outcome logic: blackjack. In more than 1,000 trials, we uncover statistically significant evidence suggesting behavioral misalignment in the learned representations of LLM.

Cited by

Related