vix.ing · top · new · best · stats

Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering

2024/12/10 by Rumi A. Allbert, Rumi Allbert, Allbert, Rumi +4 · 3 voices · 5 citations
Computer Science · Psychology · #Big Five personality traits #Online Learning and Analytics #Personality #Political science #Psychology #Social psychology

paper · pdf · doi:10.48550/arxiv.2412.10427

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2024/12/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

The field of large language models (LLMs) has grown rapidly in recent years, driven by the desire for better efficiency, interpretability, and safe use. Building on the novel approach of "activation engineering," this study explores personality modification in LLMs, drawing inspiration from research like Refusal in LLMs Is Mediated by a Single Direction (arXiv:2406.11717) and Steering Llama 2 via Contrastive Activation Addition (arXiv:2312.06681). We leverage activation engineering to develop a method for identifying and adjusting activation directions related to personality traits, which may allow for dynamic LLM personality fine-tuning. This work aims to further our understanding of LLM interpretability while examining the ethical implications of such developments.

Citations

Cited by

Discussions

Related