vix.ing · top · new · best · stats · spec

Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions

2025/09/04 by Faruk Alpay, Alpay, Faruk, Taylan Alpay +1
Engineering · #68T05 #68T50 #Advanced Control Systems Optimization #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fault Detection and Control Systems #I.2.11 #I.2.6 #I.2.7

paper · pdf · doi:10.48550/arxiv.2509.04549

openalex publication_date 2025/09/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Transformer-based language models excel in NLP tasks, but fine-grained control remains challenging. This paper explores methods for manipulating transformer models through principled interventions at three levels: prompts, activations, and weights. We formalize controllable text generation as an optimization problem addressable via prompt engineering, parameter-efficient fine-tuning, model editing, and reinforcement learning. We introduce a unified framework encompassing prompt-level steering, activation interventions, and weight-space edits. We analyze robustness and safety implications, including adversarial attacks and alignment mitigations. Theoretically, we show minimal weight updates can achieve targeted behavior changes with limited side-effects. Empirically, we demonstrate >90% success in sentiment control and factual edits while preserving base performance, though generalization-specificity trade-offs exist. We discuss ethical dual-use risks and the need for rigorous evaluation. This work lays groundwork for designing controllable and robust language models.

Citations

Related