Episode Details
Back to Episodes
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
Description
This research paper investigates how feature entanglement in large language models prevents precise, localized interventions on specific concepts. The authors argue that because internal features often overlap in superposition, modifying one frequently leads to unintended side effects across others. To solve this, they propose an orthogonality regularization method that forces features to remain nearly independent, aligning with the Independent Causal Mechanisms principle. Theoretical analysis shows that reducing feature interference provides an upper bound on the errors caused by model interventions. Empirical experiments demonstrate that this technique allows for the successful swapping of concepts—such as changing a character's name—without degrading the model’s reasoning performance. Ultimately, the study suggests that promoting geometric orthogonality creates more modular, interpretable, and controllable representations.