Episode Details

Back to Episodes
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

Published 2 days, 16 hours ago
Description

This research paper investigates how feature entanglement in large language models prevents precise, localized interventions on specific concepts. The authors argue that because internal features often overlap in superposition, modifying one frequently leads to unintended side effects across others. To solve this, they propose an orthogonality regularization method that forces features to remain nearly independent, aligning with the Independent Causal Mechanisms principle. Theoretical analysis shows that reducing feature interference provides an upper bound on the errors caused by model interventions. Empirical experiments demonstrate that this technique allows for the successful swapping of concepts—such as changing a character's name—without degrading the model’s reasoning performance. Ultimately, the study suggests that promoting geometric orthogonality creates more modular, interpretable, and controllable representations.

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us