Episode Details
Back to Episodes
Revisiting Superficial Alignment Hypothesis
Published 1 year, 5 months ago
Description
- The paper revisits the Superficial Alignment Hypothesis.
- It studies post-training scaling behavior with finetuning examples.
- Performance scales as a power law with more finetuning examples.
- Model performance correlates with reasoning ability, not just style.
- Language models can integrate new knowledge post-pre-training.
- Results suggest the hypothesis is an oversimplification.