Tagged interpretability
2 essays
- Mechanistic interpretability as generative art Concepts inside a neural network have addresses. Point an image generator at the address for 'ocean' and let it run, and what comes out is the model's own idea of ocean, rendered by the model rather than charted.
- Measuring the platonic representation There's an idea that as models get better they all end up learning the same picture of the world, whatever they were trained on. You can test it: describe one thing as text, an image, a sound and a video, then see how close those four land.