Robuta

https://hai.stanford.edu/research/finding-monosemantic-subspaces-and-human-compatible-interpretations-in-vision-transformers-through-sparse-coding Finding Monosemantic Subspaces and Human-Compatible Interpretations in Vision Transformers through... We present a new method of deconstructing class activation tokens of vision transformers into a new, overcomplete basis, where each basis vector is... human compatiblein visionfinding