Contact
Privacy
DMCA
Robuta
https://mechinterpworkshop.com/
Mechanistic Interpretability Workshop at ICML 2026
The Mechanistic Interpretability Workshop at ICML 2026. How can we use the internals of neural networks to understand a model better?
mechanistic interpretability
workshop
icml
https://arxiv.org/abs/2504.13151v2
[2504.13151v2] MIB: A Mechanistic Interpretability Benchmark
Abstract page for arXiv paper 2504.13151v2: MIB: A Mechanistic Interpretability Benchmark
mechanistic interpretability
mib
benchmark
https://arxiv.org/abs/2512.05794
[2512.05794] Mechanistic Interpretability of Antibody Language Models Using SAEs
Abstract page for arXiv paper 2512.05794: Mechanistic Interpretability of Antibody Language Models Using SAEs
mechanistic interpretability
language models
antibody
using
saes
https://arxiv.org/html/2504.13151v2
MIB: A Mechanistic Interpretability Benchmark
mechanistic interpretability
mib
benchmark
https://nemiconf.github.io/summer25/
The 2nd New England Mechanistic Interpretability (NEMI) Workshop
new england
mechanistic interpretability
nemi
workshop
https://employment.ku.dk/phd/?show=160571
PhD fellowship in Mechanistic Interpretability for LLM Security
phd fellowship
mechanistic interpretability
for llm
security
https://arxiv.org/abs/2503.06269v2
[2503.06269v2] Using Mechanistic Interpretability to Craft Adversarial Attacks against Large...
Abstract page for arXiv paper 2503.06269v2: Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
mechanistic interpretability
adversarial attacks
using
https://par.nsf.gov/biblio/10657157-towards-global-level-mechanistic-interpretability-perspective-modular-circuits-large-language-models
Towards Global-level Mechanistic Interpretability: A Perspective of Modular Circuits of Large...
This page contains metadata information for the record with PAR ID 10657157
global level
mechanistic interpretability
towards
https://jobs.inria.fr/public/classic/fr/offres/2026-09862
2026-09862 - PhD Position F/M Mechanistic Interpretability and Problem-Space Adversarial Attacks...
Offre d'emploi Inria