https://devinterp.com/
Developmental Interpretability
Website for the developmental interpretability research agenda.
developmentalinterpretability
https://arxiv.org/abs/2501.15740
[2501.15740] Propositional Interpretability in Artificial Intelligence
Abstract page for arXiv paper 2501.15740: Propositional Interpretability in Artificial Intelligence
interpretabilityartificialintelligence
https://www.anthropic.com/research/team/interpretability
Interpretability Research \ Anthropic
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
interpretabilityresearchanthropic
https://hisku.substack.com/p/a-taxonomy-of-training-time-interpretability?open=false
A Taxonomy of Training-Time Interpretability
On designing models to be interpretable by construction, not by post-hoc analysis.
training timetaxonomyinterpretability
https://mechinterpworkshop.com/
Mechanistic Interpretability Workshop at ICML 2026
The Mechanistic Interpretability Workshop at ICML 2026. How can we use the internals of neural networks to understand a model better?
mechanistic interpretabilityworkshopicml
https://machinelearningauthority.com/explainable-ai-services/
Explainable AI and Model Interpretability Services
Explainable AI XAI and model interpretability services address the technical and regulatory challenge of making machine learning model outputs understandable...
explainable aimodel interpretabilityservices
https://docs.google.com/document/d/1qdNP2VHiE9-AObzdgSGAz69AJpaWU6ON3_NlW2zzuTE/edit?tab=t.0
WhiteBox AI Interpretability Fellowship Primer (Cohort 2) - Google Docs
We suggest you view this primer on a laptop/desktop. If you're on mobile, we recommend viewing it on your Google Docs app. WhiteBox AI Interpretability...
ai interpretabilitywhiteboxfellowshipprimercohort
https://pmc.ncbi.nlm.nih.gov/articles/PMC8373813/
Applications of interpretability in deep learning models for ophthalmology - PMC
In this article, we introduce the concept of model interpretability, review its applications in deep learning models for clinical ophthalmology, and discuss...
deep learning modelsfor ophthalmologyapplicationsinterpretabilitypmc
https://www.kaggle.com/discussions/questions-and-answers/519788
Balancing Model Accuracy and Interpretability: How do you navigate this trade-off? | Kaggle
"In machine learning, achieving high model accuracy often conflicts with the need for interpretability. While accurate models may perform well, they can be c...
how do you
https://thezvi.substack.com/p/ai-33
AI #33: Cool New Interpretability Paper - by Zvi Mowshowitz
This has been a rough week for pretty much everyone.
cool newaiinterpretabilitypaperzvi
https://arxiv.org/abs/2504.13151v2
[2504.13151v2] MIB: A Mechanistic Interpretability Benchmark
Abstract page for arXiv paper 2504.13151v2: MIB: A Mechanistic Interpretability Benchmark
mechanistic interpretabilitymibbenchmark
https://new-savanna.blogspot.com/2022/05/beyond-interpretability-developing.html
NEW SAVANNA: Beyond interpretability: developing a language to shape our relationships with AI
Abstract : AI arrived in our lives, making important decisions affecting us. How should we work with this new class of co-workers? The ...
https://www.coursera.org/learn/responsible-ai-for-developers-interpretabilitytransparency
Responsible AI for Developers: Interpretability & Transparency | Coursera
Offered by Google Cloud. This course introduces concepts of AI interpretability and transparency. It discusses the importance of AI ... Enroll for free.
ai for developersresponsibleinterpretabilitytransparencycoursera
https://collaborate.princeton.edu/en/publications/enhancing-interpretability-using-human-similarity-judgements-to-p/fingerprints/?sortBy=alphabetically
Enhancing Interpretability using Human Similarity Judgements to Prune Word Embeddings - Fingerprint...
word embeddingsenhancinginterpretabilityusinghuman
https://ch.mathworks.com/help/stats/interpretability-regression.html?s_tid=CRUX_topnav
Interpretability - MATLAB & Simulink
Train interpretable regression models and interpret complex regression models
interpretabilitymatlabsimulink
https://pubmed.ncbi.nlm.nih.gov/31944251/
Responsiveness and Interpretability of 2 Measures of Physical Function in Patients With...
Our findings suggest that ASPI is preferable over BASFI when evaluating physical function after exercise interventions in patients with axSpA.
physical functionin patientsresponsivenessinterpretability
https://ri.diva-portal.org/smash/record.jsf?pid=diva2:1965872
Interpretability versus performance of analytical and neural-network-based permeability prediction...
neural networkinterpretabilityversusperformanceanalytical
https://arxiv.org/abs/2512.05794
[2512.05794] Mechanistic Interpretability of Antibody Language Models Using SAEs
Abstract page for arXiv paper 2512.05794: Mechanistic Interpretability of Antibody Language Models Using SAEs
mechanistic interpretabilitylanguage modelsantibodyusingsaes
https://arxiv.org/html/2504.13151v2
MIB: A Mechanistic Interpretability Benchmark
mechanistic interpretabilitymibbenchmark
https://arxiv.org/abs/2002.09192v1
[2002.09192v1] An Investigation of Interpretability Techniques for Deep Learning in Predictive...
Abstract page for arXiv paper 2002.09192v1: An Investigation of Interpretability Techniques for Deep Learning in Predictive Process Analytics
an investigation
https://boris-portal.unibe.ch/entities/publication/22d96bb9-f51b-43f7-9c68-2e6ec7d7c927
INFORMER- Interpretability Founded Monitoring of Medical Image Deep Learning Models
medical imagedeep learninginformerinterpretabilityfounded
https://arxiv.org/abs/1802.00614
[1802.00614] Visual Interpretability for Deep Learning: a Survey
Abstract page for arXiv paper 1802.00614: Visual Interpretability for Deep Learning: a Survey
deep learningvisualinterpretabilitysurvey
https://iris.cnr.it/handle/20.500.14243/303092
Insights into Interpretability of Neuro-Fuzzy Systems
insightsinterpretabilityneurofuzzysystems
https://sites.libsyn.com/54799/disentanglement-and-interpretability-in-recommender-systems
Data Skeptic : Disentanglement and Interpretability in Recommender Systems
dataskepticdisentanglementinterpretabilityrecommender
https://www.anthropic.com/research/team/interpretability?ref=en.gusewski.me
Interpretability Research \ Anthropic
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
interpretabilityresearchanthropic
https://ee-damp.github.io/2025-08-07-Abhijat_Bharadwaj_DDP/
(Multiresolution) Signal Processing for Interpretability and Economy in Generative AI
signal processinginterpretabilityeconomygenerativeai
https://cais.usc.edu/tag/interpretability/
interpretability Archives - USC Center for Artificial Intelligence in Society
center forartificial intelligenceinterpretabilityarchivesusc
https://repository.lib.ncsu.edu/items/967c3800-8858-460e-a0b0-917dc6be8985
Innovative strategies for strengthening interpretability of covariance analysis by use of...
innovative strategiesfor strengtheninginterpretabilitycovarianceanalysis
https://ieeetv.ieee.org/fuzzy-rule-based-classifier-design-accuracy-interpretability-and-explanation-ability
Fuzzy Rule-Based Classifier Design: Accuracy, Interpretability and Explanation Ability | IEEETV
Hisao Ishibuchi, Southern University of Science and Technology (SUSTech), Shenzhen, China (hisao@sustech.edu.cn)
fuzzyrulebasedclassifierdesign
https://sites.libsyn.com/54799/interpretability-practitioners
Data Skeptic : Interpretability Practitioners
dataskepticinterpretabilitypractitioners
https://techcommunity.microsoft.com/tag/model%20interpretability?nodeId=board%3AEducatorDeveloperBlog
Tag:"model interpretability" in "Educator Developer Blog" | Microsoft Community Hub
Find all posts, articles, and events tagged with "model interpretability" within Educator Developer Blog in Microsoft Community Hub. Stay informed with the...
model interpretabilitydeveloper blogmicrosoft communitytageducator
https://repository.tudelft.nl/record/uuid:af650ca3-09c9-4cca-a76f-3906bc33d495
Relation between prognostics predictor evaluation metrics and local interpretability SHAP values |...
evaluation metricsrelationprognosticspredictor
https://research.facebook.com/publications/neural-basis-models-for-interpretability/
Neural Basis Models for Interpretability - Meta Research
We propose an architecture denoted as the Neural Basis Model (NBM) which uses a single neural network to learn these bases. On a variety of tabular and image...
neuralbasismodelsinterpretabilitymeta
https://live-cltc.pantheon.berkeley.edu/publication/an-interpretability-study-of-llms-for-code-security/
An Interpretability Study of LLMs for Code Security - CLTC
Large language models (LLMs) such as ChatGPT have greatly advanced coding tasks but often fail to generate secure code. Current approaches to improving code...
code securityinterpretabilitystudyllmscltc
https://rescience.github.io/bibliography/Mohorcic_2023.html
[Re] Hierarchical Shrinkage: Improving the Accuracy and Interpretability of Tree-Based Methods
https://www.ndph.ox.ac.uk/research/research-groups/eph/research/genomics-and-economics-theme/genomics-and-economics-projects/evaluating-the-content-validity-construct-validity-responsiveness-interpretability-feasibility-and-acceptability-of-outcome-measurement-instruments-in-the-context-of-genome-sequencing-for-rare-disease-diagnosis-a-longitudinal-mixed-methods-multi
Evaluating the content validity, construct validity, responsiveness, interpretability, feasibility,...
the contentevaluatingvalidityconstructresponsiveness
https://arxiv.org/abs/2001.09876v1
[2001.09876v1] The POLAR Framework: Polar Opposites Enable Interpretability of Pre-Trained Word...
Abstract page for arXiv paper 2001.09876v1: The POLAR Framework: Polar Opposites Enable Interpretability of Pre-Trained Word Embeddings
https://www.meetup.com/topics/machine-learning-interpretability/rs/
Machine Learning Interpretability groups | Meetup
Find Meetup events so you can do more of what matters to you. Or create your own group and meet people near you who share your interests.
machine learninginterpretabilitygroupsmeetup
https://nemiconf.github.io/summer25/
The 2nd New England Mechanistic Interpretability (NEMI) Workshop
new englandmechanistic interpretabilitynemiworkshop
https://pmc.ncbi.nlm.nih.gov/articles/PMC10707658/
Research on customer churn prediction and model interpretability analysis - PMC
In recent years, with the continuous improvement of the financial system and the rapid development of the banking industry, the competition of the banking...
customer churn predictionmodel interpretabilityresearchanalysispmc
https://eprints.illc.uva.nl/id/eprint/306/
PP-2008-32: Interpretability in PRA - ILLC Preprints and Publications
ppinterpretabilitypraillcpreprints
https://jp.mathworks.com/help/deeplearning/visualization-and-interpretability.html?s_tid=CRUX_topnav
Visualization and Interpretability - MATLAB & Simulink
Plot training progress, assess accuracy, explain predictions, and visualize features learned by a network
visualizationinterpretabilitymatlabsimulink
https://www.illc.uva.nl/Research/Publications/Publications-by-year/publication/3049/Provability-Logics-for-Relative-Interpretability
Provability Logics for Relative Interpretability | Institute for Logic, Language and Computation
provabilitylogicsrelativeinterpretabilityinstitute
https://ch.mathworks.com/fr/discovery/interpretability.html
Interpretability - MATLAB & Simulink
Learn about interpretability: how it works, why it matters, and how to use MATLAB to perform interpretability. Resources include videos, examples, and...
interpretabilitymatlabsimulink
https://thesequence.substack.com/p/the-sequence-radar-531-the-need-for
The Sequence Radar #531: The Need for AI Interpretability
Anthropic's CEO message about one of the most important challenges in generative AI.
the sequenceneed forradaraiinterpretability
https://exec-ed.berkeley.edu/tag/ai-interpretability/
AI interpretability Archives - UC Berkeley Professional Education
ai interpretabilityuc berkeleyarchivesprofessionaleducation
https://inventions.techventures.columbia.edu/technologies/evaluating--CU21011
Evaluating robustness and interpretability of AI models for disease detection
A standardized, interpretable AI evaluation framework for robust disease detection in medical imaging, aligning models with expert feedback.
ai modelsevaluatingrobustnessinterpretabilitydisease
https://profiles.wustl.edu/en/publications/the-interpretability-of-family-history-reports-of-alcoholism-in-g/
The Interpretability of Family History Reports of Alcoholism in General Community Samples: Findings...
family history
https://jobs.inria.fr/public/classic/fr/offres/2026-09862
2026-09862 - PhD Position F/M Mechanistic Interpretability and Problem-Space Adversarial Attacks...
Offre d'emploi Inria
https://developer.nvidia.com/gtc/2019/video/s9249
GTC Silicon Valley-2019: Practical Machine Learning Interpretability Techniques | NVIDIA Developer
silicon valleymachine learninggtcpractical
https://employment.ku.dk/phd/?show=160571
PhD fellowship in Mechanistic Interpretability for LLM Security
phd fellowshipmechanistic interpretabilityfor llmsecurity
https://lilywenglab.github.io/cvpr2026-principled-interpretability-tutorial/
Principled Interpretability in Vision Models | Principled Interpretability in Vision Models
in visionprincipledinterpretabilitymodels
https://hspop.uw.edu/publication/presentation-approaches-for-enhancing-interpretability-of-patient-reported-outcomes-pros-in-meta-analysis-a-protocol-for-a-systematic-survey-of-cochrane-reviews/
Presentation approaches for enhancing interpretability of patient-reported outcomes (PROs) in...
Devji T, Johnston BC, Patrick DL, Bhandari M, Thabane L, Guyatt GH. Presentation approaches for enhancing interpretability of patient-reported outcomes (PROs)...
patient reported outcomespresentationapproachesenhancinginterpretability
https://bg.copernicus.org/articles/21/2051/2024/
BG - Interpretability of negative latent heat fluxes from eddy covariance measurements in dry...
Abstract. It is known from arid and semi-arid ecosystems that atmospheric water vapor can directly be adsorbed by the soil matrix. Soil water vapor adsorption...
https://pmc.ncbi.nlm.nih.gov/articles/PMC9763801/
Condition-based maintenance using machine learning and role of interpretability: a review - PMC
This article aims to review the literature on condition-based maintenance (CBM) by analyzing various terms, applications, and challenges. CBM is a maintenance...
condition based maintenance
https://arxiv.org/abs/2304.06919
[2304.06919] Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary...
Abstract page for arXiv paper 2304.06919: Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary Defense
https://www.kellogg.northwestern.edu/academics-research/research/detail/1987/improper-solutions-in-the-analysis-of-covariance-structures-their/
Improper Solutions in the Analysis of Covariance Structures: Their Interpretability and a...
A Monte Carlo approach was employed to investigate the interpretability of improper solutions caused by sampling error in maximum likelihood confirmatory...
in the
https://www.cambridge.org/core/journals/review-of-symbolic-logic/article/when-biinterpretability-implies-synonymy/00B8CAF9978904070D017C303308F414
WHEN BI-INTERPRETABILITY IMPLIES SYNONYMY | The Review of Symbolic Logic | Cambridge Core
WHEN BI-INTERPRETABILITY IMPLIES SYNONYMY - Volume 18 Issue 4
the review
https://msclogic.illc.uva.nl/theses/recent/publication/4161/Supremum-in-the-Lattice-of-Interpretability
Supremum in the Lattice of Interpretability | Master of Logic
in thelatticeinterpretabilitymasterlogic
https://arxiv.org/abs/2503.06269v2
[2503.06269v2] Using Mechanistic Interpretability to Craft Adversarial Attacks against Large...
Abstract page for arXiv paper 2503.06269v2: Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
mechanistic interpretabilityadversarial attacksusing
https://jobs.apple.com/nl-be/details/200626484-1242/aiml-research-scientist-ai-interpretability-visualization
AIML - Research Scientist, AI Interpretability & Visualization - Vacatures bij Apple (BE)
research scientistai interpretabilityaimlvisualizationvacatures
https://academiccommons.columbia.edu/doi/10.7916/d8-hfry-nr98
Network Structures, Concurrency, and Interpretability: Lessons from the Development of an AI...
This thesis describes the development of the SmartGraph, an AI enabled graph database. The need for such a system has been independently recognized in the...
https://nlp.stanford.edu/~wuzhengx/boundless_das/index.html
Scaling interpretability with LLMs
We propose a new method based on the theory of causal abstraction to find representations that play a given causal role in LLMs
scalinginterpretabilityllms
https://bcmullins.github.io/economic_methodology_interpretable_ml_blackboxes/
Economic Methodology Meets Interpretable Machine Learning - Part I - Interpretability,...
This post is the first entry in Economic Methodology Meets Interpretable Machine Learning and briefly introduces the ideas of black boxes, explainability, and...
economic methodologymachine learningmeetspartinterpretability
https://www.techtarget.com/searchenterpriseai/feature/Interpretability-vs-explainability-in-AI-and-machine-learning
Interpretability vs. explainability in AI and machine learning | TechTarget
Learn the key differences between interpretability and explainability in AI and machine learning, and explore examples, techniques and limitations.
ai and machine learninginterpretabilityvsexplainabilitytechtarget
https://par.nsf.gov/biblio/10657157-towards-global-level-mechanistic-interpretability-perspective-modular-circuits-large-language-models
Towards Global-level Mechanistic Interpretability: A Perspective of Modular Circuits of Large...
This page contains metadata information for the record with PAR ID 10657157
global levelmechanistic interpretabilitytowards
https://explaining.ml/
StrategyAtlas: Strategy Analysis for Machine Learning Interpretability
strategy analysismachine learninginterpretability
https://colah.github.io/notes/interp-v-neuro/
Interpretability vs Neuroscience [rough note] -- colah's blog
A list of advantages that make understanding artificial nerural networks much easier than biological ones.
interpretabilityvsneuroscienceroughnote
https://emploi.cnrs.fr/Offres/Doctorant/UMR5217-MAXPEY-002/Default.aspx?lang=EN
Portail Emploi CNRS - Job offer - PhD Thesis: Interpretability and Evaluation of LLMs and Agentic...
https://codelabs.developers.google.com/codelabs/responsible-ai/lit-on-gcp
LLM prompt debugging with the Learning Interpretability Tool (LIT) on GCP | Google Codelabs
https://www1-prod.cs.uchicago.edu/events/event/distinguished-lecture-series-been-kim-google-deepmind-alignment-and-interpretability-how-we-might-get-it-right/
Distinguished Lecture Series: Been Kim (Google DeepMind)- Alignment and interpretability: how we...
Part of the 2024-25 DSI Distinguished Speaker Series and the Computer Science Distinguished Lecture Series. Abstract: The main goal of interpretability is to...
distinguished lecture seriesgoogle deepmind