for interpretability maybe without all of those expensive human subject evaluations?
So interpretability is this thing that you do, mostly probably in the service of something else.
And interpretability is really important because we are looking at large amounts of data, but at the end of the day, we actually
The interpretability problem, in some sense, is done.
And interpretability has succeeded if it allows you to do the higher-level task.
of mechanistic interpretability , which is an attempt to understand what's going on inside AI models.
The other is local interpretability , which is the ability to say, for this particular decision, the five most important things were these.
How can you measure interpretability not necessarily in the absence, but in the abstract,
But mechanical interpretability is good enough that you may be able to like find what is a solid theory for depression
like scaling laws applied to interpretability , or scaling laws applied to post-training, or just seeing how does this thing scale.
One is what's called global interpretability , which is the ability to say at an aggregate level, at a population level, this deep learning model
Hence this growing interest in interpretability .
And we don't have that in interpretability .
Another reason why you might want interpretability is to be able to debug your system.
I think the case where you need interpretability is where there's something wrong with the specification.
Because the whole point about interpretability is that it makes sense to human beings.
He's one of the pioneers of the field of mechanistic interpretability , which is an exciting set of efforts that aims to reverse engineer neural networks
and then you see soon that there's interpretability teams elsewhere as well.
- And we should say this example of the field of mechanistic interpretability is just a rigorous, non-hand wavy way of doing AI safety,
And there's a lot of emphasis on two ideas of interpretability .
But the big question is what exactly is interpretability .
Again, there's no interpretability required there.
So there's this really popular example in the interpretability of the wolf versus husky.
I gave a bunch of reasons why you might want interpretability .
So this is an example of a task where you wanted interpretability to be able to check whether the classifier was making
Another way of doing this is saying that interpretability has to be evaluated in the context-- in terms of a real application.
There was some task in service of which you needed the interpretability , and the HCI, the visualization, the human factors communities have definitely
So we had him and one of our early teams focus on this area of interpretability , which we think is good for making models safe and transparent.
When folks come to Anthropic, interpretability often a draw, and I tell them, the other places you didn't go, tell them why you came here,
So let's actually think about what reasons you might need for interpretability .
But we need to be able to come up with proposals of what we think interpretability should mean and what the reasonable definition should be.
So before I go into talking a little bit more about ways to measure interpretability , I'm just going to give you a little teaser or a taste of the sorts
But now I'm trying to connect back to the title of the talk around interpretability , which is the classifier was doing its job.
So this is just an example of a situation where you might want interpretability , but there are many.
And so this is a related way of saying it-- we have interpretability by fiat.
And it's important for us to think about the fact that interpretability might not be the one-- there might not be one thing that is
on large language models and the significance of those, whether it ranges from questions of interpretability or to questions of the environmental cost
And for these other things, accountability, interpretability , and even more as you go further down the list-- and people are working on these things, and they're important.
And then the final chapter of the book before we get to the catch all chapter that discusses everything from sort of interpretability to every AI
And today, she'll be speaking to us about interpretability towards more rigorous evaluation.
But more importantly, the thing that I want to chat about is how we can make interpretability more rigorous.
So just to get started-- you wouldn't be here if you weren't interested in interpretability .
But you see there's huge numbers of publications in the last several years under this guise of interpretability and machine learning.
So first, let's start out with this question of what is interpretability .
And I think that explanation is actually a much more human, tangible word than interpretability .
And I'm encouraging you all to think about, when you think of interpretability , is the question what's the quality of this explanation.
And before I go forward, I also want to distinguish the word "interpretability ," or the process of providing explanation, from a lot of other words that we also
So there's a pressing need here from a legal perspective to come up with operational definitions of interpretability .
So if you're living in that scenario, you don't really need interpretability .
So the focus of this talk is going to be around the definition and the evaluation of interpretability .