The individual dataset, which is used to compute the precision, and the ensemble dataset, which is used to build the ensemble classifier through precision weighted voting, which I will explain in the next slide.So we end up with 420 individual Random Forest classifiers , and we built them as follows.
They are illuminated by the spotlights. the classifier you built based on the earlier data, may not be relevant for future data.
were building a category of, you're building a recognition system for cats, a cat recognition system, you would develop a classifier that could learn the features that cats have in common that distinguish them from other animals, like dogs and birds and fish and so on. And this CAT-egory... get it? My daughter made
And then in the child's speech, I definitely think people are waking up to the fact that it's a problem today. A classifier --? --for kids.
So one option is to say, well, interpretability, it's like porn. a classifier that is.
And then in the child's speech, I definitely think people are waking up to the fact that it's a problem today. And for the classifier , how do you classify?
And so without the information of why it's making the decision, we can't tell for those three things. But if the classifier can tell you why it's making the decision and we can find multiple classifiers who make the same decision for different reasons, then maybe we can try to give those to someone and have them decide which is the best one, which is the one that they like.
And so you could build a classifier that says if there's a tree in the picture, it's probably a wolf. And that's a classifier that, on your dataset, will do really well. Because your dataset happened to have this bias in it that you were not aware of.
So one option is to say, well, interpretability, it's like porn. If I have a classifier that predicts malignant tumor every time the sky is blue, I can check that and see how good
Sorry, not 95. 99.5% are thrown out, and the half a percent are randomly sampled and kept as the safe state. So the classifier itself is a Random Forest classifier . This is a well-known machine learning technique based on decision trees.
The dark blue is the 0.25. How do the classifiers perform in practice?
And then in the child's speech, I definitely think people are waking up to the fact that it's a problem today. Is it more like a classifier for a kid's voice, or is it speech cognition including kids?
And so there's two different ways of doing the classification that are going to lead you to exactly the same classification accuracy. So in the first round of the classifier -- this is accuracy-- in the first round of the classifier , we're doing-- again, this is a simple classification problem, so you get 100% performance, and it's mostly picking out these corners.
The second round, we say you're not allowed to use the information that was highly weighted in the first round. Find me a completely different classifier -- or as best as you can. You have a parameter that you can tune.
The huskies are usually in indoor scenes. And so you could build a classifier that says if there's a tree in the picture, it's probably a wolf. And that's a classifier that, on your dataset, will do really well.
and safe, which is the negative class. So in terms of the classifier language, the positive class corresponds to the failures and the negative class corresponds to safe or nonfailure.
So these are the-- we have 15 benchmarks to do the testing. So going back to the ensemble classifier , we had 420 classifiers , and we merge all of them into an ensemble classifier . So we have ensembles of ensembles.
So the weighted value now can be seen as a likelihood of failure. So it's not a binary classifier . It's not one zero, but it's a real number.
Closer to 0 is less likely to fail. So this is the ensemble classifier . So now the prediction-- yeah.
So these are the results. So the results for a binary classifier is typical to display these ROC curves that I'm sure you're familiar with. So here I'm plotting the true positive rate in the vertical axis as a function of the false positive rate.
But obviously life is more complicated, so we get these curves. So the solid line is the ensemble classifier . And I'm showing you this for two different benchmarks, the left one being the best case.
So let's look at the best case. So the solid line is the ensemble classifier , it and the individual colored dots are the individual pieces of the ensemble. And the colors correspond to different ratios among the false fail and safe events.
And he's published a big catalog. classification based on the 65 classifiers , and it is then released to the astronomical community." So this is the sort of standard game.
So when you do all of the math you end up with 420. So we end up with 420 RF classifiers , and we train each of them on 10 consecutive days, and then we test on the next disjoint day. So remember the Google data trace has 29 days of data.
Because there's no screen in sight. We're training our system on classifiers for patterns, primarily because most of the money's in brands.
But at least you can skim through the factors that are being used and look at the prediction performances, and you can decide. You can say I really just don't like that classifier . I want a different one.
And we combined them by doing a weighted sum based on the precision. So we take-- if you take the Ith random forest classifier and take its output, which is shown as sigma IJ, which is a binary value indicating it's either in the fail or the safe state.
So now the prediction of failure now becomes a classification problem. And by applying the model you just evaluate the ensemble classifier for a particular data point. And then you interpret the result as the likelihood of a failure.
So we ran this study. Also, if the Google classifier that's on the phone indicated that you're currently walking, we didn't want the phone to ping-- audibly ping-- light up,
I've just taken 10 days off. Yeah, the World Archery Classifiers , yes.
So we'll use something like a kernel density estimator. But the basic idea here is that with relatively simple classifiers that have no information about the labels, we can actually pick out these abnormal events. And finally, instead of simply returning all of these points that we think are anomalous or unusual, and asking users, hey, figure this out for us--
I'll go very fast here, since I assume most of you are familiar with these techniques. So we generate a large set of Random Forest classifiers , and we train them, and then we compute their precision of each one, based on a dataset which is different from the dataset we used to do the evaluation.
classifier through precision weighted voting, which I will explain in the next slide.So we end up with 420 individual Random Forest classifiers , and we built them as follows. We varied the number of decision trees between two and 15, so there are 14 of them.
allowed to use the gradients that were large in the next round. So round 1, you create your own classifier . Round 2, either someone provides some annotations and says, actually, these were bad, don't use this information when making this decision, or you just
And one thing that we're working on now is, can we at least limit the number of things that are sensitive to so you can at least look at them-- so not only give us multiple qualitatively different classifiers , but can you produce a classifier that, for any test point, is going to give you a succinct explanation. And that's something that we're actively working on right now and is relatively easy to build into the loss function.
and father genomes, if you will. And they just used this very simple Bayesian learner called the naive Bayes classifier .
and father genomes, if you will. OK? And so what the naive Bayes classifier does is it incorporates that evidence.
So now the prediction-- yeah. So the weighting is done only on the outputs of the individual classifiers as I have defined the states. So you're right that some of the outputs may be missing, because we simply did not consider-- those are excluded.
So this is work that was just accepted to , which we were calling "Right for the Right Reasons." And here's the main idea. So let's say we're interested in creating some sort of-- we have a differentiable classifier , such as neural network. And so we have some set of outputs, and we have some inputs.
I'm just trying to understand. When you say you do a first pass, you find some classifier that seems to work, and you say, oh, I don't like that classifier , it's right for the wrong reason-- I'm just curious, does that mean it's not really right, like it's not predictive?
I gave a bunch of reasons why you might want interpretability. So this is an example of a task where you wanted interpretability to be able to check whether the classifier was making its decisions for the right reasons and be able to reject it if it was making it for the wrong reasons.
So if we were to keep all of the safe states, then the two classes would be extremely unbalanced, and most of the machine learning techniques that we use do not perform well when you do a classifier with such different ratios of the two states. So we do a half a percent subsampling, meaning that we throw away 95% of the safe states, and we pick a random half a percent sample.
is a binary value indicating it's either in the fail or the safe state. So for the data point J. And we compute the precision of the Ith classifier , that's indicated as PI.
And then you interpret the result as the likelihood of a failure. OK? Then we can now convert this back to a binary classifier by setting a threshold. So you can say, if the likelihood is more than 80%, I'll consider the output to be a failure.
So we ran this study. So for example, in Heart Steps, if the Google classifier indicated that you might be driving the car, we would not have the phone light up,
And then we have some set of weights that it's parameterized by. And then if we have that, the question could be are there multiple qualitatively different classifiers . And why might we be interested in that?
And what you see, again, is that there's many classifiers that have reasonable performance. And you see that the test performance varies quite a bit depending on the different choices of classifiers . Because sometimes they end up picking up weird confounds in the data.
that's indicated as PI. And then the weighted sum of the output with the precision summed over all of the 420 classifiers is labeled SJ. So the SJ value is no longer a binary value.
The dark blue is the 0.25. But anyway, the take away is that the ensemble typically performs better than the individual classifiers , which
The dark blue is the 0.25. is good. So it means that by putting individual classifiers together under a single ensemble we do better than any one of them