has to learn to construct categories on the fly. Perceiving emotions is not a clustering problem - it's a category construction problem. And it's a category construction problem whether you'remeasuring facial movements, or bodily movements, or the acoustics of someone's voice, or whether you're measuring the
in cancer biology. For her post-doctoral work, also at Berkeley, she developed bioconductor open-source software packages for clustering and multiple hypothesis testing. Katie was also part of the chimpanzee sequencing analysis consortium that published a sequence of the chimp genome, and she used the sequence to identify
And the largest node, here we rank by the betweeness centrality metric is over here, Politico, which is the political blog in Washington and it's followed both by Democrats and Republicans. The clustering algorithm painted with the blue Democrats but it's as you can see a good bridging organization. These others are highly conservative, the high betweeness centrality groups otherwise beyond that are conservative groups. And we've repeated this for the Pew
So play this -- take this technology and play it forward with much greater capability. the clustering length of the dark matter. Different forms of dark matter will cluster on different scales. And by, by having dark matter cluster on a
Like, the Fitbit UI is cleaner and prettier, no question. But Whoop is clustering way more information in here. The live graph of my heart rate—some of that stuff you might find useful.
I'm using an athletic endeavor here. And the second is clustering because I'm suggesting the right tail is coming in. That means the performance is getting clustered of people.
Ben Shneiderman: If you look within each box all the nodes are the same color so we apply a clustering algorithm to it and we have several embedded in NodeXL, so we'll do the clustering for you and then you've got one, two, three, four, six, seven, eight clusters here and then we'll draw the boxes to include where the size the biggest one's in the upper left and it's the most number of nodes in selection from Twitter. In this case, this one's a little messy because it shows the edges go across the boxes as well.
people who I often photograph at CHI Conferences, Terry Winograd may be familiar to some of you and a variety of other people. But the clustering structure and algorithms are very powerful to tease out structure in complex networks. This was done by a student just this semester as part of a homework assignment, but it was so nice I wanted to include it here. He took the 600 tech reports from the HCL Website,
who are also friends with each other. What they found is that clustering remained high for much longer. - The world immediately gets as small as a random graph, but it stays as clustered as if it were still regular.
- The world immediately gets as small as a random graph, but it stays as clustered as if it were still regular. So you could simultaneously have the clustering that we know is real and the small world that we know is real. - Now, in Watts and Strogatz's model, they looked at 1,000 nodes, but if you apply their model to the eight billion people on Earth,
And that's why I'm here today to talk to you, and it's also why I've changed what I'm going to talk about, OK? So what is dynamic quantum clustering ? It is a really different way of looking at data.
you find out instead of expressing 48 chains, they only express 24. And that's reflected in the clustering in the first three dimensions, OK?
But I want to show you some other ways in which I think it was done. But you can see that some clustering that's been marked out, the result of clustering are the overlaps here.
those who don't know, or the Republic Party, but all the Tweets that had GOP in them. And then we used the clustering algorithm and the layout to produce this dramatic and very stereotypic picture of controversy. You have one large cluster which we made red, which are all the Republicans who are well connected with each other and then the people who were saying negative things about GOP, and we painted them blue, they were the Democrats.
If you calculate the fraction of people you know who also know each other, that is a measure of the clustering in the network. So let's try a model with a high degree of clustering . Imagine all eight billion people on Earth are arranged into a circle, and say each person knows the 100 people closest to them.
And I like what you said about making sure that when you're asking for-- when you're looking at the data, to ask for dispersion around the number that we're clustering on.
This is work that was done by a former student of mine, Victor Gonzalez, with great effort, who categorized these smaller actions, clustering them into different tasks. And it took him a great deal of time.
And then I went back and I said, well, it's been a long time since I went to Wikipedia and looked up the definition of clustering . I said, clustering didn't mean looking for long, complex, extended structures in the data. Well, it turns out that's wrong.
Well, it turns out that's wrong. I was wrong. Clustering , since I looked last, got that as part of the thing. Not too many things do it, but that's what it is.
But I want to show you some other ways in which I think it was done. So you can see that there's actually a lot of clustering in this.
that are important, that could lead to interesting decisions. Here's another sort of great example of the clustering . On Flickr, we took all the tags that had the word "mouse" in them and then did the cluster and layout and you find a yellow group over here which is the computer mouse and the blue group which is the animal mouse and the red group which is Mickey Mouse and these are pictures that come from each
Not too many things do it, but that's what it is. But the typical ways of looking for unsupervised stuff will be hierarchical or k-means clustering . There are things called DBSCAN, HDBSCAN.
But I want to show you some other ways in which I think it was done. And particularly this group of squares over here suggests some sort of clustering .
being able to get access to resources are based on local connections. It goes further than that. I mean, one of the reasons that I think that we see a lot of clustering in industries is that In order to be a successful entrepreneur in an industry, you usually need to have some prior experience in an industry.
like London like Stockholm and the Bay Area, and so on, and those people demanding more specialized in-person type of services. And the result of this is that more people are clustering in these places. And that has been great for cities like London.
I lost one to Adam. He said, well, it's generally clustering . And I said, no, it's not.
They tend to cluster by visible characteristics like race. But if you actually interview, you'll find out that their unconsciously very good at clustering themselves by religion, by socioeconomic status. One of my favorite papers in this space looks at students walking into a computer lab and finds that students end up sitting
And then things really started taking off. And you see this, uh, clustering . In a short period of time of wireless digital devices, sequencing social networks, cloud and super- computing, all setting up the potential for this, uh, era of a great inflection in
over here shows a well-organized, coherent group that actually talks to each other. By 2011, we had polished our techniques still further and so we're able to apply clustering algorithms and then put the different groups into separate boxes whose size was determined by the number of nodes in each box, and so you get a set of researchers including our
: The boxes are interesting but they have no apparent connection to the obvious underlying structure of the visualization. How did you get the boxes? Ben Shneiderman: If you look within each box all the nodes are the same color so we apply a clustering algorithm to it and we have several embedded in NodeXL, so we'll do the clustering for you and then you've got one, two, three, four, six, seven, eight clusters here and then we'll draw the boxes to include where the size the biggest one's
Ben Shneiderman: Yeah -- Dan: radiating down below -- Ben Shneiderman: Tim Berners-Lee and I'm sorry, that's Tim Berners-Lee. This is, the way it's going out these are his close collaborators. The clustering algorithm actually put them in a separate cluster , not together with him which is that's the way clustering algorithms go. But you can see there's a preponderance he's connected to these people very well and Jim Hendler, Nigel Shadbolt, and others I can't see them right now. But we tend to do
Most of the people you know live close to you, and they also have a higher probability of knowing each other. If you calculate the fraction of people you know who also know each other, that is a measure of the clustering in the network. So let's try a model with a high degree of clustering .
Randomness has patterns in it, or at least the appearance of patterns, and it has clustering .
And you find jobs like Digital Marketing Specialist and also Beachbody Coaches. And I think this reflects a broader trend that we are actually seeing in recent years, with tech jobs clustering in skills cities, like London like Stockholm and the Bay Area, and so on, and those people demanding more specialized in-person
So I mean, why the hell did I retire to do this nonsense? And the answer is, because six years ago, I co-invented dynamic quantum clustering , or DQC. C And I created the company-- I mean, Stanford patented it right away.
And I said, no, it's not. And then I went back and I said, well, it's been a long time since I went to Wikipedia and looked up the definition of clustering . I said, clustering didn't mean looking for long, complex, extended structures in the data.
But number two-- what's more impressive is the structure of what I'm showing you, because I don't know of any other way to see a complex, interlocking structure of filaments like that-- the real structure of the data-- using any other clustering algorithm. That's not possible, in my knowledge.
He is author or co-author of more than 200 technical papers, and has worked with the Sloan Digital Sky Survey collaboration on galaxy clustering , shared
How do you put all those on-- you can't imagine putting those on the same coordinate system and doing, let's say, clustering and checking for anomalies.
And we learned later as we as this world woke up to us as we could see what was all around us in a different way, that you had a clustering effect throughout the country.