// HACKER NEWS — CYBERSECURITY
How to win a beer with high-dimensional statistics
My longtime labmate-turned-student/friend1 Dhruva Karkada recently wrote a sick paper on data statistics which deservedly went viral on Twitter, in part because it has one of the prettiest scientific figures I have ever seen:
Say you’ve got a bunch of words that live in a vocabulary $\mathcal{V}$. We’re here studying models $f$ that map $f: \mathcal{V} \rightarrow \mathbb{R}^d$: that is, they map every word to a $d$-dimensional vector. We’re letting $\{v_i\}_{i=1}^{12} = \{\texttt{January}, \texttt{February}, \ldots\}$ be the months of the year, taking the 12 associated embedding vectors $\mathbf{w}_i = f(v_i)$, and computing two things:
Reading the rows of this figure from top to bottom,2 they find that:
This is a big deal because it connects data statistics to representational geometry with a really simple mathematical theory.
After seeing this a bunch of times and staring at it for a while, I was feeling in the mood to poke a hole in this beautiful result, and so I bet Dhruva a beer that I could find a collection of other, seemingly-unrelated words that form a circle + circulant matrix in the same way. He (and most others I told) thought this was crazy, since the circle clearly comes from the special relationship between the words. We settled on the terms of the bet: I had to find ten random-seeming words whose word2vec embeddings, when plotted as the above, made a clear and compelling circle.
Why’d I think this was possible? Well, we have vocabulary of $25000$ words to choose from. That gives you $N = \binom{25000}{10} \approx 3 \times 10^{37}$ sets to select from. I figured that if you threw ten darts at a board that many times, you’d definitely make a circle at least once. Info-theoretically speaking, you have $\log_2 N \approx 124$ bits of information, and surely you can make a decent 10-point circle with that amount of resolving power. The question’s just how you find a set of ten good words in the haystack.
Note that the embedding size $d = 10000$ never entered into this. This all took an afternoon with a coding agent.
That’s circular. You can just find other sets of random-looking words that form circles!
This raises certain open questions, including “how can one man be so wrong?”, which I am not qualified to answer.
But seriously: clearly we can find spurious geometric patterns. Should this change our understanding of representation geometry? I’d note a few caveats first: