Skip to content
back to the archive page
#Security

Harnessing the Universal Geometry of Embeddings (Category Security)

Just a few days ago, I spoke at PHDays about reliability and security as fundamental parts of our systems, including new AI risks. I thought I had a reasonably clear understanding of the threat model. Yesterday, however, while reading Sergey Nikolenko's Sinekura channel (@sinecor), I came across a discussion of “Harnessing the Universal Geometry of Embeddings”. The paper introduces vec2vec, a method for transforming vectors from one embedding space into another. Sergey explains the implications in his post:

There seems to be nothing extraordinarily ingenious about the method itself, but the quality of the result is remarkable: Figure 3 shows how five almost mutually orthogonal vectors became vectors with dot products ranging from 0.8 to 0.95.

In practice, this has many implications, mostly unwelcome. Databases containing vector representations need to be protected as carefully as the original text—which, as far as I understand, nobody currently does. In one experiment, the researchers took embeddings of Enron's corporate emails, mapped them into a known model's space (Figure 4), and extracted sensitive information—names, dates, and amounts—from 80% of the documents.

Another interesting attack vector, this time targeting embedding databases rather than databases holding the original material. Perhaps I will get around to reading this paper soon :)

P.S. I subscribed to Sergey's channel after writing a review of the 2018 book Глубокое обучение. Погружение в мир нейросетей, which he coauthored. A reader mentioned that he had an interesting Telegram channel. I checked it out, liked it, and now read it from time to time. Thank you for the recommendation!

#Security #AI #Software #Engineering #Database #Data