The 2017 transformer paper was originally written as a machine translation architecture, but it turned out to be a general breakthrough in recognizing long-range correlations in text, letting a model track what a given word relates to across everything that came before it. This context-modeling ability became the basis for the large language models that followed, including ChatGPT, Claude, Gemini, Llama, and Deepseek.
136 words · auto-generated from the episode video
2:09:38this is an architectural breakthrough. The paper itself is just about translating language. Okay. And they used it to translate language. Pretty soon it turned out this is something that can um recognize and learn long-term correlations in text. Meaning, what does this word have to do with all of the words that came before it? Um, and recognize context and really try to distill the understanding of what the text is. This is the breakthrough that catalyzed the modern artificial intelligence revolution. It enabled large language models. Now we have Chat GPT, Claude, Gemini, Llama, Deepseek. All of that
2:10:20comes from this 2017 paper. >> No European models, just everyone else apparently. Well, no, they have a what's it called? They have a la leat. No, mist. >> No, >> we've never >> Moving on.