AI-Generated Genomes, Retinal Implants, and Palomar's Mystery Lights Explained
EP 15
·8:50

Why “precision” is hard; data-sharing vs. privacy trade-offs

Watch AI-Generated Genomes, Retinal Implants, and Palomar's Mystery Lights Explained

Transcript

This chapter, from the episode video's captions · 4,958 words

8:51pre-existing conditions are, we're going to charge you more because you're a higher risk and liability for us as the insurance provider. >> Exactly. Yeah. You don't want that. That's just blatant nonsense. Right. Yeah. >> The the other thing actually that it's tackling is um if you have a AI generated tool, whenever we make AI, you want to have a ground truth data set that tells you what is right or wrong, right? So that you can compare one AI model from another. Well, with with cancer data sets, it's really hard to do that kind of benchmarking because the the real patient samples have sequencing errors. you don't actually know what the tumor complexity is.

9:32There's no actual ground truth. So if we were to generate data sets that have a ground truth, then we can easily benchmark all of these other AI models. So when someone says, "Hey, I've got an AI model that I think does better. We've actually got a data set where we can test that." >> Yeah. >> Right. So So that's that's a separate issue from the privacy thing, but this thing is actually tackling both of them at the same time. And obviously there's a huge economic and human imperative, right? Like there's a global c cancer burden, 20 million new cases every year. Um 9.7 million deaths every year. There's a huge economic toll. Um drug development is long and expensive. Takes

10:13hundreds of millions of dollars. And if we have this kind of data set, we can train models that do it way better, right? We can we can limit all of that and democratize this approach. This is actually totally an organic anecdote, but literally I was in a work meeting earlier today and unfortunately one of my co-workers, you know, has a partner that um was just diagnosed with stage 4 >> and like you know I mean it is it it happens every day. Yeah. >> It it is present >> personally for millions of people. >> One two degrees of separation for so many people. you would know someone >> and and literally today I mean I I I I

10:55so this is clearly obviously >> uh a highly >> important and prevalent issue. >> Yes. Exactly. And so let's talk about like cancer itself. >> Okay. >> And why it's even like hard to do this. >> Okay. >> It's hard to like make these genomic data sets that are synthetic the way that we want. Okay. Cancer is a disease of the genome. It's caused by DNA errors. At the end of the day, that's what's happening. Okay, there's 3 billion DNA letters ATGC in our genome. And we're going to use this analogy of the genome being a library of cookbooks. Every single chromosome is a single

11:35cookbook that let's say one cookbook is for Indian food, one cookbook is for desserts and so on and so forth, right? So the genome is a library of these cookbooks. Each cookbook is a chromosome and then each recipe is a gene inside that cookbook. >> Okay. >> Okay. If we have errors in critical recipes, that's going to lead to cancer. >> Mhm. >> And there's really two types of big mutations that happen, these typos that happen in our recipe. Okay. There's germline mutations which is basically when a mutation is right when a letter gets substituted for another letter or some mistake happens in the genomic

12:18reading of the thing or I delete a bunch of letters so on and so forth right there there's two types of mutations that happen there's a germline mutation that's something that's inherited from your parents and that's present in every single one of your cells right because that's the you became you came from a zygote which is a single cell and so whatever mutation was in that zygote >> permeates >> that's all of you right then there's somatic mutations these are acquired during your lifetime it happens within a single cell and then the progeny of that single cell inherits that somatic mutation right cancer is actually primarily driven by accumulated somatic mutations line

12:59>> it's it's there are obviously cases where germline mutations create the cancer off the bat but usually It happens with age and there's these accumulated somatic mutations that cause your cells to do weird things, become tumorous, become carcinogenic and then create the cancer, right? And one key aspect is ankoan which is this generative AI it's trained on and it generates only somatic mutations. It doesn't get trained on the germline info. So it's it's already getting rid of that sort of privacy >> issue which is important to understand that there's two entry points

13:39>> one of which has less of a privacy concern >> by default. Um and that's where this uncle gan is focused on. >> Exactly. Exactly. Yeah. Now when it comes to the architects of cancers there's two types of cancers. There's a driver and a passenger. So the driver mutation these are critical alterations that give a growth advantage to some kind of cell. Okay. It can either be something like you turn on an ankco gene which is something that turns on the cancer part of the genome or it turns off a tumor repressor right a tumor suppressor. So then if if the cell wants to go become tumorous, this gene isn't being active telling it no, you don't

14:20want to do that. Not >> right. And then so these are the driver mutations. Then there's passenger mutations which are kind of like hitchhikers. They accumulate by chance. There's no growth advantage. It's kind of like just there, but it doesn't really affect the the cellular machinery in the way that it like turns it on and becomes cancerous. Right? tumors have very few driver mutations amidst thousands of passenger mutations. Okay. So, there's this distribution where there's a few that are targeted for these ankco genes or these tumor suppressor genes and then there's a bunch of noise. >> Yes. >> Right. And any kind of generative model

15:02like ankan needs to model that discrepancy where there's a few of the stuff that matters and there's a lot of noise, right? And Uncle Gan does exactly that, right? It replicates this tumor architecture by making a bunch of random passengers and not that many driver mutations. >> Okay. And the final thing I want to touch on is these mutations have fingerprints when it comes to cancer. Okay. There's characteristic patterns of mutations that give certain processes a leg up that becomes cancerous. For example, like UV light, right? UV light causes um a C to

15:43T mutation. Basically, you get this dmerization where like one leg of your DNA, you know, your DNA is a twisted ladder, but one part of your ladder is going to like staple itself. >> Like one leg of your ladder, it's got two legs. One leg of your ladder is going to staple itself onto itself and become a dimer. And then that's going to cause melanoma. >> Yep. >> For example, right? Tobacco smoke is linked to another singlebase substitution that'll cause lung cancer. >> So ankan which is this thing that people have made. >> It accurately reproduces that tissue specific signature right where it's like

16:23the lung cancer is going to do this and the skin cancer is going to do that so on and so forth. It's actually learned how to do all of that stuff. >> That's fascinating. >> Right. And then the the other thing, we're we're not done with all of the stuff that can cause cancer. There's also something called large scale vandalism, right? There's copy number alterations, CNAs. This means you've taken a whole chunk of your chromosome and you've duplicated it. Oh, okay. Yeah. Okay. So, now there's a duplication or even a deletion sometimes of large chromosomeal segments. So, you've got large parts like large chapters of your cookbook that are now been copied or completely deleted.

17:03You've also got structural variance where you take a big chunk of your DNA and you flip it and now it's reversed. Now, it's just all sorts of chaos, right? >> Like there's so many different things that can happen >> that with cancer genomes, right? >> That are the driver and the cause of it, >> right? And the real and the real sort of challenge is how do we make a single architecture, a single framework that captures all of these different myriad effects >> right >> into a single pipeline >> that then can be used to train other AI to train other tools, right? And that's what this paper is doing. So

17:44>> they use something called um generative AI. We've seen this a lot, right? like uh a cat on a thing. This one, this particular architecture is called a GAN. It's a generative adversarial network. Okay, there's two competing neural networks. This is a very interesting way to actually train AI and to architecture it up. I think it's very cool. It's been used actually for image generation as well. >> Yep. >> But this time what they're doing is they're using it to generate genomic data. It's interesting because I've I've been watching this space for a while and I've I've seen GANs being used in the the image generation context

18:25now mapping it to this problem set is an interesting use of the architecture of >> general uh generative adversarial networks. That's interesting >> and and the basic idea behind general adversar generative adversarial networks is you've got two networks, okay? You've got one that's the forger and one that's the discriminator. Okay, the forger or the generator. This is the guy who's going to create forgeries from random noise and try to attempt realism. And then there's the discriminator, which is like the judge who's trying to tell whether the thing that was generated is a real thing or a fake thing. So both of these guys have

19:07the training set which is all the stuff that is actually real. And the the generator is creating fake versions of that. >> Yep. >> And over many many iterations both of these guys are going to get really good at their jobs. The generator is going to get really good at creating fake data and the discriminator is going to get really good at telling whether it's fake or not. And it's this competing adversarial. That's where we get adversarial from. This competing effect that actually causes the generator to get really really good at its job. So there's this substrate there's a substrate of truth that these two processes have access to. One's job is

19:49to create synthetic data based on that substrate. Yeah. >> The other is to say whether >> you're fake or not. And and because there's this sort of iterative back and forth between these two, you get to a place >> where you know the generator has now generated something that the discriminator views as being >> real because it's gone through that back and forth process. >> Exactly. And now it's like really good at making real looking >> real looking data data which goes back to our privacy issue and some of the other things we just >> Exactly. Yeah. Yeah. And so this cycle repeat repeats until the generator produces stuff that's indistinguishable

20:30from the real data. And this can be used for tabular data which is a lot of this genomic data, right? It's not like images and things like that. It's like tables of like how much of this is there, how much of this is there. Enko GAN uses something called Cabab GAN plus which is just a specific architecture that's used for tabular data. >> Got it. >> And it handles mixed data types, imbalanced distributions. So like the distribution is not like completely uniform. You have some stuff happening all the time. Some mutations happen all the time. Other mutations are rare. So on and so forth. So this takes care of that. >> It handles the complexity of of this kind of data set. >> Exactly. Exactly. Now the other thing that they're using is something called

21:11variational autoenccoders. These are simpler than GANs. And what they do is they basically take whatever input you have, they try to squish it down into a tiny little latent space, right? And that latent space has fewer dimensions than the original data. >> And then from that latent space, we have an decoder that >> tries to recreate >> what was originally compressed. >> Okay. Yeah. So this is used a lot. This is used actually a lot in like image generation. When we have let's say a 256

21:54x 256 image that has 256* 256 dimensions, right? Because each pixel has a value. And you can imagine if if there's only three pixels, you can imagine putting that on an xyz plane, right? where every single pixel's value has to do like the first pixel is how far it is in the x direction. The second pixel is how far it is in the y direction and so on and so forth. Well, now you have 256 times 256 dimensions. Every single image is highly dimensional. But if I were to take a bunch of photos of just a bunch of faces, right? The essence of the face is

22:36actually not that highly dimensional. There's always a nose in the middle. There's always a mouth. There's always eyes, right? Maybe maybe one of the directions could be how dark is the face, right? For us, it would be like maxed out. >> Very dark. >> Yeah, very dark. Like there would be one dimension that would tell the variational autoenccoder this is a very dark face, right? Then there would be one dimension that would tell it how far up in the nose is the nose and how how wide is the mouth and things like that. But the fact that there is a mouth is not something that needs to be encoded in the latent space because it's always there, right? The stuff that's varying

23:18is the stuff that's going to be encoded in that compressed middle part, right? And then you have a decoder that takes that middle part that says okay how much is the how much is where is this and where is that and puts it puts it into the actual image that it reconstructs. And this is something that we're seeing in emnest for example the emnest data set is a bunch of handwritten digits. >> Yeah. The key part about variational autoenccoders is that it's a continuous latent space. Meaning that all of my real data points are somewhere in my latent space, but I can pick any point in between and that'll result in something real. >> Okay? It's continuous. It's not discreet like

23:58>> this and this and there's nothing in between, right? There is actually meaning in between. And so you can see in that latent space of the emnest, it's going from the thing that looks like a nine, >> which is over here. >> Yeah. >> To a three to a six to a zero, right? And it's continuously deforming. >> Yeah. >> And all they're doing is reconstructing what happens if I move on that latent space like as a walk and I see what is the decoder reconstruct. >> Reconstruct. >> Does that sort of make sense? >> No. No. 100%. It's like it's like it's taking all of the information of the stuff that it's trained on and saying what is the essence >> of the thing.

24:39>> I'm going to put that into a landscape that I can traverse. >> You can tra Yeah, exactly. >> And then and then as I move through that landscape, it's going to create meaning with my decoder >> because there's there's different points on that landscape, for example, in the numbers that we find meaning in, which is like the nine, the three, the zero. But the there is this there is this this transformation continuous in between there where you can say oh this is like >> a 93ish >> thingy. >> Yeah. On my way from 9 to three. >> Exactly. Yeah. Exactly. >> Exactly. Yeah. So, so, so they're using both. And, and the key to using both is

25:20that you've got a GAN that's used for complex interdependent features like mutation counts, driver co-occurrence that that's there due to highfidelity matching. And then you have this variational autoenccoder that's there for continuous variables like genomic position. The position is a continuous thing, right? But the mutation count, that's a discrete thing. >> That's a discreet thing, >> right? So you're you're combining both into this hybrid approach to create this like realism in the data set. >> That's f that's actually really really interesting. >> It's they're not they're not limiting themselves to like one type of architecture right? >> They're really harnessing the true

26:02power. >> It's like there's this blended architecture which is using the the value of both of these like modalities >> and taking the best of both. >> Yeah. and and sort of using them in the context that they that they're good at >> in in in parallel or together. >> Yeah. Exactly. Exactly. Because the genome is is both continuous and discrete. >> Discrete, right? And so now we're >> we're getting like two soldiers here that that are doing the work, right? In order to train enco data from can panc cancer analysis of whole genomes PC AWG it's a data set that has 2,658

26:44whole cancer genomes like entire genomes. Okay. All 3 billion base pairs. >> Massive. >> Massive. Yeah. And this this data set is treated like sensitive information because it should be right because this has actual markers that can identify patients and things like that. But using that they've now created something that will generate data that doesn't have those limitations. >> They've created a derivative product from source data that is personal identifiable information that now has that degree of separation uh to to not run a foul of the practice. >> Exactly. Yeah. Yeah. Exactly. And from

27:24what I counted in the paper, there's five GAN models and one TVA model. So there's five GAN generative adversarial network models and then one um variational autoenccoder model. All of these guys are working together to create this giant genome right multiple giant genomes and obviously the question is to ask is like you know okay how real are you making this right the photo that you're making could it fool me? This becomes the with photos it's easy because we sort of have the eye test, right? This is the whole uncanny valley concept which is like, okay, I can look at something and be like, "Oh, the fingers >> Yeah. >> aren't right." >> I mean, now, dude, I'll be honest, dude.

28:06There's like videos on Instagram that I get fooled. I I do a double take and I'm like, "Oh, this is this is definitely >> there's a there's a Neil Degrass Neil Degrass Tyson just did a video of himself talking about how Earth is flat and he he held up an iPad to the camera. You didn't know like the video starts as just Neil Degrass Tyson talking about how the Earth is flat and then like he just moves the iPad and he's like that wasn't me." >> And it's >> it's like and it's the same it's like him in the same space. >> Oh my god. in the same outfit, in the same voice, >> and it's effectively indistinguishable. >> Yeah, dude. It's It's getting It's getting really good, right? So, at least we're harnessing it for like something

28:46>> something useful. >> Yeah. Um, so one of the things they looked at to see, okay, is the stuff that's being output by >> this neural network, is it something that can be real? Right? They looked at mutational density and type, they found a high similarity between the synthetic and the real data. Mhm. >> They looked at mutational signatures. They found that it's like pretty nicely correlated. Um, the real thing that got me was the genomic distribution. So, you can look at the entire genome of the human from, you know, chromosome 1 all the way to chromosome 23 with the sex chromosomes and you can follow where mutations are happening. Okay? And in

29:27cancer, mutations do not happen >> uniformly. >> Okay? There's going to be parts of the genome where it happens all the time >> and there's parts of the genomes where there's very few mutations. For example, if you're closer to the center of the chromosome, you're not going to get that many mutations. But if you're >> closer to the edge, you're going to get a lot more mutations just because of the physics of like, you know, you got a little bit more freedom towards the edge than towards the center, right? And so they they actually showed that mutations in their data set would follow nonuniform distributions. So on the on the top they have the distribution of mutations in a real sample.

30:08>> Yes. >> On the bottom you've got your generated. It looks pretty much the same. >> The eye test, you know, >> the eye test. Yeah. Exactly. And what they've done is like the axis on the x-axis is all the genes. Yeah, >> like all 23 chromosomes worth of genomes and they've binned it in like I think 1 kilobase >> bins and they've counted how many mutations are happening. Right. >> That's incredible. >> This is actually it's very it's very similar to what we would actually see. >> What we what we've been basically able to do is now create a a system that on the fly it's not fixed. Yeah. It's it's a it's a it's a generative model.

30:50>> Yeah. So on the fly any team can now if you take this and say we want this kind of dynamics >> and they can generate now these cancer data sets that mimic real cancer genomes to a degree of accuracy or or comparable >> Yeah. or uh high fidelity that is usable Yeah. in in practical work in this space. >> Exactly. And that's that key that you just touched on, usable, right? What does it mean to be usable? Well, they actually tried to test it with something called deep tumor. Deep tumor is a

31:32>> again another AI that's been developed to identify cancers. Okay? They could fool deep tumor really well. But here's the key. Here's the key. >> Deep tumor is really bad at rare cancer subtypes. >> Okay? >> Okay. like lymph mlll there's only 35 samples in the training set okay and if you have low mutations you're going to struggle with like creating an AI that can identify that well they substituted generated samples into the training set and deep tumor did better once deep tumor trained on those the F1 scores improved >> that's fascinating so the the synthetic

32:13data from this GAN this ankco GAN when fed into an existing detection model >> that's like kind of established. People use it all the time. >> It's it's it's credible all these things. it has now improved its ability for the these edge cases >> these like edge niche cases because now I can I can generate genomes right for that specific thing that's not really prevalent >> in a lot of the real world data sets >> because the idea is you can take the deep tumor and then use it on like real world patient and like and and know that they do or do not have itactly and then if it comes out with a result that is >> exactly really powerful >> it's it's really it's really augmenting the data sets that we already have

32:55>> right >> with these synthetic data sets which by the way this is something that people do all the time in AI research right you do data augmentation where you take whatever data you have you like apply transforms to it when it comes to image data you'll like squish it you'll like turn it from black and white to red white and blue or vice versa you'll try to make the model more robust >> to the data >> quality >> but here they're actually able to create new data that can target these really rare cancer subtypes and establish a new paradigm for medical AI. Right. We can generate more data instead of we just need to collect more data.

33:36>> Right. Right. Right. Which which you know the there is a bottleneck on the collecting data piece. >> Exactly. And so you can much more quickly scale uh if you can uh create usable synthetic data sets that actually map to the real world uh use cases in a way that is demonstrabably true. >> Yes. >> Which is kind of the point part of the point of this. >> Yeah. Exactly. Exactly. And there's no one toone correspondence with the real patients right? >> You don't have to worry about HIPPA. You you do you solve the privacy issue while accelerating the ability to both create the synthetic data sets and then have those applied to existing tool

34:17sets or create new tool sets around all this process. And the team already they've got 800 synthetic genomes that are openly available. No ethics approval needed, right? So anyone can go and try to train >> a model based on that. if they've got a new architecture in their mind that really is specific to genomic data, they can test it on this. It's a new gold standard for benchmarking, right? Because now when I generate some kind of data set, I have a ground truth on what type of cancer it is, what are the mutations that should be flagged, so on and so forth. So, it's it's a catalyst for >> medical oncology. I think it's I think it's very cool. It's a huge new tool in

34:59the toolbox for oncology studies. >> Yeah. I mean, there's still limitations, right? Like the fact that we're doing this factorization where we first create the mutations and then we have another model that goes in and distributes it across the genome. It's not fully capturing this like complex interplay between features. Like for example, if you have a copy number alter alteration where you have these giant chunks of chromosome that get duplicated or deleted, that's going to influence local mutation rate. This is not something that captures that, right? It can't simulate complex events like um chromoth cis and like tumor subclones, but it's

35:41still a huge huge step forward, right? We we have to it's it's one step at a time. Yes, >> this has obviously been an issue that is not only affecting so many people, but it's been happening for so long and any any new step towards >> our ability to solve this more robustly is like hugely important. >> Yeah. Yeah. Yeah. Because like cancer is just like it's just such an insanely difficult problem, right? It's not a single cause, >> right? >> Um there's no single like panacea, >> right? Solution. >> Yeah. It's just yeah so so any any incremental and this this I think is actually a huge step because I think this opens up

36:23>> other people to do this kind of stuff >> to create large data sets to then have >> really AI going in full-fledged >> right right >> to try and help out >> right right in a in a way that again keeps in mind some of the aspects of data privacy that that are you know are very important when you talk about medical studies >> for all the reasons we talked talked about at the top of the story. >> Uh this is actually this is a really really again we said we're going to make sure we give Toronto >> Yeah. University of Toronto. I mean their AI is uh >> this is a big deal. >> Their their AI is always like on point. You know their AI research is is

37:03ridiculous. So >> this is a really big deal. >> This is a really big deal. I mean you Yeah, you guys lost the World Series, >> but >> we'll see you next year hopefully. You know, that would be great. >> You know, round three. >> Uh round three. you've given us an incredible new tool in the box ability to to do cancer detection um and expand the amount of people that can start really working at this issue in in earnest. Uh that's our story number one. >> Yep. >> Number two, the super story. The Avengers came together uh for for this wireless retinal chip >> that's helping solve for blindness. Yeah. >> This is in the New England Journal of Medicine. And again, we have Science

37:44Corporation Stanford UCSF University of Pittsburgh, and University of Bun that all collaborated on this story. You were really excited to talk to me about this story.

From AI-Generated Genomes, Retinal Implants, and Palomar's Mystery Lights Explained

AI cancer genomes, bionic vision, and Palomar transient skepticism.