1,646 words · auto-generated from the episode video
1:29:01three four five copies six copies right that'll be one axis the other axis will be how How many of chromosome 2 do I have? How many of chromosome 3 do I have? The fourth axis. And so you have 22 dimensions each chromosome. And the point where you are is how many copies of each chromosome you have. Right? So if you only had three chromosomes, let's say you were some organism and you had three copies of one, two copies of of number two, and four copies of number three. Then my point on this axis would be three this way, two this way, four up. >> That's where I am. That's where this particular cell is. Yes. >> Right. So now we can create an actual quantitative >> landscape
1:29:41>> and we can measure at each point what is the slope of this landscape because we can track that cell at that particular point and ask which way did it go. >> Right? And we can reconstruct >> the fitness landscape that way. >> That's what they're doing very quantitatively. They're like reconstructing this in a very real sense. It's no longer a metaphor. the the initially this concept was strictly meant as a mental model. >> Yeah. >> As a way to gro or to kind of imagine >> how this happens. Yeah. What we're saying is now scientists are actually taking real world data and creating a literal fitness landscape
1:30:21>> particularly in oncological studies. >> Yes. Yes. And that's why it's called adaptive local fitness landscapes or annuploid karotypes. They're literally making that fitness landscape. And then from that construction now we can tell which way is it going to go. Is it going to go down the valley? Is it going to go up this way to this peak? Or is there another peak that's nearby? These are the questions that can they can answer. This is not a trivial problem to do, right? Because and it really comes from technology. So before we used to have the standard single cell sequencing where what we do something called um MDA. What we would do is effectively have these polymerases, these DNA replicators, they would exponentially amplify certain
1:31:04parts of DNA. So if we wanted to sequence the DNA of a single gene, single cell, there's not a lot of DNA, right? So in order to sequence, you need to amplify the amount of DNA so that then we can sequence it. when you amplify things, you could get exponential speed up on certain chromosomes and then you know the the protein that's doing the DNA replication maybe got to another protein got to another chromosome a little bit later. So there's not a lot of five because it started on five later, but on one it started right away. And so one, there's like a thousand copies of one, but there's two copies of two. That's not going to help for us, especially for annuployy, right? If we've got like a thousand copies of gene one, which is on
1:31:45chromosome 1. Well, I don't know if that's because there's actually a thousand copies >> or if he just started >> or if or if the guy just like started there first, right? Instead, these guys used the DLP plus solution. there was already um a paper that was out that had a direct library preparation of single cells. What you do in with DLP plus is you don't care about the sequences of the chromosomes. You only care about the number, >> right? Cuz that's what you want for this particular study. >> Yes. >> And so what you do here is you dilute your tissue. So you've got the tissue with a bunch of cells. You dilute it so that each drop that comes out of
1:32:26your pipette has a single cell in it. >> You put each drop into a micrfluidics like um array. It's like a chip and you can now weigh each droplet to figure out how much of this chromosome is there, how much of this because each combination is going to get you a unique weight, >> right? If I have like four 1 kilogram weights and two 5 kilogram weights, so on and so forth, I can like figure out and back construct how much of each is there right? >> Mhm. >> And that's what they used. That's the data set that they used. it was already there and they're actually using this data set to now create this 22dimensional space >> because what they're doing there is it
1:33:07removes the distinguishing this the distinguishing the problem of how to distinguish >> uh because everything is sort of universal like it's it's it's like a clean base. >> Yeah, it's a clean blaze. It's like flat. There's uniform coverage. It's not like oh chromosome one got lucky. It's kind of like, you know, when you want to just figure out how many pages there are in a book, >> right? The MDA analogy would be you you have a noisy photocopier that sometimes it'll copy a thousand copies of page one, two copies of page two, 5,000 copies of page three. Fine, you'll get to read the book. >> Yeah. Yeah. But it's >> But I don't know if the original book
1:33:48had a thousand. I don't know. You know, maybe the author was crazy, right? >> But >> with DLP Plus, all you're doing is weighing the book. Mhm. >> And from that inferring how many pages it has because the only thing you care about is the pages. You don't want to read the book. >> Right. >> Right. For this particular use case. >> Understood. Yep. >> You only care about the number of chromosomes that you've got. >> Yes. >> And so from that from that discrete data point now we've constructed this 22dimensional space. This is the original paper from 2021 in nature that actually put out that data library. Yes. And this is the data library that they're using. So that again this goes back to what I always like to talk about is this science has this compounding effect and you know new discoveries can
1:34:29enable others. So basically what we what we're saying is there is a >> uh a library that this DLP plus library that the alpha team was able to use to deal with this uh problem of number of different uh express the volume of expression that was not what they were measuring for. They just needed a clean base in order to be able to look at um this specific issue as it relates to how cancer like how cancer is making these selections about number of chromosome types as a means to traverse this fitness landscape. >> Exactly. And now we can finally get into
1:35:11reconstructing the fitness fitness landscape. Right. Because now we can follow single cells. >> Right. >> And we can ask what are they doing? >> Right. >> Right. Right. How are they going about in this fitness landscape? You can ask for a frequency change based on you know individual fitness and you can look at what is the slope of a particular cell at a certain point right and from that infer okay this is actually a really steep slope because the cell is >> is reproducing way faster that means that it's going up in fitness right and so on and so forth. So it's it's it's actually very very cool to think about. Um the other really cool technique that they used was something called crigging. Have you heard of crigging? >> No. >> Yeah, I I hadn't heard of it either.
1:35:52It's it's called gausian process regeneration. Effectively, you've got a bunch of points and you ask, okay, how do I smoothly fit a bunch of gausian processes such that you can interpolate between these different points. Okay, so it's not just like standard interpolation where you just like add lines between points. You want there to be some kind of flow. Um this was originally found because um they were looking for how to infer where the gold was in a South African mine >> of course. >> Okay. >> Like they were they they drilled bore holes everywhere >> and they they found the amount of gold in each bore hole and then they figured okay a gold line is going to be kind of a gausian sort of smooth thing. So if I
1:36:34have data here how do I infer where the gold is in between? Right? And that's what it was originally used for. But now people are using it on all sorts of stuff. Um it's used in finance a lot too. But in this case for um cancer research which I thought was very cool. >> That is interesting. >> Um so now we got to prove that this map is real, right? We've constructed this map, >> right? We understand how we were able to generate this fitness landscape. >> Um and now Okay, great. That's cool. Nice diagram. >> Yeah, nice diagram. Nice like 22dimensional space. >> Can you actually predict stuff? Yes. Okay, they did um incilico validation. So they they made agent based models where they simulate sort of cancer and
1:37:16they saw that the cancer was great. But again that's simulation going to get you a paper in nature communications. Okay. You've got to have some experimental data to back it up. So what they did was something called sister passages. What you do is you've got a a cancer tissue. Yeah. >> You split that up into two. You sequence one of them and you train your model or you you parameterize your model that way