Go back

Sequencing the genome of SARS-CoV2, featuring Grant Hall

31m 37s

Sequencing the genome of SARS-CoV2, featuring Grant Hall

In this podcast interview, Grant Hall from the COG-UK consortium explains the critical role of sequencing the SARS-CoV-2 genome. The initiative, building on the Arctic Network's experience from past outbreaks like Ebola, rapidly sequences virus samples from across the UK. This genomic data reveals mutations (about 1-2 per month) and, when integrated with hospital clinical data, helps track transmission patterns and viral spread geographically. The sequencing process itself involves converting viral RNA to more stable cDNA, amplifying it via multiplex PCR, and using nanopore technology to read the genome within 36-48 hours. Importantly, this work is done safely with non-infectious extracts. While sequencing cannot definitively trace every infection or fully prove the virus's natural origin, it provides powerful real-time evidence to inform public health decisions and counter misinformation. The decentralized, collaborative model between academic nodes and local hospitals enables fast, adaptable responses to the pandemic.

Transcription

4953 Words, 27854 Characters

English
Welcome to the Blue Side Podcast brought to you by Cambridge University Science magazine. I'm Ruby and I'm Shimane. Every two weeks we speak to local researchers, university staff and students and anyone who works in science to learn about their research and activities, hear about the work that they do and uncover what goes on behind the scenes. If you want to get in touch with a question and suggestion or just want to be featured on the podcast, just drop us a tweet. Our handle is @BlueSidePod and you can also email us at [email protected]. Hey everyone, welcome back to our series of episodes related to coronavirus. Today we're speaking to Grant Hall who is working as part of the COG UK initiative to sequence the genome of the SARS-CoV-2 novel coronavirus. So this is the virus that causes the COVID-19 disease. He'll tell us all about why sequencing is important and what we can learn from sequencing the virus's genome. So welcome Grant. Thank you so much for joining us today to talk about your work with COVID sequencing. Can you just tell us a little bit about yourself, the background and how you ended up on this project specifically? Yeah, absolutely. So I am an infill student right now in the Department of Pathology and I work in the Division of Biology as a member of the Good Fellow Lab. We traditionally do research on neuroviruses and this degree has been my introduction into the world of biology for the most part. I am a chemist by training, I completed my undergrad at the United States Military Academy working on synthetic drug development for leshmaniasis and I realized I was more intrigued on how the drugs interacted with the pathogen rather than the method for creating them. And so I went out pursuing a way to explore those interactions and found myself in the world of biology. I definitely didn't realize that I would be presented with the opportunity to work on genomic sequencing for SARS-CoV-2 when I join Professor Ian Good Fellow's lab. But now I find myself with the opportunity to be working and learning alongside a plethora of amazing scientists in the Cambridge node of the COVID-19 genomics UK consortium, Professor Ian Good Fellow as well as some of the postdocs in the lab, including Dr. Luke Meredith, have been involved in the past with an organization of universities known as the Arctic Network that have experienced doing real-time genomic sequencing and analysis during outbreak response. The Arctic Network ended up being the foundational body that would help upstart the COG-UK consortium. The prior expertise definitely shows as our node here in Cambridge was able to, from a standing start, sequence over 300 genomes in the first two weeks of setting the lab up. And after six weeks of work, we just submitted over 1,000 genomes now to the COG-UK consortium. Can you just a bit about the sequencing itself? What kind of information do you look for? What kind of information can we learn from sequencing something like a virus? So I think what's really great about when we look at sequencing data is it's not any sort of absolute data set. When we sequence viruses in mass, it's not to give us all the answers. We can't track where a virus has been through a bunch of patients just with the genetics data. But a genetic sequence can provide us with additional information that can help us paint a picture into an outbreak response. So for the case of a virus, all viruses, there's a lot of genetic diversity between them. Every time the virus replicates itself in a different host, there is the chance that you're going to incur a mutation. And so over time, viruses will slowly progress down a specific pathway, and you'll see a preponderance of a certain genome being different than when you say see it at the beginning of an outbreak. So for a virus like SARS-CoV-2, we see right now an average mutation rate of one to two changes per month. And so this slow genetic shift by mapping it in real time with the outbreak allows us to kind of map the progression and see what are the certain kinds of viral sequences that we're seeing in certain regions globally. And likewise, even more regionally within the UK is there a specific diversity that's starting to be found up north more than down south. And then when you have a great organization like the COVID Genomics UK Consortium that we are one single node apart of, and we start working closely with Public Health England, you can take this data like a side-impaired with clinical data or epidemiological data. And you start to be able to extrapolate and find more information about an outbreak in general. So you can start to track transmission routes and be better informed with the number of transmissions, potentially cases introduced into the UK and then how those propagated. Or you can start to notice if there are certain kind of phenotype of the disease, a certain way that patients present is it associated with a certain genetic sequence. And so at this point there's nothing flashier exciting, but by providing this information in real time rather than in retrospect, you can better inform Public Health Policy in a real time manner and hopefully be able to make fast decisions that will save lives. It's so amazing thinking about the fact that they've already calculated a mutation rate already for the virus so quickly. That's amazing. Yeah. Yeah. No, it truly is. The speed of the research is incredible. Yeah. And for like non-biological background people, what kind of how big is the genetic code of virus compared to like let's say like human, I don't know, genetic codes. So yeah, in comparison, the human genome, it is tiny. But even when we're looking at RNA viruses, coronaviruses in general are on the larger end of the spectrum. And this is just by nature, RNA is a lot more unstable in comparison to DNA. So SARS-CoV-2 has a genome that is just under 30,000 base pairs. So big in terms of RNA viruses, but small comparison to the human genome. Yeah. And so when you're talking about the accumulation of mutations or changes, what is the implications of this in terms of thinking about vaccine development or it's ability to transmit? So I think it comes back to what I was touching on a little bit in the sense that there are no absolute implications. These mutations could lead to a significant change in potentially the presentation of a specific protein or it could lead to a variation in the disease's ability to spread or potentially the phenotype of disease that's presented by the viral infection so you could see more or less asymptomatic cases. But especially with the virus that we don't understand, we don't necessarily know the implications of these mutations and they are random. There are probably plenty of skilled mathematicians that could start to predict and create models but that requires data. And so what's great about again this mass approach towards compiling genomic data that the COG UK consortium right now is attempting is that you're really providing a chance to develop a robust data set and sample of what the UK case load is looking like. And how does the interaction with other like COG UK like nodes work or is it that you all provide the data for your local regions but you also collaborate on like other things? So COG UK actually is a government funded consortium and initiative. So it was initially funded and set up with the help of DHSC UKRI and the welcome trust. It consists of four different public health, England institutions, the singer institute as well as 12 universities across the UK. And they are all working alongside local hospitals to sequence these samples and then process and upload the genomes while pairing it with epidemiological data collected by the hospitals and then clinical data as well. And so this then gets compiled into a massive data set and is handled by public health England and the government. But I think it's important to note that this wouldn't have been able to start as quickly as it did if it wasn't for the work of the Arctic Network which is this subset of universities that had been involved with some of the genomic sequencing during the Ebola outbreak and had developed these protocols, optimized protocols for outbreak response genomic sequencing. And so by nature of having these well established protocols in the context of viruses like Zika, Ebola and measles, they were able to then adapt them to the current outbreak and thus provide results in a very fast manner and have a really well established and uniform protocol that these labs and organizations could then work from. It sounds like extremely well organized considering how quickly this has all been taking off. And so if you're speaking about these nodes and although it's a sort of network approach and do you think it's important to sort of pair genomic sequencing with local hospitals to try and pair that, you know, is it important that that interactions happening so that we know what's happening locally? I absolutely think that's completely right. So if we're looking at issues that you see anytime you're trying to put on a massive collaboration or put together a massive initiative and you take academia, the government, different nonprofits and other organizations, anytime you have these bureaucratic organizations working together, there's going to be friction because everyone has a different operating procedure. And so anytime you can, I think, localize these interactions, there's an easier way to overcome those points of friction. You're able to work together more closely face to face and find solutions to the problems that you are encountering but at the same time provide reliable data. So the Stinger Institute obviously has this massive capacity for genomic sequencing and so they definitely are going to pull the main front when it comes to the COGUK's sequence database. But obviously they're going to work in a lot slower manner because they're receiving samples from across the UK. These small nodes are like ourselves here in Cambridge are able to provide data directly back to the hospital. And again, it comes into, it can help inform the epidemiological data and clinical data that they're getting to compare. And it's the genetic diversity that they're seeing in case loads. It explores the question, is there a way that genomic data can provide real time information about noseocomial transmission of the virus? And so while they're currently, we're still working out those kinds of protocols and people are trying to figure out whether or not this information is something that's needed to just aid the epidemiological data or if it can be more absolute, it does provide hospitals with the information and they can choose to then use it as they find fit to inform their own decision making processes. That sounds really helpful, yeah. And it's definitely a thing that like we, because we just started this series of answers about the COVID-19 situation in last, last time we were speaking to Professor Stephen Baker who was working on kind of like helping the outbreaks hospital with diagnostics. So obviously there they were running, you know, they're not a diagnostics lab, but they were kind of taking the load off of the low hospital for that, for the screening and it just worked. So it's definitely, it definitely seems like finding those local solutions is really constructive in the worthwhile. And one other thing that we talked to him about was also this idea of people staying informed and people finding reliable sources of information so that they weren't like misinformed during this situation and also like the appropriate like the fact that we have all this fake news going around and people that really know who like a reliable scientist is and so on. There's also a lot of like conspiracy theories I guess about like where the virus has come from, if it was made in the lab, if it was leaked from somewhere, I guess what can sequenced say and tell us about that? Can we use the genetic information of the virus to kind of trace it back the way at the same way instead of tracing it like forward to see where it's going? Can we look back and see where it's come from? So absolutely, when we actually look at genomic data analysis, we often actually look at it in a retrospective manner. So we can see the flow of the virus over time, how did it mutate and eventually try to work our way back to potentially that patient zero, that initial jump where a virus moved from a reservoir of some sort into a human population and began to start transmitting. There are definitely limitations from that. So we can gather as much information as we can and make definitely really well informed hypothesis on where what kind of virus was the initial source, what kind of host did the virus eventually make that jump from. But without that exact sample, there's always the possibility that there's something else that we aren't aware of. There's some sort of jump that occurred that we're not accounting for. And so it definitely, I think, can put people's mind at ease that it likely isn't some bioengineered warfare weapon because there are indicators within the genetic sequence that we'd be able to denote that this really doesn't seem right. Whereas it's not an absolute. We can't just look at the sequence and say, ah, yes, this absolutely came from this sample or this host and it's been through 12 people since then. It would be great if we had that capacity. And it does provide us with the information to I think debunk some of these conspiracy theories. And I think put faith in our health institutions. And so we've sort of spoken about the importance of the sequencing effort and why it's being done. Could you more like tell us now about what it's like doing the actual work itself, presumably you're placed within a lab, is that within Cambridge and, you know, for non-biological people out there. Could you sort of maybe briefly explain how sequencing actually happens, because it's kind of like this big black box, you just put it in and like sequence, but, you know, it's quite an interesting process. So no, yeah, absolutely, because it's definitely something, I think every scientist can think about the first time they're introduced to a new technique, and it definitely does feel like a black box. So our lab is well positioned to kind of aid the hospital. We're actually in Adam Brooks. So we're one floor below the diagnostics laboratory. And so what's really great about the technique that we use, we're able to work with their extracting samples upstairs. So at no point are we that is really real time. Yeah, it's fantastic. So we're never handling infectious material, which is great. So that means that all of the work that we're doing in the lab is done at a class one class two level. So we're not having to work in a BSL three facility under a respirator. We're not dealing with patients directly, having concern with PPE. So as long as we're enforcing a really good social distancing practice within our lab, we're not at risk of our samples. I think we're more of a risk to our samples of anything because there's a lot of handling that has to be undergone. So after the diagnostics lab identifies a positive sample, we are notified and then given some of the extracts, and it's with these extracts that we can start our process. So we use what's considered a multiplex PCR amplicon sequencing process, which is a lot of words, but to break it down simply, RNA viruses can degrade really easily, or we potentially have a really low sample. So we'll get positive samples from upstairs from the diagnostics lab. And the very first thing we'll do is we will take whatever sample we have and convert it into CDNA. So this immediately takes our sample and makes it more stable. And then we'll use a multiplex PCR. And what this is, is we take these short primer regions, and we amplify small fragments of the RNA. And so it's, think about it, making a massive jigsaw puzzle. And so we'll take our small amount of sample and amplify so that there's more of it to work with. And we'll produce these small regions, which great about this approach is we can have a partially degraded sample, and be able to still recover a significant part of the genome. And as a result, too, there is this high risk of cross contamination. And so we have to handle for that with our lab protocol. This is probably one of the bigger difficulties with this genomic sequencing technique, because we have all these tiny puzzle pieces that we have to avoid having them swap boxes, because obviously it would then ruin our complete image. In order to help deal with that, we have implemented some techniques in the lab to help avoid this. And one of them, a relic of this project really starting in a field environment is we work out of these black tents. They were designed to be able to set up a field sequencing unit or lab in any part of the world that doesn't have access to traditional lab resources. So we have certain steps that are done inside these black tents that are originally used for hydroponic plant growing, but they are perfect because material allows them to be DNA's apt, as well as take a UV light treatment so we can continually keep them sterile and help them keep two different libraries separate and to help them again prevent that cross contamination issue that can occur from working in this type of multiplex PCR. But then it, too, it makes it really accessible. You don't need to in theory have this fancy lab to work in. You can see this approach in these publicly available protocols on the Arctic Network and associated with Cog UK and a lab in the US or in South America or in Africa or in Southeast Asia could pick up these protocols, use the basic resources that they have access to and be able to begin sequencing and adding to what is right now in international cause, which is really great. It's awesome that you can take it anywhere as well. When they're using the field, what makes them really great is that they're lightweight, but obviously that's not a concern for us here in Cambridge. So we have our lab designed in a way that samples work in a progressively dirtier manner or cleaner manner so that our fresh samples are handled in one room and then aggressively moved through the process to prevent the risk of cross contamination. So after amplifying this data, we then prepare it so that it can interface with our Oxford Nanopore grid ion sequencing system. So it's a series of steps in which we just prepare the individual DNA sequences so that they can interface with the instrument and be able to be sequenced. And so that involves adding a barcode onto them so that we can pool the samples together and run more than one sample at once. So we can run up to 24 samples on a single flow cell with this kind of technology. And then it also then allows for us to clean up and remove any excess material and wash buffers and other DNA that's going to be in the sample that we get that's been extracted because it's not as simple as just poking the virus in particular. We're going to get other genetic material that is associated with you as a person as a patient. And so then the workflow roughly could be handled within a 12 to 24 hour period. But what we traditionally do is follow about a 36 to 48 hour turnaround. So we'll get a sample from upstairs around noon, run it through the PCR step and that will take us through overnight and the next day prepare it to be sequenced and then run an overnight sequencing. And at that point, we'll have a genome by midday the next day. And so it's pretty amazing to know that this kind of technique is rather recent 10 years back. This kind of technology wasn't there yet. So it's definitely our ability to respond to this outbreak is enabled by the advances, by both the Arctic Group to get this protocol to a point where it can be well established in efficient and optimized, but also by the massive rates that have made these protocols to start with. And yeah, impressive. Yeah, I know that's that's a really good point actually because knowing how much the sequencing world has just taken off over the past 10 years, I mean, the little Oxford nanopore things are great. Yeah. They're constantly getting improved as well. And yeah, it is kind of scary to think, yeah, if this had happened 10 years ago, we wouldn't have nilious as much knowledge as we do now already. So I think that's a really good positive that I hadn't really considered at this time. I guess it also means that there's hope for this kind of thing does happen again. Not only will we have better technology probably, but also it'll help us identify what the things that we need to improve and are lacking the most are. So hopefully that will motivate those gaps to be filled and push us in the right direction for the future. Exactly. And I was going to ask as well, because obviously you said you're doing your Enfield. I'm not sure if your Enfield was meant to be on the SARS-CoV-2 virus, probably not. So what made you, first of all, what made you volunteer for this kind of scheme? But also, what is it like to really be working on something where you can kind of see the impact that it's having on a very fast turnaround to what you're used to? Doorbelly in science, at least in like, you know, just your master's or graduate research. And yeah, kind of what that's like. So bottom line up front, yeah, it's very exciting to be involved with it. I remember a couple of weeks ago, a friend asking the question, like, is this the time for viral just to be alive? Like, is everyone just really up and excited? And I'd argue with no person is excited about an outbreak. But and especially any kind of scientist definitely wants to be doing the research of interest by I do research on neurobiosis traditionally. But anytime you feel like you have a specific skill set potentially that could be helpful in any way, you want to provide help when you see people in need. I think there were some like a thousand people that signed up within Cambridge to help out with a lot of the initiatives online. And they were just overwhelmed with volunteers with really, like, practical and like high level skill sets. And so having this unique opportunity to work within genomics, which definitely wasn't my background, but having the basics and PCR that could allow experts to come in and teach is really, really neat. And it's been able to in some ways, I think, diversify my install and give me a more broad exposure than I definitely would have got. So in that, it's I think really satisfying to be able to see the impact because like you said, scientists don't always get to see the direct impact. They get to know what it potentially could use for. So there's, I don't want to say selfishly, but selfishly, there's something slightly satisfying about it, but obviously no one wants to be in the circumstances that we are. No, definitely not, but yeah, it's sort of quite an amazing opportunity to be involved with all of it because of the fact that you're involved in it and I guess your family and friends must know. I mean, I don't even really work in viruses and more of a bacteria kind of gal. And I've been like flooded with questions, as if I know all the answers, and I was just wondering, you know, especially now considering you're working with it, but are you always getting questions from your family and relatives and weird articles and videos and is there a day, leisure things like that or? Absolutely. And I think most people involved with science tend to get those questions in general when anything happens. It's just one of the things you answer the questions that you can and then you do the ethically right thing to do and not answer the questions you don't know, because that just does exactly that prevents the spread of misinformation. And so I think that's what's really empowering for even the scientists that maybe don't have the opportunity to right now be in the lab and do something is be scientific advocates. Science and paper, proper reading of a journal article, because so many, I don't know about you guys, but so many of my friends and family have now taken to science for journals trying to read through an article and make sense of it, but not recognize, oh, this is a preprint. It hasn't been reviewed yet or you take it as a data point, not as the absolute new standing information change, all your practices and policy, this one single journal article. And so any person that's involved with any form of science has the ability to help inform these basic techniques to their family and friends, and so I think that's really important. Yeah, no, it's definitely a really good opportunity to do that. I think it's the confidence as well to critically appraise what you're being told. Yeah, and definitely like you said, to be able to say, look, I'm not the right person to ask for this and point people in the direction of the experts that can actually answer those questions. And certainly the patients too, like when I just think about even being in this lab, obviously in no way am I well versed in any of this, I've been taught so much of the protocol and the approach, but I've been able to do my part now in this lab group, thanks to like the great minds and scientists that had the patience to work with me. So it's the minds of like Professor Ian Goodfellow who is helping run the group. Dr. S. Torek, who is over in the Department of Clinical Medicine, who's working really closely with us and helping us pair our genomic data with the epidemiology. It's the long list of scientists like Dr. Luke Meredith and Dr. Sarah Katty, Dr. Charlotte Holdcraft, who are coming from either within the lab or different parts of the UK with their own individual backgrounds with either genetic sequencing or outbreak response or just general biology background. And their experience is helping inform how our lab group works, helping inform how other groups work, and their patients to then teach new students who can then one day hopefully fill their roles. And so if we can take that same level of patients even in our local communities, we can help make that same level of impact even if we're not directly involved. It's a real testament to how science can work so efficiently and so well. Absolutely. Well, thank you so much for chatting to us. I appreciate you're massively busy at the moment, and we really appreciate it. I mean, our listeners can't see you, but you're like in the lab, in the office right now. I'm actually sitting right next to the flow cells that we have right now. Amazing. One step removed, don't we, somebody? Yeah, yeah, definitely virtually in social distancing away from the samples. Okay. Well, thank you so much again. Well, thank you for having me. Thanks for tuning in. We hope you enjoyed the episode and learn something about sequencing the genetic code of the arts of the RNA. We have another interesting coronavirus accident lined up for you in two weeks time. So do hit the subscribe, follow whatever button it is on the whichever podcasting platform you're listening to us on. If you want to get in touch with us with any questions or suggestions for what you want to hear from us, especially at this time where we are kind of responding to what's going on in the world real time. We would love to hear about what you're curious about. You want to know about the COVID-19 situation. So please get in touch with us using either Twitter to contact us at DulucidePond. Or you can send us an email or email is [email protected] and yeah, just don't forget to follow us and leave a review. If you've enjoyed it, leave a little rating, and we'll be back with another start.

Podcast Summary

Key Points:

  1. The COG-UK consortium rapidly sequences SARS-CoV-2 genomes to track viral mutations and transmission in real-time, aiding public health responses.
  2. Genomic data, when combined with clinical and epidemiological information, helps map outbreak progression, identify transmission routes, and investigate disease phenotypes.
  3. The sequencing effort builds on prior protocols from the Arctic Network, uses accessible techniques like multiplex PCR, and involves decentralized nodes collaborating with local hospitals for timely data.
  4. While sequencing can debunk some conspiracy theories about the virus's origin by showing natural evolution, it cannot provide absolute answers about every transmission event.
  5. The work is conducted safely using non-infectious samples and portable methods, allowing for potential global adoption of the sequencing protocols.

Summary:

In this podcast interview, Grant Hall from the COG-UK consortium explains the critical role of sequencing the SARS-CoV-2 genome. The initiative, building on the Arctic Network's experience from past outbreaks like Ebola, rapidly sequences virus samples from across the UK. This genomic data reveals mutations (about 1-2 per month) and, when integrated with hospital clinical data, helps track transmission patterns and viral spread geographically.

The sequencing process itself involves converting viral RNA to more stable cDNA, amplifying it via multiplex PCR, and using nanopore technology to read the genome within 36-48 hours. Importantly, this work is done safely with non-infectious extracts. While sequencing cannot definitively trace every infection or fully prove the virus's natural origin, it provides powerful real-time evidence to inform public health decisions and counter misinformation.

The decentralized, collaborative model between academic nodes and local hospitals enables fast, adaptable responses to the pandemic.

FAQs

The Blue Side Podcast, produced by Cambridge University Science magazine, features interviews with local researchers, university staff, students, and science professionals to discuss their work and behind-the-scenes activities.

Listeners can get in touch by tweeting @BlueSidePod or emailing [email protected] with questions, suggestions, or to be featured on the podcast.

The COVID-19 Genomics UK (COG-UK) consortium is a government-funded initiative that sequences the genome of SARS-CoV-2 to track mutations and understand the virus's spread, involving universities and public health institutions across the UK.

Sequencing the virus's genome helps track mutations, map transmission routes, and inform public health policies in real time, aiding outbreak response and potentially guiding vaccine development.

Genomic data can indicate if the virus is naturally occurring rather than bioengineered, providing evidence to counter claims about lab origins, though it cannot pinpoint an exact source with absolute certainty.

The Arctic Network, a group of universities with experience from outbreaks like Ebola, provided optimized protocols for real-time genomic sequencing, enabling rapid setup and data collection for COG-UK.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.