Bioinformatics Research Center Expands Campus Access to Microbial Data
A partnership with the Research Facilitiation Service made a first-of-its-kind database more accessible and opened the door for future opportunities.
A vast database of more than 900,000 microbial genomes is now available to researchers across campus on the High-Performance Computing (HPC) Hazel cluster.
The Bioinformatics Research Center (BRC) and Research Facilitation Service (RFS) teamed up to make the database more widely accessible.
Classifying bacterial genomes, which have historically been grouped by physiological traits, is a challenge. The Genome Taxonomy Database (GTDB) standardizes microbial taxonomy based on entire genomes.
“This is really the first of its kind that is based on entire genomes. And in addition to that, it’s really bringing together, not just our reference genomes that we get from cultured microbes, but also these uncultured microbes which make up the vast majority of brand new microbes that are out there,” said BRC Director Bonnie Hurwitz. “It’s making it so researchers at NC State can really talk about these uncultured genomes in the same language as the rest of the world.”
Automation for Innovation
In addition to putting the database in a shared location, this partnership automates a lengthy and complex process researchers would otherwise have to handle themselves when accessing the GTDB. It also ensures everyone is using the most up-to-date data.
“The combination of the downloading and the pre-processing steps could take a couple weeks to do every time. It’s quite the process,” said Ford Fishman, an RFS research integration consultant. “Automating this saves researchers the time, effort and troubleshooting of trying to figure out how to do it on their own. It streamlines the whole process.”
The RFS is a collaboration between the Office of Information Technology, the NC State University Libraries and the Office of Research and Innovation that serves as a single point of contact for researchers to learn about available research computing and data resources.
According to Seth Weaver, a computational scientist who has taken over the technical reins on the GTDB project, having these databases and processes on the HPC is an efficient use of compute resources.
“We have a big cluster, but things fill up quicker than you would expect. And so when we can identify places where we can save space is always a plus,” Weaver said.
A Bridge to Discovery
The BRC hopes to continue working with the RFS to make more resources available on Hazel. One possibility BRC Assistant Director Louis-Marie Bobay believes could be valuable is a workflow to pre-digest and compare genes.
“There are all sorts of databases and all sorts of processes that have to be computed quite regularly by a lot of people. So having some of these pipelines set up could be extremely useful,” said Bobay.
The BRC’s collaborative vision includes using bioinformatics as a bridge to discovery. With the support of the RFS and services like HPC, the BRC is setting researchers up for the future by expanding data access.
“In the past, data have really been thought of as just kind of this byproduct of research. And now I think we’re starting to realize that data are really the first class citizen of where we need to go in this world of AI,” said Hurwitz. “We need to make sure our data sets are incredibly robust and readable and really, really ready for the possibilities that are waiting for us.”
- Categories: