Origins of the IBDverse: how a conversation in a carpark became one of the largest gut cell studies
Sometimes research is born in unexpected places—like an underground carpark.
In 2015, Tim Raine and Carl Anderson had just published the first transcriptomic analysis of T cells in the human gut using expression microarrays, in a handful of patients and at the scale of a few thousand cells. “At the time, this was genuinely cutting edge,” said Tim, a consultant gastroenterologist at Cambridge University Hospitals, based at Addenbrookes Hospital. “Nowadays, no one uses this technology any more.”
For their next project, they wanted to be ambitious. Over the course of a casual conversation in the Wellcome Genome Campus’ underground carpark, they came up with the idea of using single-cell sequencing to gain insights into the genes and cell types causally driving Inflammatory Bowel Disease (IBD), a project christened the “IBDverse”.
A gamble on single-cell sequencing
Genome Wide Association Studies (GWAS) have identified many variants associated with IBD risk, but these tend to lie outside of genes in regulatory regions of the genome, making it difficult to unpick their effect on disease. Tim and Carl, Head of Human Genetics and Senior Group Leader at the Wellcome Sanger Institute, were working on the assumption that such variants must have their effect on IBD by determining the expression of nearby genes. To identify those genes, they would need very detailed maps of gene regulation, linking regulatory variants to the genes they control.
The team’s main gamble was to suggest that regulatory variants must only have their effect in restricted cellular contexts. A variant with a profound effect on the expression of an important gene in many cell types would be too disruptive, and would be selected against. Therefore, to identify the effector genes for their variants of interest, they would need to map the variants in specific, disease-relevant cell types.
Tim and Carl argued that the best, and most high-throughput way to map regulatory variants was to perform single-cell RNA sequencing, letting the data cluster to define the cell types present in the tissue sample, and then mapping the genetic variants that alter gene expression within each of those clusters.
Single-cell techniques were just emerging, but were mostly being applied to model systems—idealised conditions compared to human tissue which is more complex to analyse, but much more informative about disease biology. Tim and Carl hypothesised that some of the variants’ effects would only be apparent in disease conditions, so it was crucial to obtain some samples from patients, rather than lab grown tissue models. Also, since some of the cells in the gut are present in very small numbers, some of the genetic variants are very rare, and some of the changes in expression they were looking for might be very slight, the team would need hundreds of samples to get enough data for a meaningful analysis. “This was orders of magnitude harder than anything we’d done up until this point,” said Tim. “I think it’s fair to say that it wasn’t at all clear that it could be done.”
Another argument in favour of using a new technique was that a project of this scale would take years, but sequencing technologies move fast. “The samples would all need to be processed with the same protocol, so if we’d started the project with an older, more established technology, by the time we finished, our data might be obsolete,” explained Carl. “We were almost duty-bound to start with something new and shiny, something that was a bit of a gamble. Single-cell sequencing was cutting edge technology, and Sanger was at the forefront of that. If you had to take a leap of technological faith, Sanger was the place to do it.”
To get support for a project of this scale and ambition, Tim and Carl approached multiple potential sources of funding—Open Targets, the Wellcome Sanger Institute, and the Crohn’s and Colitis Foundation—each of whom financed a section of the project. The data from the IBDverse was shared in real time with Open Targets’s pharmaceutical partners, ensuring that the research would be translated through to patient benefit.
A long-standing collaboration
Tim and Carl, who have now been working together for around 16 years, have an unusually engaged model of scientific collaboration. “The tried and tested model of clinical/academic collaboration is for the clinical partner to collect, curate, then hand off the samples to the academic team for analysis in one or multiple batches,” said Tim. “That wouldn’t work in this case. We were analysing the data as it was coming in, and our ongoing results fed back into the project, for example informing the types of patients we recruited, or our approach to biopsy collection.”





Photos from the IBDverse team in action
Tim explained that one of the reasons their collaboration works so well is that he has worked as a computational biologist and Carl has worked in the clinic. “Carl’s lab has developed a culture of working with clinicians where each team goes into each other’s environments, sometimes out of their comfort zone, to have that mutual understanding and respect.”
In this project, the clinical and academic teams needed to have a shared understanding of the work each was undertaking, and live discussions of the analyses. This type of information sharing meant that the academic research questions addressed the most relevant issues in the clinic, and that results could be interpreted using clinical context. “We would regularly go down rabbit holes in our project meetings, because people would ask fascinating and thought-provoking questions about the methods we were using or the interpretation of the data,” said Tim. “That comes from having these different perspectives and a culture of open exchange. And that’s where some of the best bits of the paper come from.”
Setting up the project
“The IBDverse is special to me for another reason,” explained Carl. “My work up to that point had been entirely computational, but we set up our own wet lab to do this project.”
Since single-cell technology was very new, the protocols had to be established and troubleshooted from scratch. “We had to get single cells from solid tissue. But as soon as they are out of the body, the cells can become stressed and die,” said Carl. “We had to process them in a way that preserved the transcriptome enough to get meaningful results. That takes a lot of thought and skill and care.”
To get the most accurate picture of gene expression, biopsies had to be processed as soon as possible. The short, 20-minute drive between Addenbrookes hospital and the Wellcome Sanger Institute was key to this. The Sanger team would get a call when eligible patients who consented to the study arrived at Addenbrookes. As soon as the sample was ready, it was placed on ice and couriered to Sanger to be processed. Rebecca McIntyre, then a senior staff scientist at Sanger, now a VP at Relation Therapeutics, developed a processing method by testing multiple kits and techniques to find the one that would best preserve gene expression. The choice of method was crucial since all the samples would have to be processed in the exact same way to minimise the impact of technical variation on the analysis.
“This also wasn’t something that we could pipeline,” noted Carl. “Each of the samples had to be processed individually, whenever they arrived. And they could arrive at very unsociable hours.” Over the course of six years, the wet lab team processed a total of 1193 samples from 646 individuals.
Halfway through the project, the Covid pandemic brought patient recruitment to a halt. Although they’d already collected a large number of samples, more than anything that had been published at that point, the team still had nowhere near enough to do the genetic analysis they were aiming for. “Tim and I never lost faith that we could get there,” said Carl. “During the pandemic, we had to convince our funders that we would be able to catch up and still deliver on our original project, but as soon as we saw that we could generate robust data from the samples, we knew that it was just a matter of time, and patience.”
“What we couldn’t have predicted, what we could only hope for, is that our original hypothesis was correct,” added Tim. “There were no guarantees until we saw the final data.”
The largest single-cell RNA sequencing dataset of IBD-relevant tissues
Two publications drawing on IBDverse data were published in Nature and Nature Genetics earlier this year, investigating different aspects of the disease.
By sequencing the RNA in each cell, the team was able to quantify which genes were expressed, and how strongly. By contrasting this with each individual’s genotype, they identified variants that alter the expression of a gene, known as eQTLs.
This analysis nominates effector genes and cell types at over half of known IBD loci, including 74 for which this is the first candidate effector gene. The pattern of effector genes suggests that IBD is underpinned by immune system dysregulation and a failure of the gut lining to repair itself. Many of the genes regulate pathways that were previously underappreciated in the context of IBD risk, and accumulated in specific cells which are not commonly associated with the disease.
Notably, many eQTLs were only found when analysing cell-type level data. Cell-type level eQTLs were more than twice as likely to colocalise with IBD GWAS loci than tissue-level eQTLs.
The IBDverse is a key step to resolve the biology of IBD and nominate targets for drug development or repurposing, and demonstrates the power of the single-cell approach to resolve the role of genetic variants. “What’s really exciting about this, in addition to the many lessons we’ve learned about IBD biology, is that the same approach can be used to unlock the biological mysteries of many different diseases,” said Carl. “Single-cell sequencing at scale provides a high-resolution view of disease biology, and by combining that with genetic variation, we can now make the insights needed to drive better drug target identification.”
Future work
Tim and Carl are already excited about further analyses they can undertake with new technologies. “We generated these samples and analysed them with a technique that was cutting edge when we conceived the project,” said Carl. “But we now have the possibility of applying new technologies to the data in ways that we couldn’t have conceived of in 2015.”
Their recently published work uses short read RNA sequencing, which quantifies how much a gene is expressed. They would now like to apply long read sequencing, to uncover which versions of the genes are expressed. “That’s something we’d never thought about at the start, but when designing the project we future-proofed it as much as possible, and so we’re now in a position where we have the right samples and the right data to be able to take on these new challenges.”