The answer was already there: how long-read sequencing is ending the diagnostic odyssey

When rare disease cases go unsolved, the answer isn't always missing - sometimes it's just hard to find. In a new study, Prasun Dutta and colleagues used long-read sequencing to diagnose three families whose cases had defied standard analysis, revealing the crucial variants had been in their data all along. Analysing 3.7 trillion bases of sensitive genomic data called for secure, high-performance computing - and EIDF was proud to provide it. Here's the story behind the science.

In some cases, the answer is already present in the genomic data. The challenge is finding the right computational approaches to translate that massive scale of information into meaningful insights.

From DNA sequencing to possible answers

For families living with a rare genetic condition, getting a diagnosis can take years. Patients may see multiple specialists and undergo repeated tests without ever receiving a clear answer. This is often described as a “diagnostic odyssey”.

Some rare disease cases remain unresolved not because the genomic data are missing, but because the relevant variant is difficult to detect, prioritise or interpret. Short-read genome sequencing (SRS) has transformed our ability to diagnose rare genetic disorders, particularly when the underlying cause is a small change in the DNA sequence. However, it is less effective at detecting some larger and more complex changes, known as structural variants (SVs).

In our recent study, published in the European Journal of Human Genetics (https://www.nature.com/articles/s41431-026-02210-x), we looked at whether long-read sequencing (LRS) could help resolve some cases that had remained undiagnosed after extensive SRS analysis. Working through the Scottish Genomes Partnership (SGP), we studied 24 families in whom a genetic cause had not previously been identified.

Standard SRS typically reads DNA in fragments of around 150 base pairs. This works very well for identifying small variants, but SVs, including larger insertions, deletions, duplications, inversions and translocations, can be much harder to detect and interpret, particularly in repetitive regions of the genome.

To overcome this, we used Oxford Nanopore Technologies (ONT) LRS. This technology reads thousands of bases in single, continuous runs, easily spanning those complex, repetitive regions, making it easier to examine complex regions of the genome and identify SVs.

Using this approach, we found a genetic answer in three out of 24 families:

  • Family 1: A 1.4-megabase inversion on chromosome 7 disrupting the AUTS2 gene, explaining a child's severe neurodevelopmental delay and epilepsy.
  • Family 2: A 5.2-megabase inversion near the DLX5/DLX6 genes in a mother and her daughter, explaining their split hand/foot malformation. The inversion altered the relationship between these genes and their regulatory elements, providing a likely explanation for the condition.
  • Family 3: A 59.8-kilobase deletion in the FN1 gene in a boy with unexplained skeletal abnormalities.

Crucially, when we retrospectively re-analysed the original SRS data using SVRare (https://www.medrxiv.org/content/10.1101/2021.10.15.21265069v2), we found that all three variants were present there all along. They had simply been missed because systematic pipelines to look for SVs weren’t standard practice at the time. This shows that, in some cases, answers may be found by re-analysing existing genomic data with updated tools and approaches.

Handling this massive volume of genomic data requires substantial computational resources. To put that scale into perspective, our raw input alone consisted of 24 FASTQ files containing over 3.7 trillion bases of DNA sequence. Because sensitive genetic information demands strict privacy controls, EIDF (operated by EPCC), supported through the University of Edinburgh’s Data-Driven Innovation Programme, provided the highly secure, high-performance computing (HPC) infrastructure needed to safely store, handle, and analyse this massive genomic dataset.

The work was funded by the Chief Scientist Office of the Scottish Government Health Directorates and the Medical Research Council. I am also grateful for personal support from the NIHR Biomedical Research Centre: Oxford (BRC), which made my involvement possible.

For the families involved, identifying the underlying genetic cause can bring an end to years of uncertainty. A diagnosis is a real turning point. It helps families get the right support, allows doctors to focus on the treatments that will work, and opens the door to new therapies or clinical trials that could make a real difference in their lives. Our findings suggest that SVs deserve more systematic attention in rare-disease genome analysis. In some cases, the information needed to make a diagnosis may already be present in existing sequencing data. The challenge isn't collecting the data but developing the right methods to finally make sense of it.

The Scottish Genomes Partnership was funded by the Chief Scientist Office of the Scottish Government Health Directorates (SGP/1 and SGP/2) and the Medical Research Council Whole Genome Sequencing for Health and Wealth Initiative (MC/PC/15080), awarded to
Timothy J. Aitman
.
Prasun Dutta
was additionally supported by funding awarded to 
Jenny C. Taylor
. (Grant Ref: MR/W01761X/1) and by the NIHR Biomedical Research Centre: Oxford (BRC).