Longevity & AgingForschungsarbeitOpen Access

Multi-omics screen pinpoints 32 candidate causal genes and three subtypes in Long COVID

A computational framework combining Mendelian randomization and network control theory flags 32 likely causal Long COVID genes and three symptom-based subtypes.

Freitag, 9. Oktober 2026 0 Aufrufe
Veröffentlicht in PLoS Comput Biol
Glowing protein-interaction network with highlighted hub nodes around a DNA double helix and a stylized coronavirus particle

Zusammenfassung

Long COVID affects an estimated 10–20% of people infected with SARS-CoV-2, yet the genetic drivers are poorly understood. Researchers combined transcriptome-wide Mendelian randomization, which tests whether gene expression causally affects disease, with network control theory, which finds key control points in protein-interaction networks. They integrated eQTL, GWAS, RNA-seq and PPI data and prioritized 32 candidate genes, 19 previously reported and 13 novel. The genes cluster in SARS-CoV-2 response, viral carcinogenesis, immune regulation and cell cycle control. The genetic architecture overlaps with metabolic, autoimmune, connective tissue and syndromic disorders. Expression of these genes also separated patients into three symptom-based subtypes. The team released an open-source Shiny app for exploring the results. The findings are computational and hypothesis-generating, and need experimental and clinical validation.

Detaillierte Zusammenfassung

Long COVID, or post-acute sequelae of SARS-CoV-2 infection (PASC), affects an estimated 10–20% of people who have had COVID-19. It causes persistent, multisystem symptoms. Age, sex, comorbidities, vaccination status and smoking all modify risk, and inflammatory, immune and coagulation biomarkers have been linked to the condition. Which genes actually drive it, as opposed to merely correlating with it, remains unclear. That gap limits targeted treatment, diagnosis and outcome prediction.

The authors built a framework that combines two complementary strategies. The first is transcriptome-wide Mendelian randomization (TWMR). It uses genetic variants that regulate gene expression (eQTLs) as instruments to test whether altered expression causally influences Long COVID risk, using GWAS data. The second is control theory (CT) applied to a human protein-protein interaction network. It identifies network driver genes, the nodes that most strongly influence the stability of the disease-related network. These were integrated with RNA-seq data to prioritize genes that either causally affect susceptibility or help control the network.

The pipeline prioritized 32 candidate genes. Nineteen had already been reported in the Long COVID literature, which supports the approach, and 13 were novel. Among them were genes acting as risk or protective factors and genes acting as network drivers. Enrichment analyses pointed to the SARS-CoV-2 response, viral carcinogenesis, immune regulation and cell cycle control. The genes also showed shared genetic architecture with syndromic, metabolic, autoimmune and connective tissue disorders. This overlap may help explain the heterogeneous, multisystem presentation of Long COVID.

Using the expression profiles of the causal genes, the authors clustered patients into three distinct subtypes with different symptom profiles. This suggests different underlying mechanisms and a possible basis for stratified diagnosis and treatment. They also released an open-source Shiny application. Users can adjust the MR and CT parameters, generate gene lists and explore the results interactively, and analysis code is available on GitHub.

For clinicians and researchers, the main value is a ranked set of mechanistic hypotheses and potential therapeutic targets, not a clinical tool. The gene lists are statistical inferences from existing datasets, and the 13 novel genes have not been experimentally confirmed. MR carries known risks from weak instruments and pleiotropy, and the subtypes need replication in independent cohorts. This summary draws on the abstract, introduction and framework overview, so detailed cohort sizes, thresholds and effect estimates are not reported here.

Wichtigste Erkenntnisse

  • The framework prioritized 32 putative causal Long COVID genes, 19 already reported in the literature and 13 novel.
  • Candidate genes were enriched in SARS-CoV-2 response, viral carcinogenesis, immune regulation and cell cycle control pathways.
  • Long COVID showed shared genetic architecture with syndromic, metabolic, autoimmune and connective tissue disorders.
  • Causal gene expression profiles separated patients into three symptom-based Long COVID subtypes.
  • Genes were classed as risk factors, protective factors or network drivers, and an open-source Shiny app lets users explore them.

Methodik

Computational study integrating transcriptome-wide Mendelian randomization (eQTL instruments with GWAS data) and control theory on a human PPI network, with RNA-seq data used for expression profiling and patient clustering. Enrichment analyses and literature cross-checking supported the gene prioritization. Details here come from the abstract, introduction and framework overview.

Studienlimitierungen

Findings are purely computational and inferred from existing datasets, with no experimental validation of the 13 novel genes. MR methods depend on strong instruments and can be confounded by pleiotropy, and the subtypes and gene lists need replication in independent cohorts. Ancestry, sample size and Long COVID case definitions, which vary by organization, may limit generalizability.

Hat dir diese Zusammenfassung gefallen?

Erhalte die neueste Longevity-Forschung jede Woche in deinen Posteingang.

E-Mail-Adresse zum Abonnieren eingeben: