Brian Zhang's blog

Statistics and other topics

Notes on an EBV DNAemia paper

Posted Aug 3, 2026 · 4 min read

Notes on “Population-scale sequencing resolves determinants of persistent EBV DNA”, Nyeo et al., Nature 2026.1

Usually in a large biobank, your phenotypes come from electronic health records or surveys of the studied population, including longitudinal follow-ups. In my PhD paper, we tested our method on serum biomarker traits in the UK Biobank, like mean platelet volume and LDL cholesterol: these were measured as byproducts of the blood samples collected for sequencing. This paper (Nyeo et al. 2026) introduces an ingenious new phenotype for biobanks: the level of persistent Epstein-Barr virus (EBV) that can be detected in the collected blood sample.

The win here and the reason this rises to a Nature paper is that you get the phenotype for free in UK Biobank and All of Us. It reminds me of Po-Ru Loh’s paper, “Insights into clonal haematopoiesis from 8,342 mosaic chromosomal alterations”, that went back to the source intensity data of UK Biobank genotyping and used it to call mosaicism events.2

The design comes from a serendipitous fact, from the section “Rationale of EBV detection” in the Methods:

The 171,823-nucleotide EBV genome (NC_007605.1) was first included in December 2013 (hg38 version GCA_000001405.15) as a sink for off-target reads that are often present in sequencing libraries, to account for pervasive EBV reads present from the immortalization of LCLs (as with the 1000 Genomes Project and related consortia). Importantly, WGS in the UKB and AOU consortia was performed on whole blood, reflecting that EBV reads detected would derive from viral DNA from past infections.

I suspect this made it easier for them to get the needed data from UKB and AOU. However, they say it should be possible to extend their approach “to a broad range of viruses … from the Polyomaviridae, Adenoviridae, Parvoviridae and Anelloviridae families”, it just might require more work.

My favorite figures in the paper were Fig. 1b and 1c. After they align reads to the EBV genome, they need to mask out two regions in order to get a stable estimate of circulating EBV DNA, what is called EBV DNAemia in the literature. This ends up correlating very well with EBV serostatus.

I thought the implications of their EBV DNAemia phenotype with other diseases were less striking than I would have expected. Possibly it is because this reflects a single time estimate.

I enjoyed their peptide presentation analysis though. They find that HLA alleles which tend to present EBV peptides more strongly, as measured by NetMHC scores, are associated with lower EBV DNAemia (Fig. 5e and Extended Data Fig. 6e-g). I trust this result. Intuitively, patients with HLA alleles that present EBV more easily will see less circulation of EBV in blood (Fig. 5f). It is interesting that this effect is stronger for class II than class I.

Some additional peptide presentation results:


  1. I had seen this paper earlier but took a closer look after seeing it in the Genetics Podcast list of episodes↩︎

  2. “The core intuition is to harness long-range phase information to search for local imbalances between maternal vs. paternal allelic fractions in a cell population (Extended Data Fig. 1).” Free text here↩︎

comments powered by Disqus