Research news
Researchers in China have developed a blood-based machine-learning model that combines fragment patterns, inferred methylation and copy-number changes to identify head and neck squamous cell carcinoma, although prospective studies are still required before it can support routine screening
A research team in China has developed a blood-based analysis that combines features of cell-free DNA to identify people with head and neck squamous cell carcinoma.
The findings suggest that complementary molecular signals can improve cancer detection when assessed together. However, the study has not established that the approach is suitable for routine population screening, where disease prevalence, participant characteristics and routes for follow-up differ from those in a case–control study.
The research team included Yakui Mou, Han Fang, Wenbin Zhang, Xicheng Song and Caiyu Sun at Yantai Yuhuangding Hospital, Qingdao University, Yantai, China. Collaborators also included Yang Shao and colleagues at the Geneseeq Research Institute and Renhong Guo at Jiangsu Cancer Hospital, both in Nanjing, Jiangsu Province, China.
The researchers called the model DECIPHER-HNSC. DECIPHER stands for ‘Detecting Early Cancer by Inspecting Plasma High-throughput sEquencing Reads’. And, HNSC stands for ‘head and neck squamous cell carcinoma’. Because all cancer participants had squamous-cell disease, the reported performance cannot be assumed to apply to other head and neck malignancies.
DECIPHER-HNSC analysed cell-free DNA (cfDNA) released into the bloodstream. In people with cancer, a proportion may originate from tumour cells but this contribution can be small at an early stage of tumour development.
The model combined focal copy-number variation, fragment-size profiles and a genome-wide methylation feature. Copy-number variation describes gains or losses of genomic material, while fragment-size profiles capture the physical distribution of DNA fragments. The study did not measure methylation marks directly. Instead, the researchers applied a deep-learning method to DNA cleavage patterns around cytosine–phosphate–guanine sites to infer a methylation status. All three feature groups were derived computationally from the same low-coverage whole-genome sequencing data.
The investigators trained the model with samples from 144 patients and 148 controls. An internal validation cohort contained 100 cancer cases and 96 controls, while an external cohort contained 95 cases and 98 control subjects.
The model achieved areas under the receiver operating characteristic curve of 0.962 in internal validation and 0.966 in external validation. At a threshold intended to provide 95 per cent specificity in the training cohort, it identified 78 of 100 internal-validation cases and 77 of 95 external-validation cases. Observed specificity was therefore above the study target threshold 95.8 and 96.9 per cent, respectively.
Sensitivity was however seen to be lower among patients with the earliest stages of disease. The model identified 54.5 per cent of stage I cases in internal validation and 68.4 per cent in external validation. Detection exceeded 80 per cent for stage III disease, and all stage IV cases produced positive results. This pattern matters because a test intended for early detection must be seen to be performing at a high standard before cancer has advanced to stage IV.
The study remained a retrospective comparison of people with confirmed cancer and healthy participants, rather than a prospective assessment of people whose disease status was unknown. Controls underwent extensive medical review, laboratory tests, chest imaging and specialist head and neck examination. The researchers also excluded recent infection, inflammation and surgery, and retained controls only if they remained free from cancer or major health deterioration for 12 months.
This process reduced the risk of undiagnosed cancer in the control group but produced a comparison with particularly healthy participants. People assessed in routine care may have benign conditions, inflammation, suspicious symptoms or premalignant lesions that could affect cfDNA and reduce apparent performance.
Sensitivity and specificity do not reveal how many positive results would represent cancer in the intended population. When a disease is uncommon, a test with high specificity can still produce substantial numbers of false-positive results. These could lead to imaging, biopsies and anxiety, while false-negative results could delay necessary investigations. Prospective studies must therefore establish predictive values within a defined population and clinical pathway.
All extraction, library preparation and sequencing took place at Nanjing Geneseeq Technology. Although samples came from two hospitals, the external validation did not test transfer between laboratories or sequencing platforms. Seven authors, including corresponding author Yang Shao, declared that they were Geneseeq employees, which strengthens the evidence for independent replication.
Notable limitations in the study include under representation of oral cavity cancers, incomplete tobacco and alcohol data, and the absence of systematic tests for both human papillomavirus and Epstein–Barr virus status.
However, the study has demonstrated the promise of a multimodal computational strategy derived from one sequencing assay. DECIPHER-HNSC is a research-stage model, not a clinically validated screening programme therefore prospective studies must now test it in realistic populations, confirm performance across laboratories, define follow-up after positive and negative results, and determine whether its use improves clinical outcomes without excessive unnecessary investigation.
For further reading please visit: 10.1038/s41698-026-01689-3
ILM 51.6 Sept 2026