GenoBERT: A Language Model for Accurate Genotype Imputation

Genotype imputation enables dense variant coverage for genome-wide association and risk-prediction studies, yet conventional reference-panel methods remain limited by ancestry bias and reduced rare-variant accuracy. We present Genotype Bidirectional Encoder Representations from Transformers model (GenoBERT), a transformer-based, reference-free framework that tokenizes phased genotypes and uses self-attention mechanism to capture both short- and long-range linkage disequilibrium (LD) dependencies. Benchmarking on two independent datasets—the Louisiana Osteoporosis Study (LOS) and the 1000 Genomes Project (1KGP)—across ancestry groups and multiple genotypes missing levels (5–50%) shows that GenoBERT achieves the highest overall accuracy compared to the other four baselines (Beagle5.4, SCDA, BiU-Net, and STICI). At practical sparsity levels (⩽ 25% missing), GenoBERT attains high overall imputation accuracies (r2≈0.98) across datasets, and even maintains robust performance (r2>0.90) at 50% missing level. Experiment results across different ancestries confirm consistent gains for both datasets, with resilience to small sample sizes and weak LD. The 128-SNP (single-nucleotide polymorphism) context window (≈ 100Kb) was validated through LD-decay analyses as sufficient to span local correlation structures. By eliminating reference-panel dependence while preserving high accuracy, GenoBERT provides a scalable, and robust solution for genotype imputation and a foundation for downstream genomic modeling.

Paper

Similar papers

© 2026 NYSGPT2525 LLC