DeepSeq: High-Throughput Single-Cell RNA Sequencing Data Labeling via Web Search-Augmented Agentic Generative AI Foundation Models
Generative artificial-intelligence foundation models offer transformative potential for processing structured biological data, particularly single-cell RNA sequencing (scRNA-seq), where datasets are scaling exponentially toward billions of cells. A persistent bottleneck for supervised learning on these data is annotation: cell-type labeling remains slow, expertise-intensive, and difficult to reproduce at scale. We present DeepSeq, a modular and reproducible pipeline that applies large language models (LLMs) to automate cell-type labeling from the top marker genes of unsupervised clusters, and we use it to test whether retrieval capability or raw model scale is the dominant driver of annotation accuracy on a single benchmark dataset (PBMC3K; N=8 clusters under the standard filter). DeepSeq supports both offline local inference and online inference with real-time web-search retrieval, and couples a two-stage evaluation protocol—marker-gene verification followed by ontology-aware label scoring—to marker-inferred reference labels. An archived agentic GPT-4o configuration reaches 75.0% exact-match agreement with these labels (95% Wilson CI [40.9, 92.9]; N=8 clusters) without any task-specific fine-tuning. A controlled same-model ablation shows that adding retrieval to a small local model helps substantially (Llama3.2-1B, Δ=+25 percentage points), but same-model GPT contrasts are mixed ( +12.5 pp for GPT-4o, −25.0 pp for GPT-3.5-turbo): across nine LLMs, retrieval is model- and dataset-dependent rather than a universal driver of accuracy, and LLM labeling does not surpass a specialized reference classifier (CellTypist, 87.5% exact-match) on this benchmark. We delineate the failure modes—marker ambiguity, ontology granularity, and retrieval noise—that bound current performance, and provide full reproducibility scaffolding. DeepSeq offers a concrete, auditable substrate for high-throughput atlas construction and for the labeled corpora that virtual-cell foundation models will require, aligned with community efforts such as the Human Cell Atlas and the Human Tumor Atlas Network.
Paper
References (25)
Scroll for more · 13 remaining