How Much Do RF Drone Benchmarks Overstate? A Controlled Study and Theory of Data Leakage in UAV Signal Identification

Radio-frequency (RF) sensing is a central modality for counter-unmanned-aerial-system (counter-UAS) defence because it exploits the control, telemetry, and video links between a drone and its operator. Reported accuracies for RF-based drone detection and identification are often very high, but many are obtained using cross-validation that splits a small number of continuous recordings into short segments. This can place near-duplicate slices of the same recording in both training and test partitions, creating data leakage. We study this leakage problem through theory and measurement. We formalise the optimism of segment-level cross-validation and show, using Cover's function-counting theorem, that a classifier can exactly memorise the recording-to-label map when the number of independent recordings, R, is small relative to the feature dimension, d. In particular, this can occur when 2R is less than or approximately equal to d. Under these conditions, naive accuracy approaches 1, and the inflation gap approaches 1 - ACC*, where ACC* is the Bayes accuracy. The inflation eases only once R grows beyond this separability threshold. A controlled synthetic experiment with 10 seeds confirms the predicted curves: naive balanced accuracy rises from the Bayes level toward 1.0 as recording-specific nuisance variation grows, while honest recording-grouped evaluation declines to chance, with a gap reaching about 0.5. On the public DroneRF dataset, pooled leave-one-recording-out cross-validation shows drone type identification, AR versus Bebop, collapsing from a naive macro-F1 of 0.74 to 0.46, the two-class chance level. A leakage-pathway ablation attributes essentially all of the inflation to segment-level leakage.

Paper

References (11)

05“DroneDetect dataset: A radio frequency dataset of UAS signals for machine learning detection and classification,”2021 · IEEE DataPort
06Pool, then score. For leave-one-recording-out with single-class folds, pool out-of-fold predictions and compute one metric with fixed labels; report a bootstrap confidence interval
07Group, do not shuffle. Partition by recording ( GroupKFold or leave-one-recording-out), never by segment
08Check the regime. Report the number of independent recordings per class; if 2 R ≲ d (Corollary 1), treat identification claims as unsupported regardless of accuracy
09Report both. Always present naive and grouped scores side by side; their gap is the leakage estimate
10a controlled synthetic experiment (10 seeds) that confirms the predicted dependence on nuisance strength and on R (Section 6)
11a theoretical prediction of the leakage-inflation curves: using Cover’s theorem,the naive split memorises recordings while 2 R ≲ d , so naive accuracy → 1 and the gap →

Similar papers

© 2026 NYSGPT2525 LLC