A deep learning framework for dynamically rendering personal sound zones (PSZs) with head tracking is presented, utilizing a spatially adaptive neural network (SANN) that inputs listeners’ head coordinates and outputs PSZ filter coefficients. The SANN model is trained using either simulated acoustic transfer functions (ATFs) with data augmentation for robustness in uncertain environments, or a mix of simulated and measured ATFs for customization under known conditions. Numerical experiments are conducted using an in-house PSZ rendering system with a linear loudspeaker array in the frequency range of 100-1500 Hz. It is found that augmenting room reflections in the training data improves model robustness more effectively than augmenting system imperfections, and that adding constraints related to filter compactness to the loss function does not significantly affect isolation performance. Comparisons of the best-performing model with traditional filter design methods show that, the SANN model achieves comparable or superior robustness in an unknown room environment without requiring explicit regularization. Furthermore, benchmark tests on a laptop demonstrate that the model offers greater data compression and computational efficiency than traditional rendering approaches, supporting SANN as a viable option for real-time applications.