Federated learning (FL) has emerged as a promising framework for distributed learning, enabling collaborative model training without sharing private data. Existing wireless FL works primarily adopt two communication strategies: 1) over-the-air (OTA) computation, which exploits wireless signal superposition for simultaneous gradient aggregation; and 2) digital communication, which allocates orthogonal resources for gradient uploads. Prior work on OTA and digital FL either enforces zero bias (explicitly or via assumed homogeneous path loss) or permits uncontrolled bias, yielding high-variance updates under heterogeneous channels and creating a performance bottleneck due to devices with poor channel conditions. We propose wireless FL updates that admit a structured, time-invariant model bias to achieve low-variance gradient aggregation, and analyze their convergence in a unified framework, in both strongly convex and non-convex settings. The resulting bounds reveal a bias-variance trade-off governed by the design parameters. To optimize this trade-off, we pose a non-convex joint design problem and develop a successive convex approximation framework to tune the parameters. Extensive experiments across heterogeneous wireless settings, covering both strongly convex and non-convex image classification tasks, compare the proposed OTA and digital designs against state-of-the-art baselines. The results demonstrate that optimizing the bias–variance trade-off through a structured bias yields faster FL convergence and improved generalization over existing schemes.