The significant computational demands for foundation model training have led to a growing concern on the centralization of the ecosystem behind their development. Although training with volunteered edge devices can democratize the process, existing approaches fall short: they struggle to match cloud-based training performance, exhibit limited scalability with model size, exceed device memory capacity, and have prohibitive communication overhead. We introduce a new communication-efficient paradigm, Cleave, which finely partitions training operations through a novel selective hybrid tensor parallelism method. Our evaluations show that we can match cloud-based GPU training using a large collection of edge devices, achieving 4–10x speedup over edge methods.