FedProbe: Attacking Federated Model Ownership Verification via Fine-Grained Watermark Detection and Erasure
Federated Learning (FL) enables collaborative model training across multiple parties while preserving data privacy. It has been widely adopted in sensitive domains such as healthcare and finance. As a result, the trained models themselves have become critical intellectual property. To prevent unauthorized use or distribution, watermarking techniques are commonly employed to verify model ownership. However, these mechanisms introduce new attack surfaces. Existing attacks typically assume the presence of watermarks and focus on their removal or forgery, overlooking a critical preliminary step: detecting whether a watermark exists. Furthermore, current approaches often rely on coarse-grained perturbation strategies that degrade model performance, limiting the practicality of such attacks. This paper proposes FedProbe, an attack framework from the perspective of an internal FL participant. FedProbe aims to detect and selectively remove embedded watermarks while minimizing the impact on model utility. It first detects the presence of a watermark by combining minimal inversion perturbation optimization with outlier detection. Then, it identifies critical pathways by analyzing neuron-level response differences under reversed triggers and performs fine-grained interventions via targeted neuron-level parameter replacement. Experimental results show that FedProbe achieves over 99% confidence in watermark detection across three mainstream datasets and various watermarking methods. Simultaneously, it reduces watermark verification success rates to 24% while keeping primary task performance degradation within 3%, demonstrating strong generalizability and robustness.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex