Summary
"SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection" describes a motivation and systematic approach to disentangling different components of speaking characteristics, in order to perform audio deepfake detection. This paper demonstrates a working 2-stage training pipeline, numerous ablations, and metrics over several English language focused datasets. In addition, qualitative demonstrations of the space of features learned and used by the model form a study which backs the philosophical approach of the paper with regard to disentangling stylistic characteristics of real speakers, in order to have a generalized defense against audio deepfakes.
Strengths
The overall approach taken here is useful, the problem is relevant, and the overall experiments cover a portion of the necessary ground to match the claims.
Qualitative studies both in the main body of the paper, and the appendix are generally interesting and it is worth considering if these experiments can be fit into the unified paper "flow" for this publication, especially given final conclusion writing. Overall ablation study and description of methods are well done, and form a high quality "core" to the paper. Architecture figures and ablation tables are both outstanding in terms of clearly communicating approach, and variations in the results. Use of open source toolkits, and sharing of hyperparameters should help reproducibility.
Weaknesses
The portion of the title "... Generalized Audio Deepfake Detection" claims a robust and general method for "audio deepfake detection". However, there are a broad range of prosodic styles from people with impaired speech, new learners of languages, children developing their ability to communicate and so on.
For a paper about generalized deepfake detection, where such detection keys particularly on the prosody of at least one example, I would like to see a larger study and example of voices in-the-wild, as opposed to the current examples which seem to be reasonably fluent speakers and performers, who may have "trained" speech patterns by and large. This is despite the use of the "in-the-wild" dataset, which doesn't seem truly wild in terms of robustness testing. Defense approaches should have some example study and discussion of False Positives beyond pure metrics (though the metrics discussions here are well done) - particularly when out-of-domain speaking patterns may have heavy overlap with "deepfake" data and the features used for classification, demonstrated as part of the paper. The examples in Figure 4. along with accompanied writing are a start down this road but not sufficient.
Mozilla Common Voice has a decent amount of this type of truly-in-the-wild speech for some qualitative study, and there are existing papers which use the same dataset for few-and-zero-shot TTS and voice conversion. The "in-the-wild" dataset here seems to largely focus on imitative TTS and voice conversion, and their "real" counterparts, which would generally point to celebrities, politicians and other public figures who (very likely) do not have the types of speech patterns mentioned previously. Though dysarthric speech is mentioned briefly in the limitations section, the issues which crop up from study of dysarthric speech are also found to some extent in many "typical" speakers as well, in more subtle ways so directly addressing this with some examples would strengthen the core claim of the paper in regards to "generalized detection".
MLAAD is multilingual, but some details of the dataset construction lead to limitations in its testing (outside the scope of this paper, beyond the continued critique that broader datasets and synthetic generation methods are needed to test generalization). However, here only the EN subset appears to be used - which again reduces the claims from the title since it means the bulk of testing is on English locales. This is not a problem in terms of the experiments, but the writing and claims of the paper should be limited around this fact. Additionally, these are speech deepfakes not the broader category of "audio" per-se, so maybe some further adjustment is warranted, though other papers in this subarea tend to use "audio deepfake" to describe speech deepfakes.
As it stands, the examples shown do not convince me that the "attacks" used here are sufficiently high-quality to claim a generalized defense, though the developed method seems to perform well on the datasets used, and the overall scientific study (though limited) is well done.
Questions
What are the systems tested in Table 1? Either by name, or citation? What is the source of the speakers? If these have PII, a description of the speakers broad categorizations is sufficient. If pulled from an existing dataset, speaker ids would be good. As it stands this table is largely uninformative, without any material information beyond a general design motivation for follow-on work (since CCA shows some behavior differences between methods).
Given the importance of both sample rate and noise in audio, it would be very useful to test this approach under those forms of degradation - e.g. does the method scale down to data of narrow bandwidth, at low samplerate or under the presence of additive noise / background sound (such as music, crowd noise, applause, and so on). The prosodic example may hold under reasonable conditions, but how many detections are relying on prosodic features versus simpler acoustic artifacts? Figure 2. hints at this to some extent, but some explicit description and study would be useful.
Generally the data examples shown are extremely noisy, and the synthesis methods are not particularly high quality. Testing on both clean audio, and higher quality synthesis, as well as under controlled degradations could raise my score. After all it is plausible an attacker may use telephony as a transmission channel - especially if the degradations imposed by the channel give the attacker a further advantage.
As a general direction - it may be useful to directly answer some of the questions posed by the titles of the citations in this paper e.g. "Does audio deepfake detection generalize?" - the claim here being "yes", but demonstrations being limited to existing datasets rather than further tests with recently developed technologies / APIs and so on. "Does deepfake detection rely on artifacts?" - the claim here is also (somewhat) "yes", which hurts the counterclaim of being generalized to some degree, unless these artifacts are general across a broad swath of methods, which would be a surprising finding given existing demonstrations.
The primary concern in order to raise my score would be a more proper scoping of the generalization claims, and the domain claims around this method given limitations of the testing datasets. The conclusion also discusses a fair bit about qualitative analyses which are largely relegated to the appendix, so there is further mismatch between the chosen title and the final claim.
Larger and more diverse datasets (multi-lingual being one option, more unusual speaker styles would be another), or more particularly use of a variety of recent, high performing methods would raise my score if the writing is mostly unchanged. Some of these methods may only be available by API, which is unfortunate but perhaps necessary - additionally TortoiseTTS and spinoffs should have specific, stronger synthesis exemplars than those demonstrated, especially under the assumption an attacker may be doing manual selection given a corpus of intermediate generations to choose the best final result.
Limitations
The authors have addressed some limitations of their work, however this review is partly hinging on the gap between claims, and the effective limitations and demonstrated results. More writing on the limitations, and particularly potential harms of deploying unbalanced "defense" methods in terms of accessibility would be beneficial.