QFFN-BERT: An Empirical Study of Depth, Performance, and Data Efficiency in Hybrid Quantum-Classical Transformers

Parameterized quantum circuits (PQCs) have recently emerged as promising components for constructing compact hybrid quantum-classical neural modules. In this work, we introduce QFFN-BERT, a hybrid quantum-classical classifier that integrates a PQC-based feed-forward module into a compact pre-trained BERT model at the [CLS]-representation level for sentence classification. Rather than replacing the full position-wise Transformer feed-forward network (FFN) across all token positions, our study focuses on a computationally feasible [CLS]-level design and empirically characterizes its depth-related trade-offs under simulator-based evaluation. The proposed PQC module uses a residual connection, both $R_{Y}$ and $R_{Z}$ rotations, and an alternating entanglement strategy to improve training stability while maintaining expressive capacity. Experiments on SST-2 and DBpedia show that carefully configured QFFN-BERT variants can remain competitive with the classical bert-tiny baseline while substantially reducing trainable parameters within the inserted feed-forward module. Additional three-seed experiments on key SST-2 settings further indicate that the 4-layer QFFN-BERT maintains a stable full-data advantage, whereas the 8-layer model in the 10% few-shot setting remains competitive but does not consistently outperform the baseline. Together with an ablation study on a non-optimized PQC design, these results suggest that PQC-based modules can serve as compact alternatives for sentence-level hybrid Transformer classifiers, while also highlighting important trade-offs among depth, trainability, and computational cost.

Paper

Similar papers

© 2026 NYSGPT2525 LLC