The Monte Carlo Transformer: a stochastic self-attention model for sequence prediction

This paper introduces the Sequential Monte Carlo Transformer, an original\napproach that naturally captures the observations distribution in a transformer\narchitecture. The keys, queries, values and attention vectors of the network\nare considered as the unobserved stochastic states of its hidden structure.\nThis generative model is such that at each time step the received observation\nis a random function of its past states in a given attention window. In this\ngeneral state-space setting, we use Sequential Monte Carlo methods to\napproximate the posterior distributions of the states given the observations,\nand to estimate the gradient of the log-likelihood. We hence propose a\ngenerative model giving a predictive distribution, instead of a single-point\nestimate.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC