We derive an unbiased estimator for expectations over discrete random\nvariables based on sampling without replacement, which reduces variance as it\navoids duplicate samples. We show that our estimator can be derived as the\nRao-Blackwellization of three different estimators. Combining our estimator\nwith REINFORCE, we obtain a policy gradient estimator and we reduce its\nvariance using a built-in control variate which is obtained without additional\nmodel evaluations. The resulting estimator is closely related to other gradient\nestimators. Experiments with a toy problem, a categorical Variational\nAuto-Encoder and a structured prediction problem show that our estimator is the\nonly estimator that is consistently among the best estimators in both high and\nlow entropy settings.\n