The recent phenomenal success of language models has reinvigorated machine\nlearning research, and large sequence models such as transformers are being\napplied to a variety of domains. One important problem class that has remained\nrelatively elusive however is purposeful adaptive behavior. Currently there is\na common perception that sequence models "lack the understanding of the cause\nand effect of their actions" leading them to draw incorrect inferences due to\nauto-suggestive delusions. In this report we explain where this mismatch\noriginates, and show that it can be resolved by treating actions as causal\ninterventions. Finally, we show that in supervised learning, one can teach a\nsystem to condition or intervene on data by training with factual and\ncounterfactual error signals respectively.\n