For many low-resource or endangered languages, spoken language resources are\nmore likely to be annotated with translations than with transcriptions. Recent\nwork exploits such annotations to produce speech-to-translation alignments,\nwithout access to any text transcriptions. We investigate whether providing\nsuch information can aid in producing better (mismatched) crowdsourced\ntranscriptions, which in turn could be valuable for training speech recognition\nsystems, and show that they can indeed be beneficial through a small-scale case\nstudy as a proof-of-concept. We also present a simple phonetically aware string\naveraging technique that produces transcriptions of higher quality.\n