Ainu is an unwritten language that has been spoken by Ainu people who are one\nof the ethnic groups in Japan. It is recognized as critically endangered by\nUNESCO and archiving and documentation of its language heritage is of paramount\nimportance. Although a considerable amount of voice recordings of Ainu folklore\nhas been produced and accumulated to save their culture, only a quite limited\nparts of them are transcribed so far. Thus, we started a project of automatic\nspeech recognition (ASR) for the Ainu language in order to contribute to the\ndevelopment of annotated language archives. In this paper, we report speech\ncorpus development and the structure and performance of end-to-end ASR for\nAinu. We investigated four modeling units (phone, syllable, word piece, and\nword) and found that the syllable-based model performed best in terms of both\nword and phone recognition accuracy, which were about 60% and over 85%\nrespectively in speaker-open condition. Furthermore, word and phone accuracy of\n80% and 90% has been achieved in a speaker-closed setting. We also found out\nthat a multilingual ASR training with additional speech corpora of English and\nJapanese further improves the speaker-open test accuracy.\n