Action recognition and detection in the context of long untrimmed video\nsequences has seen an increased attention from the research community. However,\nannotation of complex activities is usually time consuming and challenging in\npractice. Therefore, recent works started to tackle the problem of unsupervised\nlearning of sub-actions in complex activities. This paper proposes a novel\napproach for unsupervised sub-action learning in complex activities. The\nproposed method maps both visual and temporal representations to a latent space\nwhere the sub-actions are learnt discriminatively in an end-to-end fashion. To\nthis end, we propose to learn sub-actions as latent concepts and a novel\ndiscriminative latent concept learning (DLCL) module aids in learning\nsub-actions. The proposed DLCL module lends on the idea of latent concepts to\nlearn compact representations in the latent embedding space in an unsupervised\nway. The result is a set of latent vectors that can be interpreted as cluster\ncenters in the embedding space. The latent space itself is formed by a joint\nvisual and temporal embedding capturing the visual similarity and temporal\nordering of the data. Our joint learning with discriminative latent concept\nmodule is novel which eliminates the need for explicit clustering. We validate\nour approach on three benchmark datasets and show that the proposed combination\nof visual-temporal embedding and discriminative latent concepts allow to learn\nrobust action representations in an unsupervised setting.\n