Authorial clustering involves the grouping of documents written by the same\nauthor or team of authors without any prior positive examples of an author's\nwriting style or thematic preferences. For authorial clustering on shorter\ntexts (paragraph-length texts that are typically shorter than conventional\ndocuments), the document representation is particularly important: very\nhigh-dimensional feature spaces lead to data sparsity and suffer from serious\nconsequences like the curse of dimensionality, while feature selection may lead\nto information loss. We propose a high-level framework which utilizes a compact\ndata representation in a latent feature space derived with non-parametric topic\nmodeling. Authorial clusters are identified thereafter in two scenarios: (a)\nfully unsupervised and (b) semi-supervised where a small number of shorter\ntexts are known to belong to the same author (must-link constraints) or not\n(cannot-link constraints). We report on experiments with 120 collections in\nthree languages and two genres and show that the topic-based latent feature\nspace provides a promising level of performance while reducing the\ndimensionality by a factor of 1500 compared to state-of-the-arts. We also\ndemonstrate that, while prior knowledge on the precise number of authors (i.e.\nauthorial clusters) does not contribute much to additional quality, little\nknowledge on constraints in authorial clusters memberships leads to clear\nperformance improvements in front of this difficult task. Thorough\nexperimentation with standard metrics indicates that there still remains an\nample room for improvement for authorial clustering, especially with shorter\ntexts\n