Authorship Attribution of Source Code: A Language-Agnostic Approach and\n Applicability in Software Engineering
Authorship attribution (i.e., determining who is the author of a piece of\nsource code) is an established research topic. State-of-the-art results for the\nauthorship attribution problem look promising for the software engineering\nfield, where they could be applied to detect plagiarized code and prevent legal\nissues. With this article, we first introduce a new language-agnostic approach\nto authorship attribution of source code. Then, we discuss limitations of\nexisting synthetic datasets for authorship attribution, and propose a data\ncollection approach that delivers datasets that better reflect aspects\nimportant for potential practical use in software engineering. Finally, we\ndemonstrate that high accuracy of authorship attribution models on existing\ndatasets drastically drops when they are evaluated on more realistic data. We\noutline next steps for the design and evaluation of authorship attribution\nmodels that could bring the research efforts closer to practical use for\nsoftware engineering.\n