Fantastic Features and Where to Find Them: Detecting Cognitive Impairment with a Subsequence Classification Guided Approach
Despite the widely reported success of embedding-based machine learning\nmethods on natural language processing tasks, the use of more easily\ninterpreted engineered features remains common in fields such as cognitive\nimpairment (CI) detection. Manually engineering features from noisy text is\ntime and resource consuming, and can potentially result in features that do not\nenhance model performance. To combat this, we describe a new approach to\nfeature engineering that leverages sequential machine learning models and\ndomain knowledge to predict which features help enhance performance. We provide\na concrete example of this method on a standard data set of CI speech and\ndemonstrate that CI classification accuracy improves by 2.3% over a strong\nbaseline when using features produced by this method. This demonstration\nprovides an ex-ample of how this method can be used to assist classification in\nfields where interpretability is important, such as health care.\n