Summary
This work focuses on
Strengths
Originality: Related work in Chen et al., IEEE SMC 2022 uses a language model for pseudo label corrections during BCI self-recalibration, though the study is focused on simulations with EEG data from a longitudinal study with participants with ALS using the P300 speller. This work is a creative combination of existing ideas to develop a new method to recalibrate BCIs for communication by using language models to improve pseudolabel quality and enhancing continuous learning during recalibration via the use of a replay buffer and data augmentation.
Quality: This paper presents results from a longitudinal online BCI study to demonstrate utility of the proposed approach, which is the gold standard in evaluating BCI algorithms. The inclusion of results from offline analysis also enhances the paper.
Clarity: The paper is very well-written and organised. Areas needing clarity and suggestions to improve readability are noted below.
Significance: This work is highly relevant to developing automated approaches to periodically recalibrate BCIs for communication for long-term BCI use with minimal user disruptions. The approach is applicable to general BCIs for communication. Results from a longitudinal online study with a BCI user from a target end user population increase the impact of the paper.
Weaknesses
- Low sample size. The paper presents results from one participant with generally high performance level, so difficult to assess the utility across a broad range of user performance levels. The low sample size is understandable given the challenge with conducting studies in target BCI end-user populations; in particular, this is an iBCI study, in contrast to a non-invasive BCI study. The authors recognise the limitation of the lack of generalisability of results given the low sample size. The authors include results from simulations using data from the current participant to investigate the impact of a broad range of character error rates on the recalibration performance (Figure 5).
- Potential order effects due to lack of randomisation of the no recalibration block (block 2) and the recalibration blocks (blocks 3 and 4). If understood correctly, the RNN decoder is updated with the data from the current calibration block and does not rely on data from the seed model block, so the testing order could be randomised daily to mitigate order effects.
- There is the confound of the recalibration blocks (blocks 3 and 4) displaying the LM-decoded outputs (“the top-scored result was displayed on the screen as the final decoded sentence.”) vs. the no recalibration blocks (block 2) displaying the RNN-decoded outputs. This difference in feedback may potentially impact the BCI user experience (mental state, motivation, etc.) and further compound order effects as the user is aware given the fixed block order.
Questions
- “The second block employed a frozen seed model, trained on a combination of data from [46] and data collected prior to this evaluation (21 sessions in total).” Are the “data collected prior to this evaluation” from the current participant? Also referred to as “newly collected data” earlier in the paper. Are these 21 sessions prior to day 0?
- “updating the decoder after every sentence.” How is the end of a sentence detected? Automatically?
- Equation 1: \theta_k, where k denotes a day implies that the recalibration uses all the data from that day. Is Equation 1 supposed to be \theta_{x, k}, where x refers to sentence?
- What are the character error rates of the LM-decoded outputs (Figure 2b)? This is to assess whether the use of LM-based correction at word level introduces errors at the character level (vs. Figure 2a with RNN-decoder outputs). (Can be inferred based on simulations in figure 5).
- Figure 2: It would be useful to include results from recalibration with the ground truth labels. Why are the amounts of data collection different on day 0 and day 105 different? If understood correctly, day 0 does not have four blocks. This needs to be specified/clarified in the text/caption.
- Inconsistency: Over an approximately 8-month period, our participant used the iBCI system monthly and wrote on average 57.7 sentences per usage session.” vs. “X’s writing speed was 69.5 ± 8.6 characters per minute on average.”
- What is “per-frame labeling”?
- x_{i, t}, y_{i, t}: define subscript i.
- Define all acronyms and variables in the captions and provide more context such that the captions standalone to understand the content of the presented information without necessarily referencing the text. Figure and table captions should be more informative to minimise confusion/misinterpreting the CER% or WER% results across figures/tables. For example, the mismatch between the average online WER % with CORP in table 1 vs. figure 2 is explained in the text and not the figure caption. Same with Figure 3. Captions should state if results are from offline vs. online analysis, specific blocks used during recalibration, etc., for clarity.
- Check that the contrast between line styles is preserved when figures are in grayscale.
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
- More discussion is needed on the societal impact of the BCI technology. In particular, how effectively the BCI communicates a user’s intent when there is no alternative. There is the potential concern that the LM may be more dominant than the user’s intent, particularly in cases with low BCI prediction accuracy.
- “we do not anticipate pseudolabel quality to be a major concern in practice. This is because future clinically viable iBCIs are expected to have a high decoding accuracy... Users are also likely to utilize the iBCI frequently, resulting in small nonstationarities most of the time. ... we believe that the pseudo-labels will have high accuracy, allowing CORP to sustain the iBCIs accuracy indefinitely.” Given the low sample sizes and no current data from “future clinically viable iBCIs”, these claims are questionable. There are issues related to recording quality with long term use of intracortical electrodes.