The task of video object segmentation with referring expressions\n(language-guided VOS) is to, given a linguistic phrase and a video, generate\nbinary masks for the object to which the phrase refers. Our work argues that\nexisting benchmarks used for this task are mainly composed of trivial cases, in\nwhich referents can be identified with simple phrases. Our analysis relies on a\nnew categorization of the phrases in the DAVIS-2017 and Actor-Action datasets\ninto trivial and non-trivial REs, with the non-trivial REs annotated with seven\nRE semantic categories. We leverage this data to analyze the results of RefVOS,\na novel neural network that obtains competitive results for the task of\nlanguage-guided image segmentation and state of the art results for\nlanguage-guided VOS. Our study indicates that the major challenges for the task\nare related to understanding motion and static actions.\n