It is important for socially assistive robots to be able to recognize when a\nuser needs and wants help. Such robots need to be able to recognize human needs\nin a real-time manner so that they can provide timely assistance. We propose an\narchitecture that uses social cues to determine when a robot should provide\nassistance. Based on a multimodal fusion approach upon eye gaze and language\nmodalities, our architecture is trained and evaluated on data collected in a\nrobot-assisted Lego building task. By focusing on social cues, our architecture\nhas minimal dependencies on the specifics of a given task, enabling it to be\napplied in many different contexts. Enabling a social robot to recognize a\nuser's needs through social cues can help it to adapt to user behaviors and\npreferences, which in turn will lead to improved user experiences.\n