Enhancing Temporal Action Segmentation with Large Language Models

Temporal action segmentation aims to assign action labels to each frame of an untrimmed video lasting several minutes, making it crucial for fields such as action understanding, evaluation, and skill acquisition support. While existing methods primarily focus on recognizing logical relationships between adjacent actions, they often fail to capture the out-of-context semantics of individual actions, namely the overall procedure, which challenges the accuracy of action label prediction. Although state transition models like Markov models and Recurrent Neural Networks (RNNs) can capture overall procedures, they require designing domain-specific state transition models for effective learning. To address these challenges, we propose a novel approach that leverages large language models (LLMs) as universal state transition models. These models, trained on diverse linguistic and image data on an internet scale, possess extensive general knowledge across various domains. By incorporating commonsense and logical action sequences, our method enhances action recognition accuracy. Evaluations conducted on multiple video datasets, including 50Salads and Breakfast, demonstrate that our proposed method outperforms existing techniques.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC