ImagineTrack: Imaging Target Appearance to Prompt Object Tracking

The majority of existing trackers model single object tracking as a one-shot detection task, which means they tackle every frame individually, ignoring the temporal information. Despite its progress in speed, they fail to model appearance variations. Searching target solely based on its historical appearance without considering its current appearance cannot effectively handle scenes with appearance changes. In human vision, we can imagine an object’s future appearance based on the past, which helps us perceive the target at the current moment. Motivated by this, we propose ImagineTrack, which imagines the target’s appearance from historical temporal information to obtain a timely template to prompt the following task. Specifically, our ImagineTrack introduces the concept of spatio-temporal predictive learning to tracking, predicting current and future templates based on historical tracking results. Due to the instability of historical results, we alse retain the initial template, treating this dynamic template as an extended prompt. This approach ensures that the target search is conducted based on both static initial template and dynamic predicted templates. Consequently, ImagineTrack leverages temporal prompts and the interactions among diverse object features, thereby achieving robust tracking performance. Extensive experiments on benchmark datasets (LaSOT, LaSOText, GOT10k, TrackingNet) demonstrate its competitive accuracy with real time efficiency.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC