No-Frills Human-Object Interaction Detection: Factorization, Layout Encodings, and Training Techniques

We show that for human-object interaction detection a relatively simple\nfactorized model with appearance and layout encodings constructed from\npre-trained object detectors outperforms more sophisticated approaches. Our\nmodel includes factors for detection scores, human and object appearance, and\ncoarse (box-pair configuration) and optionally fine-grained layout (human\npose). We also develop training techniques that improve learning efficiency by:\n(1) eliminating a train-inference mismatch; (2) rejecting easy negatives during\nmini-batch training; and (3) using a ratio of negatives to positives that is\ntwo orders of magnitude larger than existing approaches. We conduct a thorough\nablation study to understand the importance of different factors and training\ntechniques using the challenging HICO-Det dataset.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC