DOCENT: Learning Self-Supervised Entity Representations from Large Document Collections

This paper explores learning rich self-supervised entity representations from\nlarge amounts of the associated text. Once pre-trained, these models become\napplicable to multiple entity-centric tasks such as ranked retrieval, knowledge\nbase completion, question answering, and more. Unlike other methods that\nharvest self-supervision signals based merely on a local context within a\nsentence, we radically expand the notion of context to include any available\ntext related to an entity. This enables a new class of powerful, high-capacity\nrepresentations that can ultimately distill much of the useful information\nabout an entity from multiple text sources, without any human supervision.\n We present several training strategies that, unlike prior approaches, learn\nto jointly predict words and entities -- strategies we compare experimentally\non downstream tasks in the TV-Movies domain, such as MovieLens tag prediction\nfrom user reviews and natural language movie search. As evidenced by results,\nour models match or outperform competitive baselines, sometimes with little or\nno fine-tuning, and can scale to very large corpora.\n Finally, we make our datasets and pre-trained models publicly available. This\nincludes Reviews2Movielens (see https://goo.gle/research-docent ), mapping the\nup to 1B word corpus of Amazon movie reviews (He and McAuley, 2016) to\nMovieLens tags (Harper and Konstan, 2016), as well as Reddit Movie Suggestions\n(see https://urikz.github.io/docent ) with natural language queries and\ncorresponding community recommendations.\n

Paper

References (46)

Scroll for more · 34 remaining

Similar papers

© 2026 NYSGPT2525 LLC