A Hierarchical Approach to Scaling Batch Active Search Over Structured Data

Active search is the process of identifying high-value data points in a large\nand often high-dimensional parameter space that can be expensive to evaluate.\nTraditional active search techniques like Bayesian optimization trade off\nexploration and exploitation over consecutive evaluations, and have\nhistorically focused on single or small (<5) numbers of examples evaluated per\nround. As modern data sets grow, so does the need to scale active search to\nlarge data sets and batch sizes. In this paper, we present a general\nhierarchical framework based on bandit algorithms to scale active search to\nlarge batch sizes by maximizing information derived from the unique structure\nof each dataset. Our hierarchical framework, Hierarchical Batch Bandit Search\n(HBBS), strategically distributes batch selection across a learned embedding\nspace by facilitating wide exploration of different structural elements within\na dataset. We focus our application of HBBS on modern biology, where large\nbatch experimentation is often fundamental to the research process, and\ndemonstrate batch design of biological sequences (protein and DNA). We also\npresent a new Gym environment to easily simulate diverse biological sequences\nand to enable more comprehensive evaluation of active search methods across\nheterogeneous data sets. The HBBS framework improves upon standard performance,\nwall-clock, and scalability benchmarks for batch search by using a broad\nexploration strategy across coarse partitions and fine-grained exploitation\nwithin each partition of structured data.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC