Papers¶
Canonical method papers, grouped by problem setting. Each section is a short reading list, not an archive. Newest year first, same year by title. Foundations is oldest first. Surveys leads with the two standard intros, then newest first.
Surveys and Tutorials¶
- Taking the Human Out of the Loop: A Review of Bayesian Optimization - Shahriari, Swersky, Wang, Adams, and de Freitas, Proceedings of the IEEE, 2016. The standard survey of methods and applications.
- A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning - Brochu, Cora, and de Freitas, 2010. Early tutorial that still reads well.
- Active Learning and Bayesian Optimization: A Unified Perspective to Learn with a Goal - Di Fiore, Nardelli, and Mainini, Archives of Computational Methods in Engineering, 2024. Survey connecting BO and active learning.
- Recent Advances in Bayesian Optimization - Wang, Jin, Schmitt, and Olhofer, ACM Computing Surveys, 2023. Broad survey of methods through about 2022.
- Recent Advances in Bayesian Optimization (slides) - AAAI 2023 tutorial slides (Deshwal, Belakaria, and Doppa).
Foundations¶
- On Bayesian Methods for Seeking the Extremum - Mockus, 1975. The original expected-improvement argument.
- Efficient Global Optimization of Expensive Black-Box Functions - Jones, Schonlau, and Welch, 1998. EGO and expected improvement.
- Gaussian Processes for Global Optimization - Osborne, Garnett, and Roberts, LION 2009. GP surrogates for global optimization.
- Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design - Srinivas, Krause, Kakade, and Seeger, ICML 2010. GP-UCB and no-regret bounds.
- Convergence Rates of Efficient Global Optimization Algorithms - Bull, JMLR 2011. When EGO converges, and at what rate.
- Practical Bayesian Optimization of Machine Learning Algorithms - Snoek, Larochelle, and Adams, NeurIPS 2012. MCMC GPs and EI for hyperparameter tuning.
- Scalable Bayesian Optimization Using Deep Neural Networks - Snoek et al., ICML 2015. DNGO: neural-net surrogates when GPs get expensive.
- GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration - Gardner, Pleiss, Bindel, Weinberger, and Wilson, NeurIPS 2018. The GP engine under BoTorch.
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization - Balandat et al., NeurIPS 2020. The library paper behind most current PyTorch BO.
Surrogate Design¶
Kernels, input transforms, and non-GP surrogates. DNGO is under Foundations.
- Adaptive Kernel Design for Bayesian Optimization Is a Piece of CAKE with LLMs - Suwandi et al., NeurIPS 2025. LLM-driven evolution of GP kernels during BO.
- A Study of Bayesian Neural Network Surrogates for Bayesian Optimization - Li, Rudner, and Wilson, ICLR 2024. When BNNs help as BO surrogates, and when they do not.
- Bayesian Optimization with Conformal Prediction Sets - Stanton, Maddox, and Wilson, AISTATS 2023. Distribution-free uncertainty in the loop.
- Bayesian Optimization with Informative Covariance - Eduardo and Gutmann, TMLR 2023. Encode known structure in the kernel.
- Kernel Identification Through Transformers - Simpson et al., NeurIPS 2021. KITT: recommend a kernel from data in one forward pass.
- Differentiable Compositional Kernel Learning for Gaussian Processes - Sun, Zhang, Wang, Zeng, Li, and Grosse, ICML 2018. Neural kernel networks.
- Bayesian Optimization with Robust Bayesian Neural Networks - Springenberg, Klein, Falkner, and Hutter, NeurIPS 2016. BOHAMIANN: MCMC Bayesian neural nets as surrogates.
- Deep Kernel Learning - Wilson, Hu, Salakhutdinov, and Xing, AISTATS 2016. A deep net as the feature map of a GP kernel.
- Input Warping for Bayesian Optimization of Non-Stationary Functions - Snoek, Swersky, Zemel, and Adams, ICML 2014. Learn a warping so a stationary kernel fits better.
- Gaussian Process Kernels for Pattern Discovery and Extrapolation - Wilson and Adams, ICML 2013. Spectral mixture kernels.
- Structure Discovery in Nonparametric Regression through Compositional Kernel Search - Duvenaud, Lloyd, Grosse, Tenenbaum, and Ghahramani, ICML 2013. Grammar over sums and products of kernels.
Acquisition Functions¶
- Information-Theoretic Bayesian Optimization for Bilevel Optimization Problems - Kanayama, Ito, Tamura, and Karasuyama, UAI 2026. Information gain on both levels of a nested black-box problem.
- FunBO: Discovering Acquisition Functions for Bayesian Optimization with FunSearch - Aglietti et al., ICML 2025. LLM search over acquisition functions written as code.
- Unexpected Improvements to Expected Improvement for Bayesian Optimization - Ament, Daulton, Eriksson, Balandat, and Bakshy, NeurIPS 2023. LogEI, a numerically stable EI.
- Joint Entropy Search for Maximally-Informed Bayesian Optimization - Hvarfner, Hutter, and Nardi, NeurIPS 2022. Information gain on the joint optimum and optimal value.
- πBO: Augmenting Acquisition Functions with User Beliefs for Bayesian Optimization - Hvarfner, Stoll, Souza, Lindauer, Hutter, and Nardi, ICLR 2022. Weight the acquisition by a user prior over the optimum.
- Why Non-myopic Bayesian Optimization is Promising and How Far Should We Look-ahead? A Study via Rollout - Yue and AL Kontar, AISTATS 2020. How much lookahead actually helps.
- Maximizing Acquisition Functions for Bayesian Optimization - Wilson, Hutter, and Deisenroth, NeurIPS 2018. Monte Carlo acquisition via autodiff.
- Parallelised Bayesian Optimisation via Thompson Sampling - Kandasamy et al., AISTATS 2018. TS as a simple batch acquisition.
- Max-value Entropy Search for Efficient Bayesian Optimization - Wang and Jegelka, ICML 2017. Information about the maximum value rather than its location.
- GLASSES: Relieving The Myopia Of Bayesian Optimisation - González, Osborne, and Lawrence, AISTATS 2016. Non-myopic planning.
- Predictive Entropy Search for Efficient Global Optimization of Black-box Functions - Hernández-Lobato, Hoffman, and Ghahramani, NeurIPS 2014. Tractable information-theoretic acquisition.
- Entropy Search for Information-Efficient Global Optimization - Hennig and Schuler, JMLR 2012. Select queries that reduce entropy of the argmax.
- The Knowledge-Gradient Policy for Correlated Normal Beliefs - Frazier, Powell, and Dayanik, INFORMS JOC 2009. Value of information when beliefs are correlated.
High-Dimensional¶
- NeST-BO: Fast Local Bayesian Optimization via Newton-Step Targeting of Gradient and Hessian Information - Tang, Kudva, and Paulson, AISTATS 2026. Local BO that targets a Newton step.
- Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization - Xu, Wang, Phillips, and Zhe, ICLR 2025. Lengthscale init, not a fancy surrogate, is often the failure mode.
- Understanding High-Dimensional Bayesian Optimization - Papenmeier, Poloczek, and Nardi, ICML 2025. Why vanilla GP-BO fails in high-d, and a simple MLE fix.
- Vanilla Bayesian Optimization Performs Great in High Dimensions - Hvarfner, Hellsten, and Nardi, ICML 2024. Careful GP priors often beat specialized high-d methods.
- Sparse Bayesian Optimization - Liu et al., AISTATS 2023. Sparsity in the input for high-d problems.
- The Behavior and Convergence of Local Bayesian Optimization - Wu, Kim, Garnett, and Gardner, NeurIPS 2023. When local BO converges, and when it does not.
- Increasing the Scope as You Learn: Adaptive Bayesian Optimization in Nested Subspaces - Papenmeier, Nardi, and Poloczek, NeurIPS 2022. BAxUS: grow the embedding as budget grows.
- Local Bayesian Optimization via Maximizing Probability of Descent - Nguyen, Wu, Gardner, and Garnett, NeurIPS 2022. Local BO by following estimated descent.
- High-dimensional Bayesian Optimization with Sparse Axis-Aligned Subspaces - Eriksson and Jankowiak, UAI 2021. SAASBO: sparsity-inducing priors on lengthscales.
- Re-Examining Linear Embeddings for High-Dimensional Bayesian Optimization - Letham, Calandra, Rai, and Bakshy, NeurIPS 2020. ALEBO.
- A Framework for Bayesian Optimization in Embedded Subspaces - Nayebi, Munteanu, and Poloczek, ICML 2019. HeSBO, hashing into a low-d embedding.
- Scalable Global Optimization via Local Bayesian Optimization - Eriksson, Pearce, Gardner, Turner, and Poloczek, NeurIPS 2019. TuRBO: trust-region local BO.
- Optimization, Fast and Slow: Optimally Switching between Local and Bayesian Optimization - McLeod, Roberts, and Osborne, ICML 2018. Switch between local search and BO.
- Discovering and Exploiting Additive Structure for Bayesian Optimization - Gardner, Guo, Weinberger, Garnett, and Grosse, AISTATS 2017. Learn which additive decomposition to use.
- Bayesian Optimization in a Billion Dimensions via Random Embeddings - Wang, Hutter, Zoghi, Matheson, and de Freitas, JAIR 2016. REMBO: optimize in a random low-dimensional subspace.
- High Dimensional Bayesian Optimisation and Bandits via Additive Models - Kandasamy, Schneider, and Póczos, ICML 2015. Additive GP structure.
Constrained and Safe¶
- Maximally Robust Satisficing Bayesian Optimization - Kinnunen, Mikkola, Niskanen, and Klami, UAI 2026. Satisficing solutions that stay good under large input perturbations.
- Scalable Constrained Bayesian Optimization - Eriksson and Poloczek, AISTATS 2021. SCBO: TuRBO with constraints.
- A General Framework for Constrained Bayesian Optimization using Information-based Search - Hernández-Lobato et al., JMLR 2016. PESC.
- Safe Exploration for Optimization with Gaussian Processes - Sui, Gotovos, Burdick, and Krause, ICML 2015. SafeOpt.
- Bayesian Optimization with Inequality Constraints - Gardner, Kusner, Xu, Weinberger, and Cunningham, ICML 2014. Constrained EI.
- Bayesian Optimization with Unknown Constraints - Gelbart, Snoek, and Adams, UAI 2014. Constraints that are themselves expensive black boxes.
Multi-Objective¶
- Multi-Objective Bayesian Optimization over High-Dimensional Search Spaces - Daulton, Eriksson, Balandat, and Bakshy, UAI 2022. MORBO: local trust regions for high-d multi-objective BO.
- Parallel Bayesian Optimization of Multiple Noisy Objectives with Expected Hypervolume Improvement - Daulton, Balandat, and Bakshy, NeurIPS 2021. NEHVI under noise.
- Differentiable Expected Hypervolume Improvement for Parallel Multi-Objective Bayesian Optimization - Daulton, Balandat, and Bakshy, NeurIPS 2020. qEHVI for parallel MOBO.
- Efficient Computation of Expected Hypervolume Improvement Using Box Decomposition Algorithms - Yang, Emmerich, Deutz, and Bäck, JOGO 2019. Fast EHVI.
- Predictive Entropy Search for Multi-objective Bayesian Optimization - Hernández-Lobato, Hernández-Lobato, Shah, and Adams, ICML 2016. PESMO: information gain on the Pareto set.
- ParEGO: A Hybrid Algorithm with On-line Landscape Approximation for Expensive Multiobjective Optimization Problems - Knowles, IEEE TEVC 2006. Scalarization plus EGO.
Multi-Fidelity, Multi-Task, and Transfer¶
- Pre-trained Gaussian Processes for Bayesian Optimization - Wang et al., JMLR 2024. HyperBO: transfer a GP prior from related tasks.
- Few-Shot Bayesian Optimization with Deep Kernel Surrogates - Wistuba and Grabocka, ICLR 2021. Deep kernels for few-shot HPO.
- Multi-fidelity Bayesian Optimization with Max-value Entropy Search and its Parallelization - Takeno, Fukuoka, Tsukada, Koyama, Shiga, Takeuchi, and Karasuyama, ICML 2020. Information-theoretic multi-fidelity acquisition.
- BOHB: Robust and Efficient Hyperparameter Optimization at Scale - Falkner, Klein, and Hutter, ICML 2018. TPE plus Hyperband.
- Fast Bayesian Optimization of Machine Learning Hyperparameters on Large Datasets - Klein, Falkner, Bartels, Hennig, and Hutter, AISTATS 2017. FABOLAS.
- Multi-fidelity Bayesian Optimisation with Continuous Approximations - Kandasamy, Dasarathy, Schneider, and Póczos, ICML 2017. Continuous fidelity rather than a discrete ladder.
- Gaussian Process Bandit Optimisation with Multi-fidelity Evaluations - Kandasamy, Dasarathy, Oliva, Schneider, and Póczos, NeurIPS 2016. MF-GP-UCB.
- Freeze-Thaw Bayesian Optimization - Swersky, Snoek, and Adams, 2014. Pause and resume training runs using a GP over learning curves.
- Multi-Task Bayesian Optimization - Swersky, Snoek, and Adams, NeurIPS 2013. Share data across related tasks.
- Global Optimization of Stochastic Black-Box Systems via Sequential Kriging Meta-Models - Huang, Allen, Notz, and Zeng, JOGO 2006. Early multi-fidelity EGO.
Batch and Parallel¶
- GIBBON: General-purpose Information-Based Bayesian Optimisation - Moss, Leslie, Gonzalez, and Rayson, JMLR 2021. Cheap batch information-theoretic acquisition.
- Batched Large-scale Bayesian Optimization in High-dimensional Spaces - Wang, Gehring, Kohli, and Jegelka, AISTATS 2018. Ensemble batch BO in high-d.
- Batch Bayesian Optimization via Local Penalization - González, Dai, Hennig, and Lawrence, AISTATS 2016. Penalize around pending points.
- Batched Gaussian Process Bandit Optimization via Determinantal Point Processes - Kathuria, Deshpande, and Kohli, NeurIPS 2016. Diverse batches via DPPs.
- The Parallel Knowledge Gradient Method for Batch Bayesian Optimization - Wu and Frazier, NeurIPS 2016. Parallel KG.
- Parallelizing Exploration-Exploitation Tradeoffs in Gaussian Process Bandit Optimization - Desautels, Krause, and Burdick, JMLR 2014. GP-BUCB.
- Parallel Gaussian Process Optimization with Upper Confidence Bound and Pure Exploration - Contal, Buffoni, Robicquet, and Vayatis, ECML 2013. GP-UCB-PE.
Discrete and Mixed Spaces¶
- Bounce: Reliable High-Dimensional Bayesian Optimization for Combinatorial and Mixed Spaces - Papenmeier, Nardi, and Poloczek, NeurIPS 2023. Nested embeddings plus trust regions.
- Think Global and Act Local: Bayesian Optimisation over High-Dimensional Categorical and Mixed Search Spaces - Wan, Nguyen, Ha, Ru, Lu, and Osborne, ICML 2021. High-dimensional categorical and mixed spaces.
- Bayesian Optimisation over Multiple Continuous and Categorical Inputs - Ru, Alvi, Nguyen, Osborne, and Roberts, ICML 2020. CoCaBO.
- Combinatorial Bayesian Optimization using the Graph Cartesian Product - Oh, Tomczak, Gavves, and Welling, NeurIPS 2019. COMBO.
- Bayesian Optimization of Combinatorial Structures - Baptista and Poloczek, ICML 2018. BOCS.
Preferential Feedback¶
- qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization - Astudillo, Lin, Bakshy, and Frazier, AISTATS 2023. Expected utility of the best option.
- Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes - Lin, Astudillo, Frazier, and Bakshy, AISTATS 2022. Learn a utility from pairwise comparisons, then optimize it.
- Preferential Bayesian Optimization - González, Dai, Damianou, and Lawrence, ICML 2017. Optimize from pairwise comparisons rather than numeric scores.
Meta-Learning¶
- PFNs4BO: In-Context Learning for Bayesian Optimization - Müller, Feurer, Hollmann, and Hutter, ICML 2023. Prior-fitted networks as BO surrogates.
- Initializing Bayesian Hyperparameter Optimization via Meta-Learning - Feurer, Springenberg, and Hutter, AAAI 2015. Warm-start the BO prior from related tasks.
LLMs and BO¶
The LLM is the optimizer (no GP), or it occupies one slot in a BO loop. CAKE is under Surrogate Design. FunBO is under Acquisition Functions. PlugBO is under Software and Blogs.
- Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch - Brunzema et al., 2026. An LLM agent runs the loop; a BoTorch backend holds the posterior.
- LLINBO: Trustworthy LLM-in-the-Loop Bayesian Optimization - Chang, Azvar, Okwudire, and Al Kontar, 2025. Keep a GP in the loop so LLM proposals stay uncertainty-aware.
- Reasoning BO: Enhancing Bayesian Optimization with Long-Context Reasoning Power of LLMs - Yang et al., 2025. Reasoning models and a knowledge graph to guide BO sampling.
- Large Language Models as Optimizers - Yang et al., ICLR 2024. OPRO: the LLM proposes candidates from a textual history, with no GP posterior.
- Large Language Models to Enhance Bayesian Optimization - Liu, Astorga, Seedat, and van der Schaar, ICLR 2024. LLAMBO: LLM warm-start, surrogate, and sampler inside a BO loop.