Skip to content

Kernel Autoresearch for Open-Ended Model Discovery

Kernels encode the inductive biases of a wide range of machine learning models, and the choice of kernel largely determines what a model can learn from limited data.

Kernel Autoresearch (Kernaut) treats kernel design as open-ended program synthesis. Coding agents write kernels as programs, and construction contracts ensure that every accepted kernel is valid. A quality-diversity archive keeps strong kernels with distinct behaviors, and meta-evaluators test whether the discoveries generalize to tasks that the search never saw.

The name combines kernel and astronaut: Kernaut explores unfamiliar spaces of kernels, much as an astronaut navigates unknown territory.

Installation and examples Meta-evaluation protocol

Kernel design as model discovery

A fixed library of kernels limits which structures automated search can express. Kernaut searches over programs instead, and it keeps explicit rules for constructing positive semidefinite kernels. These kernels apply to any method that needs one. Examples include Gaussian process surrogates for Bayesian optimization (Wistuba and Grabocka, 2021), scientific modeling in chemistry (Griffiths et al., 2023) and for differential equations (Chen et al., 2021), and kernel-based uncertainty estimates for language models (Nikitin et al., 2024).

The supplied meta-evaluators use Gaussian processes. To use another kernel method, provide its fitting and scoring rules through a meta-evaluator. The meta-evaluation protocol tests whether discovered inductive biases transfer to unseen tasks. It follows DiscoGen (Goldie et al., 2026).

How Kernaut works

Agents propose kernel components, a trusted backend verifies their construction, and evaluation guides further proposals. A frozen kernel is then tested on tasks the search never saw. Read how Kernaut works for the framework overview, or follow the DWF example to see a discovered kernel.

Benchmarks and tasks

Benchmark or task Purpose
Black-box optimization Evaluate prediction and optimization on transformed analytic functions
Greenhouse-gas forecasting Evaluate transfer from CO2, CH4, and N2O to the held-out records SF6, CFC-12, and CFC-11
ChemBench enzyme kinetics Evaluate transfer from canonical rate-law mechanisms to distinct held-out mechanisms
GlucoseBench forecasting Evaluate kernel transfer across simulated patient groups
User-supplied regression Search on observations supplied as an input matrix and target vector
Offline sine-wave example Check the complete task and verification workflow without a model provider

The benchmark guide describes the data, metrics, and available commands. The meta-evaluation protocol specifies which tasks are available during search, selection, and final evaluation.

Methods and extensions

Use custom regression tasks for a single dataset, or define a custom benchmark for evaluation across task families. Baseline extensions support comparisons under a shared evaluator. Model adapters connect additional providers to the search interface.

The archive visualizer supports inspection during a search and analysis of completed runs. The contribution guide specifies the requirements for new tasks, reference methods, and integrations.

Research publication

Kernel Autoresearch for Open-Ended Model Discovery
Richard Cornelius Suwandi, Feng Yin, and Kevin Murphy.

Read the paper, see the project citation, related references, and verification guarantees.