Preprint · 2026
Kernel Autoresearch for Open-Ended Model Discovery
TL;DR. A fixed kernel library limits which models a search can discover. Kernaut uses coding agents and validity-preserving contracts to discover reusable kernel programs that transfer to unseen tasks without further LLM calls.
Abstract
Kernels encode the inductive bias of a wide range of machine learning models, yet automated kernel design faces a fundamental dilemma. A fixed grammar of base kernels and operators guarantees validity but limits the search to structures expressible by those building blocks. Conversely, unrestricted programs remove this limitation but no longer guarantee validity. In our stress tests, 22–58% of LLM-generated kernels that pass numerical checks on random inputs fail when evaluated at different scales or dimensions. We propose Kernel Autoresearch (Kernaut), which treats kernel design as open-ended model discovery. Coding agents write kernels as programs, while construction contracts ensure that every accepted kernel is valid. A quality-diversity archive retains high-performing kernels with distinct behaviors, and novelty screening steers agents toward functionally new candidates. Our experiments demonstrate that the discovered kernels encode reusable inductive biases that generalize to unseen tasks. On held-out black-box optimization families, a discovered kernel outperforms a meta-learned deep kernel trained on the same episodes. Furthermore, kernels discovered from ten enzyme-kinetic rate laws achieve lower error than tuned ARD and deep kernel baselines on five unseen mechanisms. The discovered kernels are also interpretable programs that human researchers can refine: a human-refined version of one further reduces the held-out predictive error by 5.7% and optimization regret by 7.8%.
Kernel Autoresearch
A kernel defines which inputs a model treats as similar. Kernaut searches for that similarity rule as a short Python program. The agent writes a component, and trusted backend code assembles the kernel through a construction contract.
The four contracts let agents write new feature maps and input transformations, as well as combine existing kernels:
- Feature map. Map each input to a vector. The interpreter takes the inner product of two vectors.
- Spectral. Propose frequencies and raw weights. The interpreter squares the weights and combines cosine components.
- Pullback. Transform each input, then apply a trusted library kernel to the transformed inputs.
- Closure. Combine trusted or accepted kernels through sums, products, nonnegative scaling, and input transformations.
These constructions preserve positive semidefiniteness (PSD), the matrix property required of a kernel. Kernel validity also requires finite outputs, correct dimensions, and pure components. A pure component depends only on its input and parameters, without batch dependence, randomness, or state from earlier calls.
Each discovery campaign runs a sequence of fresh agent sessions. A session either creates a program or revises an archived program. A MAP-Elites archive retains strong candidates with distinct kernel behavior 1. Archived examples and structured rejection reasons guide later proposals. Numerical fingerprints discourage candidates that behave like known kernels.
Search and parameter tuning use meta-training data only. Validation then selects one frozen program, which transfers to held-out tasks without structural changes or further LLM calls. Each new Gaussian process (GP) fits only covariance amplitude and observation noise.
Why numerical checks are not enough
A covariance matrix can pass a numerical PSD test on one set of inputs and fail on another. In three unrestricted-program cohorts, additional checks falsified 22–58% of the programs that passed the initial screen.
Kernaut checks execution and output shapes at Tier 0, sampled PSD at Tier 1, and construction plus consistency at Tier 2. The validity argument comes from trusted assembly under the contract assumptions. Source inspection and behavioral checks test conformance, but do not prove arbitrary Python programs pure over every input.
Experiments
The experiments test whether one frozen kernel structure transfers beyond the tasks used during search. We report the continuous ranked probability score (CRPS), which evaluates a predictive distribution against the observed value 2. Lower CRPS is better. Optimization uses regret AUC, the mean normalized simple regret over BO steps, where lower is also better.
Black-box optimization
Search and validation use five function families. Testing uses 60 tasks from six other families, with dimensions up to eight. Each predictive episode fits on 8 + 2d observations and scores 32 held-out points, where d is the input dimension. Each discovery campaign has 12 runs.
Validation selected the dual warp–fold (DWF) kernel in the first ensemble campaign. DWF combines a mild coordinate warp with a triangular fold. Its test CRPS was 0.601, compared with 0.612 for fixed Matérn-5/2. Its regret AUC was 0.671, compared with 0.693.
Across seven ensemble campaigns, six selected kernels improved test CRPS over Matérn-5/2. All seven GPT-6 Astra campaigns improved CRPS, with gains from 5.8% to 11.2%. These repeated campaigns measure discovery variability, whereas Figure 3 compares individual frozen kernels.
We also compared DWF with Few-Shot Bayesian Optimization (FSBO), which meta-learns a deep kernel 3. At the matched budget of 518 meta-training points, DWF had lower CRPS and regret AUC in both FSBO evaluation modes. At 64,000 points, neither FSBO mode differed significantly from DWF on either metric.
Does open-ended synthesis matter?
The matched ablation keeps the proposer ensemble, run budget, and evaluation protocol fixed. Restricting agents to compositions of library kernels raises test CRPS from 0.571 to 0.618. Full search wins 11 of 12 paired campaigns against this closure-only arm.
| Search | Test CRPS ↓ | Regret AUC ↓ |
|---|---|---|
| Full archive | 0.571 ± 0.031 | 0.681 |
| Independent proposals | 0.617 ± 0.037 | 0.695 |
| Closure grammar | 0.618 ± 0.021 | 0.706 |
| Greedy compositional search | 0.625 | 0.684 |
Independent proposals also perform worse, but that arm changes feedback, novelty screening across runs, and candidate retention together. The comparison measures their combined contribution, rather than isolating archive sharing.
Forecasting unseen greenhouse-gas records
Search uses monthly CO2, CH4, and N2O records. Testing uses SF6, CFC-12, and CFC-11, with 30 windows per record and a 48-month horizon. Some windows reverse time, so the benchmark includes both forecasting and backcasting.
Validation selected a residual period-bank kernel. It combines a Matérn component, a smooth residual component, and low-frequency sinusoids that act as slow trends. The selected kernel ranks first among 58 frozen candidates on both CFC records.
Across all three records, the selected kernel has mean CRPS 0.091, compared with 0.120 for AutoGP, which infers a structure for each window 4. The overall gain comes mainly from reversed windows. On forward windows, its CRPS is 0.034, behind CAKE at 0.022. Windows overlap within each record, so these are not independent record-level replications.
Transfer to unseen mechanisms and patients
Enzyme kinetics. ChemBench maps seven experimental variables to a reaction rate. Search uses ten rate-law families. Testing uses 75 episodes from five unseen mechanisms, with 20 observations for fitting and 20 for scoring in each episode.
The validation-selected frontier kernel reaches CRPS 0.312, compared with 0.567 for optimized automatic relevance determination (ARD). ARD learns a separate lengthscale for each input. The selected kernel improves over all six learned references in paired tests with Holm correction. The selected ensemble kernel is competitive with the strongest learned references.
The frontier program combines substrate-response shapes inspired by the training mechanisms, local sensitivities, and modulation by enzyme loading. These features provide an inspectable biochemical prior. Predictive transfer does not establish recovery of the true mechanism.
Glucose dynamics. GlucoseBench uses simulated patients. Search uses ten children, validation uses ten adolescents, and testing uses ten adults. Each forecast uses only training episodes from its target patient, so the experiment tests transfer of kernel structure.
With two training episodes, the five discovered kernels reduce geometric-mean normalized MSE by 38–67% against relative-time ARD-Matérn. With six episodes and a 45-minute prefix, the reductions narrow to 0–23%. The selected ensemble kernel encodes a delayed meal response and a fading insulin effect.
These glucose results describe ten simulated adults. Across 35 paired patient-level comparisons, no adjusted p-value falls below 0.05 after Holm correction. The reported percentage reductions therefore describe the observed benchmark means.
From Discovery to Human Refinement
DWF makes the discovered inductive bias visible. Its warp stays near the identity, while its fold bends the input geometry around a peak at 0.513. The original program applies a Matérn kernel jointly to the warped and folded coordinates.
A human refinement gives the warp and fold separate kernels, then averages their covariances. This additive construction preserves PSD and strengthens the reflection preference.
| Construction | Test CRPS ↓ | Regret AUC ↓ |
|---|---|---|
| Archived DWF | 0.6011 | 0.6712 |
| Additive branches, human refinement | 0.5668 | 0.6190 |
The CRPS gain has a paired 95% bootstrap interval excluding zero. The regret reduction remains descriptive because the multiplicity-adjusted test does not cross 0.05. Smoothing the fold preserves the CRPS gain, so the sharp corner is unnecessary for this improvement.
Scope and Open Questions
The experiments test small-data Gaussian processes. The BBO benchmark has at most eight observed dimensions, and the scientific benchmarks use simulated mechanisms or patients. Performance on larger datasets, higher-dimensional tasks, and clinical data remains untested here.
Construction contracts establish validity under stated assumptions. They do not guarantee predictive accuracy, scientific novelty, or causal interpretation. Kernel fingerprints compare behavior on finite probe sets, and dimension-specific code can still fail outside its tested interface.
The short programs make further tests concrete: anonymize scientific variable names to measure the proposer’s use of domain knowledge, or add contracts for state-space and multi-output kernels. The documentation describes the implementation and workflow.
Citation
@misc{suwandi2026kernaut,
title={Kernel Autoresearch for Open-Ended Model Discovery},
author={Suwandi, Richard Cornelius and Yin, Feng and Murphy, Kevin},
year={2026},
url={https://www.alphaxiv.org/abs/2610.kernel-autoresearch},
note={Preprint}
}
References
-
Illuminating Search Spaces by Mapping Elites
Mouret, J.-B. and Clune, J., 2015. arXiv:1504.04909. -
Strictly Proper Scoring Rules, Prediction, and Estimation
Gneiting, T. and Raftery, A. E., 2007. Journal of the American Statistical Association, 102(477), 359–378. -
Few-Shot Bayesian Optimization with Deep Kernel Surrogates
Wistuba, M. and Grabocka, J., 2021. International Conference on Learning Representations. -
Sequential Monte Carlo Learning for Time Series Structure Discovery
Saad, F. A., Patton, B. J., Hoffman, M. D., Saurous, R. A. and Mansinghka, V. K., 2023. Proceedings of the 40th International Conference on Machine Learning, PMLR 202.






