Preprint · 2026

Kernel Autoresearch for Open-Ended Model Discovery

TL;DR. A fixed kernel library limits which models a search can discover. Kernaut uses coding agents and validity-preserving contracts to discover reusable kernel programs that transfer to unseen tasks without further LLM calls.

Abstract

Kernels encode the inductive bias of a wide range of machine learning models, yet automated kernel design faces a fundamental dilemma. A fixed grammar of base kernels and operators guarantees validity but limits the search to structures expressible by those building blocks. Conversely, unrestricted programs remove this limitation but no longer guarantee validity. In our stress tests, 22–58% of LLM-generated kernels that pass numerical checks on random inputs fail when evaluated at different scales or dimensions. We propose Kernel Autoresearch (Kernaut), which treats kernel design as open-ended model discovery. Coding agents write kernels as programs, while construction contracts ensure that every accepted kernel is valid. A quality-diversity archive retains high-performing kernels with distinct behaviors, and novelty screening steers agents toward functionally new candidates. Our experiments demonstrate that the discovered kernels encode reusable inductive biases that generalize to unseen tasks. On held-out black-box optimization families, a discovered kernel outperforms a meta-learned deep kernel trained on the same episodes. Furthermore, kernels discovered from ten enzyme-kinetic rate laws achieve lower error than tuned ARD and deep kernel baselines on five unseen mechanisms. The discovered kernels are also interpretable programs that human researchers can refine: a human-refined version of one further reduces the held-out predictive error by 5.7% and optimization regret by 7.8%.

Kernel Autoresearch

A kernel defines which inputs a model treats as similar. Kernaut searches for that similarity rule as a short Python program. The agent writes a component, and trusted backend code assembles the kernel through a construction contract.

Kernaut pipeline: a coding agent proposes components, the backend checks and tunes them, and accepted kernels enter an archive that guides later proposals.
Figure 1. The agent proposes, revises, and interprets kernel components. The backend executes, checks, tunes, and scores each proposal. Only programs that pass Tier 2 enter the archive.

The four contracts let agents write new feature maps and input transformations, as well as combine existing kernels:

  • Feature map. Map each input to a vector. The interpreter takes the inner product of two vectors.
  • Spectral. Propose frequencies and raw weights. The interpreter squares the weights and combines cosine components.
  • Pullback. Transform each input, then apply a trusted library kernel to the transformed inputs.
  • Closure. Combine trusted or accepted kernels through sums, products, nonnegative scaling, and input transformations.

These constructions preserve positive semidefiniteness (PSD), the matrix property required of a kernel. Kernel validity also requires finite outputs, correct dimensions, and pure components. A pure component depends only on its input and parameters, without batch dependence, randomness, or state from earlier calls.

Each discovery campaign runs a sequence of fresh agent sessions. A session either creates a program or revises an archived program. A MAP-Elites archive retains strong candidates with distinct kernel behavior 1. Archived examples and structured rejection reasons guide later proposals. Numerical fingerprints discourage candidates that behave like known kernels.

Search and parameter tuning use meta-training data only. Validation then selects one frozen program, which transfers to held-out tasks without structural changes or further LLM calls. Each new Gaussian process (GP) fits only covariance amplitude and observation noise.

Why numerical checks are not enough

A covariance matrix can pass a numerical PSD test on one set of inputs and fail on another. In three unrestricted-program cohorts, additional checks falsified 22–58% of the programs that passed the initial screen.

Stacked bars show 12 of 35, 18 of 31, and 8 of 37 initially passing unrestricted programs failing hidden checks. All five diagnostic canaries fail the hidden suite.
Figure 2. Additional checks change input scales and dimensions. Green marks programs that survive those checks, not a proof of kernel validity. The four contract-based controls pass Tier 2.

Kernaut checks execution and output shapes at Tier 0, sampled PSD at Tier 1, and construction plus consistency at Tier 2. The validity argument comes from trusted assembly under the contract assumptions. Source inspection and behavioral checks test conformance, but do not prove arbitrary Python programs pure over every input.

Experiments

The experiments test whether one frozen kernel structure transfers beyond the tasks used during search. We report the continuous ranked probability score (CRPS), which evaluates a predictive distribution against the observed value 2. Lower CRPS is better. Optimization uses regret AUC, the mean normalized simple regret over BO steps, where lower is also better.

Black-box optimization

Search and validation use five function families. Testing uses 60 tasks from six other families, with dimensions up to eight. Each predictive episode fits on 8 + 2d observations and scores 32 held-out points, where d is the input dimension. Each discovery campaign has 12 runs.

Validation selected the dual warp–fold (DWF) kernel in the first ensemble campaign. DWF combines a mild coordinate warp with a triangular fold. Its test CRPS was 0.601, compared with 0.612 for fixed Matérn-5/2. Its regret AUC was 0.671, compared with 0.693.

Predictive CRPS and optimization regret AUC across kernels on 60 test tasks. DWF has mean CRPS 0.601 and regret AUC 0.671.
Figure 3. Frozen kernels on 60 held-out BBO tasks. Lower is better on both axes. Bars show paired 95% bootstrap intervals for each method’s gap to DWF. Grid spectral mixture (GSM), input warping, and random Fourier feature (RFF) results average five fitted variants.

Across seven ensemble campaigns, six selected kernels improved test CRPS over Matérn-5/2. All seven GPT-6 Astra campaigns improved CRPS, with gains from 5.8% to 11.2%. These repeated campaigns measure discovery variability, whereas Figure 3 compares individual frozen kernels.

We also compared DWF with Few-Shot Bayesian Optimization (FSBO), which meta-learns a deep kernel 3. At the matched budget of 518 meta-training points, DWF had lower CRPS and regret AUC in both FSBO evaluation modes. At 64,000 points, neither FSBO mode differed significantly from DWF on either metric.

FSBO predictive CRPS and regret AUC at 518 and 64,000 meta-training points, compared with frozen DWF and Matérn-5/2.
Figure 4. Two measured FSBO data budgets, evaluated on the same 60 test tasks. Native FSBO fine-tunes on each episode. Frozen FSBO refits only amplitude and noise. Lines connect measured budgets and do not establish a scaling law.

Does open-ended synthesis matter?

The matched ablation keeps the proposer ensemble, run budget, and evaluation protocol fixed. Restricting agents to compositions of library kernels raises test CRPS from 0.571 to 0.618. Full search wins 11 of 12 paired campaigns against this closure-only arm.

Table 1. The first three rows report 12 seeded campaigns per arm. CRPS shows mean ± sample standard deviation. Greedy compositional search is one deterministic run. This ablation optimizes CRPS alone and uses different campaigns from Figure 3.
Search Test CRPS ↓ Regret AUC ↓
Full archive 0.571 ± 0.031 0.681
Independent proposals 0.617 ± 0.037 0.695
Closure grammar 0.618 ± 0.021 0.706
Greedy compositional search 0.625 0.684

Independent proposals also perform worse, but that arm changes feedback, novelty screening across runs, and candidate retention together. The comparison measures their combined contribution, rather than isolating archive sharing.

Forecasting unseen greenhouse-gas records

Search uses monthly CO2, CH4, and N2O records. Testing uses SF6, CFC-12, and CFC-11, with 30 windows per record and a 48-month horizon. Some windows reverse time, so the benchmark includes both forecasting and backcasting.

Validation selected a residual period-bank kernel. It combines a Matérn component, a smooth residual component, and low-frequency sinusoids that act as slow trends. The selected kernel ranks first among 58 frozen candidates on both CFC records.

Median-CRPS windows for SF6, CFC-12, and CFC-11. The period-bank kernel tracks CFC turning behavior more closely than RBF. Both CFC panels run backward in time.
Figure 5. The median-CRPS window for each held-out record. SF6 is forecast forward. Both CFC panels are backcasts, with held-out observations before the dotted origin. Bands show 95% predictive intervals, and concentrations are in parts per trillion.

Across all three records, the selected kernel has mean CRPS 0.091, compared with 0.120 for AutoGP, which infers a structure for each window 4. The overall gain comes mainly from reversed windows. On forward windows, its CRPS is 0.034, behind CAKE at 0.022. Windows overlap within each record, so these are not independent record-level replications.

Transfer to unseen mechanisms and patients

Enzyme kinetics. ChemBench maps seven experimental variables to a reaction rate. Search uses ten rate-law families. Testing uses 75 episodes from five unseen mechanisms, with 20 observations for fitting and 20 for scoring in each episode.

The validation-selected frontier kernel reaches CRPS 0.312, compared with 0.567 for optimized automatic relevance determination (ARD). ARD learns a separate lengthscale for each input. The selected kernel improves over all six learned references in paired tests with Holm correction. The selected ensemble kernel is competitive with the strongest learned references.

The frontier program combines substrate-response shapes inspired by the training mechanisms, local sensitivities, and modulation by enzyme loading. These features provide an inspectable biochemical prior. Predictive transfer does not establish recovery of the true mechanism.

Left: CRPS improvement across five held-out enzyme mechanisms for ensemble and frontier campaigns. Right: adult glucose prediction error falls as training episodes increase.
Figure 6. Transfer beyond the search data. Left: positive CRPS reductions favor discovered kernels over fixed Matérn-5/2. Validation selects ensemble E1 and frontier F2. Right: adult glucose prediction after a 45-minute prefix. Lower normalized MSE is better.

Glucose dynamics. GlucoseBench uses simulated patients. Search uses ten children, validation uses ten adolescents, and testing uses ten adults. Each forecast uses only training episodes from its target patient, so the experiment tests transfer of kernel structure.

With two training episodes, the five discovered kernels reduce geometric-mean normalized MSE by 38–67% against relative-time ARD-Matérn. With six episodes and a 45-minute prefix, the reductions narrow to 0–23%. The selected ensemble kernel encodes a delayed meal response and a fading insulin effect.

These glucose results describe ten simulated adults. Across 35 paired patient-level comparisons, no adjusted p-value falls below 0.05 after Holm correction. The reported percentage reductions therefore describe the observed benchmark means.

From Discovery to Human Refinement

DWF makes the discovered inductive bias visible. Its warp stays near the identity, while its fold bends the input geometry around a peak at 0.513. The original program applies a Matérn kernel jointly to the warped and folded coordinates.

Left: near-identity warp and triangular fold peaking at 0.513. Right: covariance with an input at 0.2, comparing joint DWF and the additive refinement near the fold mirror at 0.825.
Figure 7. The frozen DWF geometry. Joint DWF remains close to Matérn-5/2 at the fold mirror. The human-defined additive split increases covariance there to about 0.70, expressing a stronger preference for similar values at mirrored inputs.

A human refinement gives the warp and fold separate kernels, then averages their covariances. This additive construction preserves PSD and strengthens the reflection preference.

Table 2. Validation-selected refinement on the same 60 held-out BBO tasks. CRPS decreases by 5.7%, and mean regret AUC decreases by 7.8%. The additive construction is a human refinement, not an agent discovery. The archived kernel has a different historical search budget.
Construction Test CRPS ↓ Regret AUC ↓
Archived DWF 0.6011 0.6712
Additive branches, human refinement 0.5668 0.6190

The CRPS gain has a paired 95% bootstrap interval excluding zero. The regret reduction remains descriptive because the multiplicity-adjusted test does not cross 0.05. Smoothing the fold preserves the CRPS gain, so the sharp corner is unnecessary for this improvement.

Scope and Open Questions

The experiments test small-data Gaussian processes. The BBO benchmark has at most eight observed dimensions, and the scientific benchmarks use simulated mechanisms or patients. Performance on larger datasets, higher-dimensional tasks, and clinical data remains untested here.

Construction contracts establish validity under stated assumptions. They do not guarantee predictive accuracy, scientific novelty, or causal interpretation. Kernel fingerprints compare behavior on finite probe sets, and dimension-specific code can still fail outside its tested interface.

The short programs make further tests concrete: anonymize scientific variable names to measure the proposer’s use of domain knowledge, or add contracts for state-space and multi-output kernels. The documentation describes the implementation and workflow.

Citation

@misc{suwandi2026kernaut,
  title={Kernel Autoresearch for Open-Ended Model Discovery},
  author={Suwandi, Richard Cornelius and Yin, Feng and Murphy, Kevin},
  year={2026},
  url={https://www.alphaxiv.org/abs/2610.kernel-autoresearch},
  note={Preprint}
}

References

  1. Illuminating Search Spaces by Mapping Elites
    Mouret, J.-B. and Clune, J., 2015. arXiv:1504.04909.
  2. Strictly Proper Scoring Rules, Prediction, and Estimation
    Gneiting, T. and Raftery, A. E., 2007. Journal of the American Statistical Association, 102(477), 359–378.
  3. Few-Shot Bayesian Optimization with Deep Kernel Surrogates
    Wistuba, M. and Grabocka, J., 2021. International Conference on Learning Representations.
  4. Sequential Monte Carlo Learning for Time Series Structure Discovery
    Saad, F. A., Patton, B. J., Hoffman, M. D., Saurous, R. A. and Mansinghka, V. K., 2023. Proceedings of the 40th International Conference on Machine Learning, PMLR 202.