Skip to content

Meta-evaluation protocol

Meta-evaluation assesses whether a discovered kernel transfers beyond the tasks used during search. Kernaut follows the separation between discovery tasks and evaluation tasks discussed by Goldie et al. (2026). The task split and the observation split within a task serve different purposes.

Two levels of separation

An episode is one small regression, forecasting, or optimization problem. Within an episode, fitting observations determine the Gaussian process parameters. Separate query observations measure predictive error. For forecasting, those queries are the withheld suffix of the episode's sequence.

At the meta level, episodes or task families are assigned to three roles:

Partition Role Permitted adaptation
Meta-training Generate and revise kernel programs, tune their parameters, and choose reference settings Search can use evaluator feedback from these episodes
Meta-validation Select among the frozen candidate programs using a prespecified metric Select a candidate without revising its code or kernel parameters
Meta-test Assess the selected program on families or patient groups withheld from search Apply the same predefined fitting procedure without revising or reselecting the program

A meta-training episode can contain query observations that are held out from its GP fit. Their evaluation results still provide feedback to search, so these observations are not a final test set. Conversely, a meta-test episode includes fitting observations that its GP may use under the fixed evaluation protocol.

In the domain evaluators, the kernel program and its parameters remain fixed during benchmark evaluation. The GP fits covariance amplitude and observation noise from each episode's fitting observations using fixed grids. Target standardization also uses those fitting observations. For GlucoseBench, parameter fitting uses the other complete training episodes, and conditioning additionally uses the query prefix.

Benchmark partitions

Benchmark Meta-training Meta-validation Meta-test
Black-box optimization Five function families New seeded transformations of the same five families Six different function families
Greenhouse gases CO2, CH4, N2O New seeded episodes from those records SF6, CFC-12, CFC-11
ChemBench Ten canonical mechanism domains New input samples in those domains Five structurally distinct domains
GlucoseBench Ten child profiles Ten adolescent profiles Ten adult profiles through the external hidden-test evaluator

The benchmark catalogue lists the families and available package interfaces. ts-benchmark --split test evaluates the three held-out gases, SF6, CFC-12, and CFC-11. The greenhouse split separates generated episodes and held-out gas records. It does not make training and validation calendar months disjoint. For custom temporal benchmarks, define disjoint time intervals if that is the intended generalization test.

Black-box optimization (BBO), greenhouse, and ChemBench episodes use the seed formula:

seed = split_offset + 1000 * family_index + episode_index
split_offset = 0 (train), 10000 (validation), 20000 (test)

Keep fewer than 1,000 episodes per family or domain with this scheme to avoid overlapping seed ranges. ChemBench uses this seed for input sampling. Oracle noise follows the external oracle's implementation. The GlucoseBench exporter uses a fixed seed derived from each patient identifier.

Configuration

Partition membership is defined in the domain adapters, rather than as a percentage split in TOML. Command-line options control the evaluated partition and episode budgets. [campaign].seed controls search and verification randomness. It does not change the built-in meta partitions or their episode seed offsets.

Interface Relevant options and defaults
meta-run --episodes-per-family 2, always meta-training
meta-benchmark --split all, --episodes-per-family 1, --bo-steps 8, --candidate-pool-size 128
ts-run --episodes-per-family 3, always meta-training, 48-month full evaluation horizon
ts-benchmark --split all, --episodes-per-family 1, --horizon-months 48
chem-run Always meta-training. --episodes-per-domain 1, --train-points 20, --test-points 20, --difficulty easy, --noise-level 0.01
chem-benchmark The same episode and oracle options, plus --split all
ChemBench scoring --worst-domain-weight 0.0 in both commands. Positive values add a worst-domain CRPS penalty
glucose-run --data selects the exported observations. Search uses child profiles
glucose-benchmark --split all means training and validation only. There is no hidden-test CLI option

Use explicit --split validation during selection because the first three benchmark commands otherwise include test results. Keep episode sizes, transformations, preprocessing, and oracle settings consistent between candidate and baseline evaluations. Changing family membership, seed offsets, or the greenhouse training-window minimum requires a custom evaluator or an adapter change. There is no general --train-ratio, --validation-ratio, or --test-ratio option. Built-in parameter-tuning tools can use reduced meta-training episodes. Their scores are not validation results.

Search, select, and evaluate

After configuring a model, run a search on BBO meta-training episodes:

kernaut meta-run --config configs/openai.toml --context examples/meta_bo_task.md \
  --archive runs/bbo/archive.sqlite --strategy evolutionary --episodes-per-family 2

Stop the search before selection. Preserve the resulting archive and evaluate its accepted candidates on meta-validation episodes:

kernaut meta-benchmark --archive runs/bbo/archive.sqlite \
  --split validation --episodes-per-family 2 --bo-steps 8 --candidate-pool-size 128 \
  > runs/bbo/validation.json

Choose the metric before inspecting results, such as splits.validation.predictive.mean_crps for prediction or splits.validation.bo.regret_auc for optimization. Select the candidate with the lowest value under that rule and record its candidate_id. The default candidate limit is 100. Increase --limit or repeat --candidate-id when your prespecified candidate set requires it.

Replace SELECTED_ID below with the chosen identifier before evaluating the held-out families:

kernaut meta-benchmark --archive runs/bbo/archive.sqlite \
  --split test --candidate-id SELECTED_ID --episodes-per-family 2 \
  --bo-steps 8 --candidate-pool-size 128 > runs/bbo/test.json

Benchmark commands include their reference baselines even when --candidate-id is supplied. Reference settings are selected using meta-training evaluations. The reported JSON is separate from the search archive's evaluation records. Do not use the order of the test report to select another program.

The same select-then-test procedure applies to ts-benchmark and chem-benchmark. For GlucoseBench, select on adolescent folds and use the external evaluator for hidden adult testing. Repeated decisions based on validation results also adapt to that partition, so record the number of candidates and the selection rule.

Metrics and reporting

CRPS evaluates a predictive distribution against an observed value. Lower CRPS is preferable. See Gneiting and Raftery (2007) for proper scoring rules. Kernaut negates loss-based scores so that larger EvaluationRecord.score values indicate better candidates. Search can also include novelty or worst-domain penalties, so report the raw predictive metrics separately from search fitness.

Report the split definitions, episode counts, candidate set, selection rule, baseline budgets, failures, and software revisions with any scientific comparison. Scores from different tasks or preprocessing protocols are not directly comparable merely because both use CRPS. For a new dataset or domain, follow custom benchmarks.