Skip to content

PMLB

The Penn Machine Learning Benchmarks (PMLB) is a curated collection of benchmark datasets for evaluating machine learning methods. HGP is evaluated on the binary classification datasets in the suite.

Data preparation

The preprocessing script lives at scripts/preprocess/pmlb_preprocess.py. It fetches a named dataset with the pmlb package and writes it to the data folder as an HDF file.

python scripts/preprocess/pmlb_preprocess.py breast_cancer --data_path data

This writes data/breast_cancer.hdf.

The list of binary classification datasets comes from the PMLB summary table. The same query is used by the batch runner to decide which datasets to process.

from pmlb.dataset_lists import df_summary

binary_datasets = df_summary[df_summary["n_classes"] == 2.0]["dataset"].tolist()

Warning

Some datasets listed in older PMLB summaries are no longer available, because they were removed as duplicates. Fetching such a dataset raises an error. Skip the missing datasets and continue with the rest.

Running a benchmark

Benchmark the Boolean GP on a single preprocessed dataset with scripts/run_benchmark.py.

python scripts/run_benchmark.py --data_path data/breast_cancer.hdf

See python scripts/run_benchmark.py --help for the full list of options.

Hyperparameter tuning

Tune a single dataset with the Optuna script. The study name keeps trials grouped, and artifacts are written to the artifact directory.

python scripts/optuna_hypertuning.py \
    --data_path data/breast_cancer.hdf \
    --study_name pmlb_breast_cancer \
    --n_trials 25 \
    --max_n_trials 25 \
    --n_runs 5 \
    --n_folds 3 \
    --artifact_dir ./artifacts

Running on all datasets

The scripts/run_on_pmlb.py script runs the Optuna tuning across every binary classification dataset. It fetches the dataset names, builds one optuna_hypertuning.py command per dataset, and runs them in parallel.

python scripts/run_on_pmlb.py --n_jobs 4

The --n_jobs argument sets how many datasets are processed at the same time. Preprocessing is not required beforehand. When a dataset file is missing, optuna_hypertuning.py downloads it from PMLB and writes the HDF file before tuning.

Results on PMLB

This case study evaluates HGP on 70 binary classification datasets from PMLB. We compare two HGP variants against two interpretable baselines.

The HGP variants are:

  • HGP (default): run with run_on_pmlb.py --model gp_default --num_epochs 10000, using the library defaults.
  • HGP (tuned): run with run_on_pmlb.py --model gp_tuned, which tunes each dataset with the Optuna script for 100 trials using the standard Optuna hyperparameter search.

The baselines are:

  • Decision trees from scikit-learn1, an interpretable tree-based classifier.
  • BoolXAI2, a local-search method that also evolves expressive boolean formulas.34

All results use F1 score on the held-out test set, averaged across runs. Each comparison counts, per dataset, which method reached the best (or tied for best) score.

Per-dataset scores

The score of each variant across all datasets.

HGP default F1 scores across PMLB datasets

HGP tuned F1 scores across PMLB datasets

Tuned vs default

Tuning helps on most datasets. HGP (tuned) was best or tied on 59 of 70 datasets, and HGP (default) on 19 of 70.

HGP tuned vs HGP default

Tuned vs decision tree

HGP (tuned) outperforms the decision tree baseline. HGP (tuned) was best or tied on 51 of 70 datasets, and the decision tree on 27 of 70.

HGP tuned vs decision tree

Default vs decision tree

Without tuning, HGP is on par with the decision tree. Each was best or tied on 38 of 70 datasets.

HGP default vs decision tree

Tuned vs decision tree vs BoolXAI

Against both baselines at once, HGP (tuned) leads. HGP (tuned) was best or tied on 48 of 70 datasets, the decision tree on 20 of 70, and BoolXAI on 12 of 70.

HGP tuned vs decision tree vs BoolXAI

References


  1. scikit-learn DecisionTreeClassifier, https://scikit-learn.org/stable/modules/generated/sklearn.tree.DecisionTreeClassifier.html

  2. BoolXAI, https://github.com/fidelity/boolxai

  3. Gili Rosenberg, John Kyle Brubaker, Martin J. A. Schuetz, Grant Salton, Zhihuai Zhu, Elton Yechao Zhu, Serdar Kadıoğlu, Sima E. Borujeni, and Helmut G. Katzgraber. Explainable artificial intelligence using expressive boolean formulas. Machine Learning and Knowledge Extraction, 5(4):1760–1795, 2023. doi:10.3390/make5040086

  4. Serdar Kadioglu, Elton Yechao Zhu, Gili Rosenberg, John Kyle Brubaker, Martin J. A. Schuetz, Grant Salton, Zhihuai Zhu, and Helmut G. Katzgraber. Boolxai: explainable ai using expressive boolean formulas. Proceedings of the AAAI Conference on Artificial Intelligence, 39(28):28900–28906, Apr. 2025. doi:10.1609/aaai.v39i28.35157