Skip to contents

Applies standard preprocessing steps to a training dataset: validates structure, coerces character columns to factors, drops unused factor levels, rejects missing and infinite values, and converts response labels to a factor.

Usage

preprocess_data(data, labels = NULL, ...)

Arguments

data

A data frame of raw training data.

labels

Optional vector of target labels. Converted to a factor with unused levels dropped.

...

Additional arguments (currently unused).

Value

A list with two elements:

  • $data: the preprocessed data frame

  • $labels: the processed factor labels (or NULL if not supplied)

Details

This function is called automatically by fit_model(). Do not call it manually before fit_model(), because doing so will result in feature metadata being extracted from the pre-processed data rather than the original data, which corrupts the imputation values used by boundary_compute() for 2D slicing.

It is exported for use in custom workflows and testing, but it is not typically needed in the standard fit_model() → boundary_compute() pipeline.