Skip to contents

Generates synthetic classification data using Gaussian mixture models constructed by the MixSim package. Useful for creating controlled datasets to explore decision boundary behavior when no real dataset is available.

Usage

simulate_mixsim(
  n,
  K,
  p,
  MaxOmega,
  class_names = NULL,
  seed = NULL,
  noise_ratio = 0,
  test_ratio = 0
)

Arguments

n

Number of training observations to generate.

K

Number of classes.

p

Number of numeric features (dimensions).

MaxOmega

Maximum pairwise overlap between mixture components (0 to 1). Smaller values create better-separated classes.

class_names

Optional character vector of class labels (length K). Defaults to "Class 1", "Class 2", etc.

seed

Optional integer for reproducibility. The global random seed is restored after the call.

noise_ratio

Numeric in [0, 1). Proportion of n to add as uniform background noise (randomly labeled). Use this to test classifier robustness to contamination.

test_ratio

Numeric in [0, 1). If greater than 0, generates an additional independent test dataset of size round(n * test_ratio) using the same mixture parameters. This is not a split of the training data.

Value

If test_ratio == 0 (default): a data frame with a Sim class column and p feature columns (X1, X2, ...). If test_ratio > 0: a list with $train and $test data frames.

Details

Class distributions are randomly generated subject to the MaxOmega overlap constraint. The simulation is fully reproducible when seed is supplied.

Test data

When test_ratio > 0, an additional independent dataset is generated using the same mixture parameters but a fresh random draw. This is not a split of the training data: the training set has n observations and the test set has round(n * test_ratio) independently generated observations. The two sets are statistically independent, so the test set is a fair evaluation sample.

See also

Examples

# \donttest{
# Generate a 3-class, 2-dimensional training dataset
train_df <- simulate_mixsim(n = 200, K = 3, p = 2, MaxOmega = 0.05, seed = 42)
head(train_df)
#>       Sim        X1        X2
#> 1 Class 1 0.9061175 1.0388990
#> 2 Class 1 0.8026125 1.0469506
#> 3 Class 1 0.9219127 0.9029906
#> 4 Class 1 0.9210607 0.9177018
#> 5 Class 1 0.9798823 0.8153331
#> 6 Class 1 0.9352758 0.9874468

# Generate training and independent test data
sim <- simulate_mixsim(
  n = 200, K = 3, p = 2, MaxOmega = 0.05,
  seed = 42, test_ratio = 0.3
)
nrow(sim$train) # 200
#> [1] 200
nrow(sim$test) # 60 (independently generated, not split from train)
#> [1] 60
# }