Generates synthetic classification data using Gaussian mixture models constructed by
the MixSim package. Useful for creating controlled datasets to explore decision
boundary behavior when no real dataset is available.
Usage
simulate_mixsim(
n,
K,
p,
MaxOmega,
class_names = NULL,
seed = NULL,
noise_ratio = 0,
test_ratio = 0
)Arguments
- n
Number of training observations to generate.
- K
Number of classes.
- p
Number of numeric features (dimensions).
- MaxOmega
Maximum pairwise overlap between mixture components (0 to 1). Smaller values create better-separated classes.
- class_names
Optional character vector of class labels (length
K). Defaults to"Class 1","Class 2", etc.- seed
Optional integer for reproducibility. The global random seed is restored after the call.
- noise_ratio
Numeric in [0, 1). Proportion of
nto add as uniform background noise (randomly labeled). Use this to test classifier robustness to contamination.- test_ratio
Numeric in [0, 1). If greater than 0, generates an additional independent test dataset of size
round(n * test_ratio)using the same mixture parameters. This is not a split of the training data.
Value
If test_ratio == 0 (default): a data frame with a Sim class column and
p feature columns (X1, X2, ...). If test_ratio > 0: a list with $train
and $test data frames.
Details
Class distributions are randomly generated subject to the MaxOmega overlap
constraint. The simulation is fully reproducible when seed is supplied.
Test data
When test_ratio > 0, an additional independent dataset is generated using the
same mixture parameters but a fresh random draw. This is not a split of the
training data: the training set has n observations and the test set has
round(n * test_ratio) independently generated observations. The two sets are
statistically independent, so the test set is a fair evaluation sample.
Examples
# \donttest{
# Generate a 3-class, 2-dimensional training dataset
train_df <- simulate_mixsim(n = 200, K = 3, p = 2, MaxOmega = 0.05, seed = 42)
head(train_df)
#> Sim X1 X2
#> 1 Class 1 0.9061175 1.0388990
#> 2 Class 1 0.8026125 1.0469506
#> 3 Class 1 0.9219127 0.9029906
#> 4 Class 1 0.9210607 0.9177018
#> 5 Class 1 0.9798823 0.8153331
#> 6 Class 1 0.9352758 0.9874468
# Generate training and independent test data
sim <- simulate_mixsim(
n = 200, K = 3, p = 2, MaxOmega = 0.05,
seed = 42, test_ratio = 0.3
)
nrow(sim$train) # 200
#> [1] 200
nrow(sim$test) # 60 (independently generated, not split from train)
#> [1] 60
# }