Fitting models with neuralnetwork

neuralnetwork fits multilayer perceptrons for tabular regression and classification. The examples below use base R data sets and keep the training runs short enough for R CMD check, but the same calls work with your own data frames and matrices.

library(neuralnetwork)

Multiclass classification

Start with the formula interface. For small tabular data, hidden = "auto" and optimizer = "auto" are enough for a first fit.

fit_class <- nn_fit(
  Species ~ .,
  data = iris,
  hidden = "auto",
  optimizer = "auto",
  epochs = 10,
  validation_split = 0.2,
  seed = 1,
  verbose = FALSE
)

fit_class
#> neuralnetwork model
#>   Task:       classification
#>   Layers:     4 -> 4 -> 3
#>   Optimizer:  lbfgs | activation: tanh | backend: rcpp
#>   Loss:       cross_entropy
#>   Trained:    13 function evaluations
#>   Convergence: code 1 | message: NEW_X
#>   Selected:   train loss 0.014447 | accuracy 0.99167
#>               validation loss 0.38622 | validation accuracy 0.9

The printed model reports the task, architecture, optimizer, loss, backend, training length, and the selected checkpoint’s training and validation metrics. They need not come from the last epoch. Epoch optimizers store their trajectory in fit_class$history; L-BFGS stores final diagnostics in one row.

Use predict() for classes or probabilities.

predict(fit_class, iris[1:5, ], type = "class")
#> [1] setosa setosa setosa setosa setosa
#> Levels: setosa versicolor virginica
round(predict(fit_class, iris[1:5, ], type = "prob"), 3)
#>      setosa versicolor virginica
#> [1,]      1          0         0
#> [2,]      1          0         0
#> [3,]      1          0         0
#> [4,]      1          0         0
#> [5,]      1          0         0

nn_evaluate() returns metrics for the fitted task. Multiclass classification includes accuracy, balanced accuracy, macro precision, macro recall, macro F1, and log loss.

ev_class <- nn_evaluate(fit_class, iris)
ev_class
#> neuralnetwork evaluation
#> Metrics:
#>             metric    value
#>           accuracy  0.97333
#>  balanced_accuracy  0.97333
#>    macro_precision  0.97359
#>       macro_recall  0.97333
#>           macro_f1   0.9734
#>           log_loss 0.088802
#> 
#> Confusion matrix:
#>             estimate
#> truth        setosa versicolor virginica
#>   setosa         49          1         0
#>   versicolor      0         48         2
#>   virginica       0          1        49

This call scores all 150 flowers, including training rows. It is a workflow check, not an independent test-set estimate. nn_cv() below fits each fold without its assessment rows. When choosing settings from the same data, keep a separate test set out of both training and tuning for the final evaluation.

For imbalanced classification, balanced accuracy or F1 is usually more useful than raw accuracy. For probability forecasts, inspect log loss as well.

When a class is never predicted

Accuracy can hide a missed minority class. In the following confusion matrix, none of the 19 Poor observations is predicted correctly, yet overall accuracy is about 83%.

confusion <- matrix(
  c(20, 0, 45, 0, 0, 19, 18, 0, 378), nrow = 3, byrow = TRUE,
  dimnames = list(truth = c("Excellent", "Poor", "Typical"),
                  estimate = c("Excellent", "Poor", "Typical"))
)
confusion
#>            estimate
#> truth       Excellent Poor Typical
#>   Excellent        20    0      45
#>   Poor              0    0      19
#>   Typical          18    0     378
class_f1 <- 2 * diag(confusion) / (rowSums(confusion) + colSums(confusion))
round(class_f1, 5)
#> Excellent      Poor   Typical 
#>   0.38835   0.00000   0.90215
round(c(accuracy = sum(diag(confusion)) / sum(confusion),
        macro_f1 = mean(class_f1)), 5)
#> accuracy macro_f1 
#>  0.82917  0.43017

The three F1 scores are 0.38835, 0, and 0.90215, giving macro F1 of 0.43017. Dropping the zero would give 0.64525 and overstate performance. nn_evaluate() computes F1 directly from counts and includes every fitted class in macro precision, recall, and F1. A class-wise zero denominator returns zero. Balanced accuracy averages recall only over classes present in the truth; it equals macro recall when all fitted classes occur in the evaluation set. When a class is absent from the truth, its reported recall of zero is a convention, not evidence that the model failed on examples of that class.

Binary classification and weights

Two-class outcomes use a one-output sigmoid model internally. The public prediction API still returns a two-column probability matrix.

iris_binary <- subset(iris, Species != "virginica")
row_weight <- ifelse(iris_binary$Species == "versicolor", 1.5, 1)

fit_binary <- nn_fit(
  Species ~ .,
  data = iris_binary,
  hidden = c(6, 3),
  optimizer = "adam",
  epochs = 8,
  batch_size = 16,
  learning_rate = 0.01,
  sample_weight = row_weight,
  class_weight = "balanced",
  gradient_clip = 5,
  validation_split = 0.2,
  seed = 2,
  verbose = FALSE
)

round(predict(fit_binary, iris_binary[1:5, ], type = "prob"), 3)
#>      setosa versicolor
#> [1,]  0.687      0.313
#> [2,]  0.729      0.271
#> [3,]  0.670      0.330
#> [4,]  0.670      0.330
#> [5,]  0.680      0.320
nn_evaluate(fit_binary, iris_binary)
#> neuralnetwork evaluation
#> Metrics:
#>             metric   value
#>           accuracy    0.97
#>  balanced_accuracy    0.97
#>    macro_precision  0.9717
#>       macro_recall    0.97
#>           macro_f1 0.96997
#>           log_loss 0.48052
#>        sensitivity    0.94
#>        specificity       1
#>          precision       1
#>             recall    0.94
#>                 f1 0.96907
#> 
#> Confusion matrix:
#>             estimate
#> truth        setosa versicolor
#>   setosa         50          0
#>   versicolor      3         47

Weights apply to the training objective even when a batch contains only one row. Here the network is just an intercept. It starts at zero, the first target is zero, and the second update carries most of the training weight.

intercept_update <- function(weights) {
  fit <- nn_fit(matrix(0, 2, 1), c(0, 10), hidden = 0,
                optimizer = "sgd", epochs = 1, batch_size = 1,
                learning_rate = 0.1, sample_weight = weights,
                scale = FALSE, y_scale = FALSE, shuffle = FALSE,
                seed = 1, verbose = FALSE)
  unname(coef(fit)$b[[1]])
}
c(weight_on_second = intercept_update(c(1, 9)),
  weight_on_first = intercept_update(c(9, 1)))
#> weight_on_second  weight_on_first 
#>              1.8              0.2

The intercepts are 1.8 and 0.2. Dividing each one-row gradient by that row’s weight would erase the difference. Training instead uses the full training weight sum and the actual batch size. With l2 = 0, multiplying every row weight by the same positive constant leaves this objective unchanged. Zero-weight training rows are omitted from learning, scalers and fitted formula bases; they still must pass input encoding checks. Scaling uses ordinary means and standard deviations among the retained rows, not weighted moments.

Regression

Regression follows the same shape. By default, regression targets are scaled for training and predictions are returned on the original scale.

fit_reg <- nn_fit(
  mpg ~ wt + hp + disp,
  data = mtcars,
  hidden = c(8, 4),
  optimizer = "adam",
  epochs = 25,
  batch_size = 8,
  learning_rate = 0.01,
  validation_split = 0.2,
  seed = 3,
  verbose = FALSE
)

fit_reg
#> neuralnetwork model
#>   Task:       regression
#>   Layers:     3 -> 8 -> 4 -> 1
#>   Optimizer:  adam | activation: sigmoid | backend: rcpp
#>   Loss:       squared_error
#>   Trained:    25 epochs | best epoch: 23
#>   Selected:   train loss 0.13226 | rmse 3.2451
#>               validation loss 0.030743 | validation rmse 1.5645
round(predict(fit_reg, mtcars[1:5, ]), 2)
#> [1] 23.12 22.42 24.74 19.54 16.42
nn_evaluate(fit_reg, mtcars)
#> neuralnetwork evaluation
#> Metrics:
#>  metric   value
#>    rmse  3.0025
#>     mae   2.281
#>     rsq 0.74381

Robust regression

Squared error is the default regression loss. If a few observations may be unusually influential, use Huber loss.

mtcars_outlier <- mtcars
mtcars_outlier$mpg[1] <- mtcars_outlier$mpg[1] + 40

fit_huber <- nn_fit(
  mpg ~ wt + hp,
  data = mtcars_outlier,
  hidden = 4,
  optimizer = "adam",
  loss = "huber",
  huber_delta = 1,
  epochs = 20,
  batch_size = 8,
  learning_rate = 0.01,
  seed = 4,
  verbose = FALSE
)

summary(fit_huber)
#> neuralnetwork summary
#>   Task:       regression
#>   Layers:     2 -> 4 -> 1
#>   Optimizer:  adam | activation: sigmoid | backend: rcpp
#>   Loss:       huber (delta=1)
#>   Epochs:     20 | best epoch: 20
#> 
#> Selected model training row:
#>  epoch train_loss validation_loss train_metric validation_metric gradient_norm
#>     20    0.18135              NA       7.6733                NA       0.41666
#>  learning_rate backtracked
#>           0.01       FALSE

Training controls

The training loop supports dropout, L2 regularization, gradient clipping, learning-rate decay, validation splits, early stopping, and callbacks. This example stops after two epochs so the callback behavior is visible without making the vignette slow.

epochs_seen <- 0L

fit_callback <- nn_fit(
  mpg ~ wt + hp,
  data = mtcars,
  hidden = 4,
  optimizer = "adam",
  epochs = 20,
  batch_size = 8,
  learning_rate = 0.01,
  l2 = 1e-4,
  dropout = 0.05,
  gradient_clip = 5,
  validation_split = 0.2,
  callbacks = function(state) {
    epochs_seen <<- state$epoch
    if (state$epoch >= 2) {
      return(list(stop = TRUE))
    }
    NULL
  },
  seed = 5,
  verbose = FALSE
)

fit_callback
#> neuralnetwork model
#>   Task:       regression
#>   Layers:     2 -> 4 -> 1
#>   Optimizer:  adam | activation: sigmoid | backend: rcpp
#>   Loss:       squared_error
#>   Trained:    2 epochs | best epoch: 2
#>   Selected:   train loss 0.52135 | rmse 6.4176
#>               validation loss 0.13126 | validation rmse 3.2201
#>   Stopped:    callback

Training choices:

Tuning and cross-validation

Use nn_tune() for a grid search. Classification metrics include accuracy, balanced_accuracy, f1, and log_loss. Regression metrics include rmse, mae, and rsq.

tuned <- nn_tune(
  Species ~ .,
  data = iris,
  grid = list(
    hidden = list(4, c(6, 3)),
    learning_rate = c(0.01)
  ),
  metric = "balanced_accuracy",
  epochs = 4,
  validation_split = 0.2,
  seed = 6,
  verbose = FALSE
)

tuned
#> neuralnetwork tuning result
#>   Candidates: 2
#>   Objective:  balanced_accuracy (higher is better)
#>   Best score: 0.97619
#> 
#> Top candidates:
#>  hidden learning_rate candidate_id success            metric   score error rank
#>       4          0.01            1    TRUE balanced_accuracy 0.97619  <NA>    1
#>    6, 3          0.01            2    TRUE balanced_accuracy 0.84392  <NA>    2
tuned$best_params
#>   hidden learning_rate
#> 1      4          0.01

When exploring a wider grid, error_action = "continue" keeps invalid candidate combinations in the result table and ranks the usable fits. The default remains strict, so a bad grid fails before it produces misleading results.

Use nn_cv() for fold-level estimates.

cv <- nn_cv(
  Species ~ .,
  data = iris,
  k = 3,
  metric = "f1",
  hidden = 4,
  epochs = 2,
  seed = 7,
  verbose = FALSE
)

cv
#> neuralnetwork cross-validation
#>   Folds:   3
#>   Repeats: 1
#> 
#>    metric    mean      sd n_scored n_folds
#>  macro_f1 0.86362 0.10213        3       3

Permutation importance

Permutation importance measures how much a metric changes when one feature is shuffled.

imp <- nn_permutation_importance(
  fit_reg,
  mtcars,
  metric = "mae",
  n_repeats = 2,
  seed = 8
)

imp
#> neuralnetwork permutation importance
#>   Metric:  mae
#>   Repeats: 2
#> 
#>  feature importance baseline permuted metric n_repeats
#>       wt    0.98792    2.281   3.2689    mae         2
#>     disp    0.57524    2.281   2.8562    mae         2
#>       hp    0.54879    2.281   2.8298    mae         2

Save, load, and inspect

Models are regular R objects. nn_save() and nn_load() add package-level checks around saveRDS() and readRDS().

model_path <- tempfile(fileext = ".rds")
nn_save(fit_reg, model_path)
fit_loaded <- nn_load(model_path)

all.equal(
  predict(fit_reg, mtcars[1:3, ]),
  predict(fit_loaded, mtcars[1:3, ])
)
#> [1] TRUE

The package includes compatibility helpers for common nnet and neuralnet tasks.

nn_class_ind(iris$Species[1:4])
#>      setosa versicolor virginica
#> [1,]      1          0         0
#> [2,]      1          0         0
#> [3,]      1          0         0
#> [4,]      1          0         0

computed <- nn_compute(fit_class, iris[1:2, ])
names(computed$neurons)
#> [1] "input"   "hidden1" "output"
round(computed$net.result, 3)
#>      setosa versicolor virginica
#> [1,]      1          0         0
#> [2,]      1          0         0

Checking scales and parameter intervals

For a formula such as log(mpg) ~ wt, the response is log-mpg. Automatic target scaling is undone before reporting predictions, but log() is not: nn_evaluate() compares those predictions with log(mpg) too. With several responses, RMSE averages squared errors over rows and output columns, while R-squared pools variation after centering each output separately. Outputs with larger units can dominate the pooled R-squared.

log_fit <- nn_fit(log(mpg) ~ wt, mtcars, hidden = 0, epochs = 100,
                  verbose = FALSE, seed = 12)
nn_evaluate(log_fit, mtcars)$metrics
#>      rmse       mae       rsq 
#> 0.1318687 0.1015628 0.7975582
sqrt(mean((predict(log_fit, mtcars) - log(mtcars$mpg))^2))
#> [1] 0.1318687

The last number should match the reported RMSE. Exponentiating predictions produces a different prediction target; it does not automatically estimate the conditional mean of mpg.

metric = "loss" means the fitted data loss throughout tuning, cross-validation and permutation importance. It excludes L2 penalties and uses the fitted target scaler. That makes loss suitable for checking one model’s objective, but it is not interchangeable with original-scale RMSE. Tuning candidates share one holdout. For a fully reproducible manual split, pass validation_rows to nn_fit() and inspect fit$validation$rows. Row weights can also be passed to nn_evaluate(); training class weights do not silently become evaluation weights.

For coefficient uncertainty, start with a model whose parameters are identifiable. A network without hidden layers and with a single linear output can be checked directly against ordinary least squares:

linear_fit <- nn_fit(mpg ~ wt, mtcars, hidden = 0, optimizer = "lbfgs",
                     scale = FALSE, y_scale = FALSE, epochs = 300,
                     verbose = FALSE, seed = 12)
nn_confint(linear_fit, mtcars)
#>      term  estimate std.error conf.low conf.high
#> 1 W1[1,1] -5.344474  0.559101 -6.48631 -4.202637
#> 2   b1[1] 37.285133  1.877627 33.45051 41.119760
confint(lm(mpg ~ wt, mtcars))
#>                 2.5 %    97.5 %
#> (Intercept) 33.450500 41.119753
#> wt          -6.486308 -4.202635

The limits agree; lm() prints the intercept first, whereas network parameters list weights before biases. Disabling scaling here makes both sets of coefficients use the same units. The regression intervals assume independent errors with constant Gaussian variance. Unregularized binary logistic fits without hidden layers use normal, asymptotic limits instead. Both require a converged L-BFGS fit and the original training observations.

Hidden-layer weights, redundant softmax parameters, weighted fits and Huber fits do not receive intervals from nn_confint(). A singular information matrix or saturated logistic probabilities also produces an error. Adding a small constant to a Hessian would make it invertible, but would not make these intervals statistically justified. nn_hessian() remains available as a mean-loss curvature diagnostic, not a parameter covariance matrix.

nn_generalized_weights() reports input sensitivities per unit before predictor scaling. Its classification outputs are probability derivatives, not the log-odds derivatives provided by neuralnet. For factors, polynomial bases and interactions, the inputs are encoded columns rather than raw variables.

Function map

Function names:

Need Use
Fit a regression or classification network nn_fit()
Fit a no-hidden-layer multinomial model nn_multinom()
Get class probabilities or numeric predictions predict()
Score a fitted model nn_evaluate()
Tune a small grid nn_tune()
Run repeated k-fold validation nn_cv()
Estimate feature importance nn_permutation_importance()
Get compute-style hidden activations nn_compute()
Get generalized weights nn_generalized_weights()
Save and reload a model nn_save() and nn_load()

Start with nn_fit(), inspect nn_evaluate(), and add nn_tune() or nn_cv() when the first model is worth more computation.

Reference help: ?neuralnetwork, ?neuralnetwork-metrics, ?neuralnetwork-callbacks, and ?neuralnetwork-objects.

mirror server hosted at Truenetwork, Russian Federation.