Contents

struct

SushiAI::Optim::LossScaleOptions

Configuration constants governing dynamic loss scaling.

Declared in
include/SushiAI/optim/loss_scaler.hpp

Public attributes

double initial = 65536.0

The scale the first step runs at; 2^16, as PyTorch starts.

High on purpose. Starting low and growing wastes the early steps at a scale that underflows, and the cost of starting too high is bounded and self-correcting: a handful of skipped steps while the backoff finds the ceiling.

double growth = 2.0

What the scale is multiplied by after a clean run; must be a power of two >= 1.

double backoff = 0.5

What it is multiplied by on overflow; must be a power of two in (0, 1].

std::uint64_t growth_interval = 2000

How many consecutive clean steps buy one growth.

PyTorch's default is 2000. It is large because growing costs a skipped step when it overshoots and gains only headroom that was not being used, so the trade is asymmetric in the same direction the backoff is.

double minimum = 1.0

The floor; the scale never backs off below this.

double maximum = 1073741824.0

The ceiling; the scale never grows above this.