struct
SushiAI::Optim::LossScaleOptions
Configuration constants governing dynamic loss scaling.
- Declared in
include/SushiAI/optim/loss_scaler.hpp
Public attributes
double initial = 65536.0The scale the first step runs at; 2^16, as PyTorch starts.
High on purpose. Starting low and growing wastes the early steps at a scale that underflows, and the cost of starting too high is bounded and self-correcting: a handful of skipped steps while the backoff finds the ceiling.
double growth = 2.0What the scale is multiplied by after a clean run; must be a power of two >= 1.
double backoff = 0.5What it is multiplied by on overflow; must be a power of two in (0, 1].
std::uint64_t growth_interval = 2000How many consecutive clean steps buy one growth.
PyTorch's default is 2000. It is large because growing costs a skipped step when it overshoots and gains only headroom that was not being used, so the trade is asymmetric in the same direction the backoff is.
double minimum = 1.0The floor; the scale never backs off below this.
double maximum = 1073741824.0The ceiling; the scale never grows above this.

