The SushiAI CLI
SushiAI ships with a small command-line tool that drives everything you do day to day: building the C++ library, running its tests, running the demo and the example, and generating documentation. It is a thin wrapper around CMake and CTest that reads your machine-specific toolchain paths from a config file so you don’t have to retype long compiler flags.
Installing the CLI
The CLI is a Python package that lives in the cli/ folder. Install it once
and it puts two commands on your PATH:
sa— short name.sushiai— long name.
They are identical; every example below uses sa.
hub install-cli sushiai # always editable, against this checkout
This installs through pipx on every platform, into an isolated venv.
sushicore, the package the stack’s CLIs share, is a PyPI dependency declared
in cli/pyproject.toml, so pipx resolves it with the rest. The install is
always editable and there is no flag to make it otherwise: a frozen copy would
stop tracking git pulls on this checkout without ever saying so.
pipx uninstall sushiai-cli # to remove it later
How the CLI finds your compiler and siblings
SushiAI builds against the shared SushiStack SYCL toolchain — it selects no toolchain of its own — and against two sibling checkouts, SushiRuntime and SushiBLAS. The CLI reads all of this from:
cli/config.toml— committed, shared defaults.cli/config.local.toml— your machine’s absolute paths (gitignored).
With sibling ../sushiruntime and ../sushiblas checkouts, or inside a
SushiStack workspace, run sa setup once: it provisions the toolchain and
writes cli/config.local.toml for you. hub install does the same for
several checkouts at once. Run sa config to see what the CLI resolved.
Command overview
| Command | What it’s for |
|---|---|
sa build / test / run / clean |
Build, test, run, and clean the C++ project |
sa docs |
Build the Doxygen API reference into build/docs/api-site/html/ |
sa docs bundle |
Pack the published manual pages and the API reference into one archive |
sa demo mlp |
Train an MLP end to end and report the compile count |
sushiai train |
Train a model described by a JSON configuration |
sushiai inference |
Evaluate a trained checkpoint over a dataset |
sa config |
Show the resolved configuration |
sa env |
Show the environment your builds run under |
sa setup |
Provision the dependencies and write the tool paths |
sa doctor |
Report whether this machine can build SushiAI |
sa link / unlink |
Add this checkout to a workspace’s module registry, or remove it |
Run any command with --help to see its options.
sa --version prints sushiai-cli and its version. sa --describe prints the
command catalogue as JSON: every command, its options and their defaults.
A failure you can act on, such as a cli/config.toml that is not valid TOML,
ends in one line naming the file and exit code 1. Ctrl+C while a build, a test
run or a launched program is running stops it and exits 130.
sa build
sa build # Release build (the default)
sa build --type debug # Debug build
sa build --type relwithdebinfo
sa build --clean # delete the build tree first, then build from scratch
The --type (-t) option accepts release, debug, or relwithdebinfo.
If either sibling can’t be found, sa build fails fast with a clear message
naming which one and where it looked, rather than letting CMake’s own
add_subdirectory error surface. sa config shows what the CLI resolved for
both SushiRuntime and SushiBLAS; set sushiruntime_dir / sushiblas_dir
in cli/config.local.toml (or SUSHIRUNTIME_DIR / SUSHIBLAS_DIR in the
environment) to point either elsewhere.
-D / --define — CMake cache variables
sa build -D SA_BUILD_TESTS=OFF
sa build -D SB_SYCL_TARGETS="spir64;nvidia_gpu_sm_61" -D SA_BUILD_TESTS=OFF
Repeatable, VAR=VALUE, passed straight through to cmake — the same flag
sb build takes. A malformed entry exits 2 before the build tree is touched,
--clean included, with the message naming what you typed.
Two things worth knowing:
- The values are cached. They persist in the build tree, so dropping the
flag on a later run does not unset them. Pass
VAR=to clear one, or--cleanto start over. - They are appended last, after every default the CLI sets, so an explicit
-Dis never silently outranked by one of them.
This is how AOT device targeting is reached: SushiAI add_subdirectorys
SushiBLAS, so SushiBLAS’s SB_SYCL_TARGETS is a cache variable of this
build tree and there is no other sanctioned way to set it.
sa test
sa test # functional (unit+integration+regression) — the default
sa test --suite unit
sa test --suite integration
sa test --suite regression
sa test --suite all # every CTest test
sa test --suite package # the out-of-tree package consumer (no CTest)
sa test --filter 'Version.*' # ctest -R regex over 'Suite.Case' test names
sa test --repeat 5 # re-run each test up to 5x, stop on first failure
The three tiers come from the GTest suite-name prefix (Unit_*,
Integration_*, Regression_*) and nothing else — see
tests/CMakeLists.txt. regression holds the guarantees whose failure is
silent (see CONTRIBUTING.md §3).
For GoogleTest-level options CTest doesn’t expose, run the binary directly
with sa run and pass flags after --.
sa test needs a build with SA_BUILD_TESTS=ON, which is the default; a tree
configured with -D SA_BUILD_TESTS=OFF has no SushiAI tests registered.
--suite package
Every other suite is a CTest label. This one is not, and runs no CTest at all.
It installs SushiAI to build/package/prefix, configures tests/package/
into build/package/consumer against that prefix, builds it and runs it —
the way a third party consumes the library, through find_package(SushiAI),
with no access to this source tree. Both intermediates live under build/,
so sa clean reclaims them.
This is the only test in the repo that can see what find_package(SushiAI)
actually hands somebody. Every in-tree test includes individual headers from
the source tree, so a public header that the umbrella names but
install(DIRECTORY) does not ship, or that only compiles because a sibling
checkout happened to sit next to it, is invisible to all of them.
What it does not cover, and this is worth being exact about since the
suite was written in response to a header bug: the suite compiles what
SushiAI/SushiAI.h names, so it catches referenced but not installed. The
inverse — installed but never referenced by the umbrella, which is the actual
bug that prompted this work, four headers missing from the umbrella — it cannot
catch by construction. A header the umbrella does not name is simply never
compiled, and nothing fails. That half is checked at configure time instead, in
src/CMakeLists.txt: every include/SushiAI/**/*.hpp must appear in the
umbrella or the configure fails with the list. The two checks are
complementary and neither subsumes the other.
It also checks what the exported CMake interface does not carry — warnings,
-Werror, sanitizers, -march=, the in-tree policy targets, and the config
layer’s private nlohmann_json. Those assertions are in
tests/package/CMakeLists.txt, because what a consumer inherits is only
observable from the consumer side.
Three things the CLI does on the consumer’s behalf, all host-environment rather than package concerns:
- One install prefix covers all three modules. SushiBLAS and SushiRuntime are
add_subdirectory’d into this build (seecmake/BLAS.cmake), so their install rules run as part of this one and thefind_dependencycalls inSushiAIConfig.cmakeresolve from the same prefix. - The consumer configure reuses the same compiler and host-tool flags as the
main configure, from one shared helper (
_toolchain_args), so the two cannot drift.CMAKE_PREFIX_PATHis the deliberate exception — the consumer must also see the install prefix.SUSHIRUNTIME_DIR/SUSHIBLAS_DIRare excluded too, for a stronger reason: handing the consumer the sibling source trees is the shortcut this suite exists to forbid. - On Windows the built binary needs
sushiruntime.dllandhwloc-15.dllbeside it or it exits0xC0000135before reachingmain, which looks like a package failure and is not one. hwloc is linked PRIVATE by SushiRuntime and so never appears on a consumer’s link line — but PRIVATE is a link-time statement and the loader does not readINTERFACE_LINK_LIBRARIES; hwloc is insushiruntime.dll’s import table either way. In-tree binaries get it from PATH (saputs the vcpkg installedbinthere, seecli/sushiai/env.py); the CLI stages both DLLs for the consumer instead, and runs it with that PATH entry removed so the staging is actually load-bearing rather than shadowed by the developer’s environment.
Both build/package/prefix and build/package/consumer are deleted before
each run: a stale prefix would let a header deleted from the tree keep
satisfying an include, and a stale consumer directory would let a DLL staged by
an earlier run keep satisfying the loader.
It is not part of --suite all, which covers only the CTest tests. A full
install plus a second configure and build costs minutes and is rarely what you
want mid-loop, so all prints a line saying the package suite is not included
rather than leaving the gap to be discovered. --filter and --repeat are
CTest knobs; passing either with --suite package exits 2 instead of being
silently ignored.
The consumer compiles at C++17 — the same standard as the library, and
deliberately not SushiBLAS’s C++20 consumer. tests/package/CMakeLists.txt
records the measured reason (SushiAI exports out-of-line functions whose
parameter type is SushiRuntime::span, so the C++17/C++20 spelling difference
reaches the mangled name); see also INTEGRATION.md §3.
sa run
Runs a built executable. With no target it runs the configured default
(sushiai_example).
sa run # run the default target
sa run sushiai_example
sa run --sort # interactively pick from the list of executables
sa run sushiai_tests -- --gtest_filter='Unit_Version.*'
sa demo — the AI-7 showcase
sa demo mlp # the defaults
sa demo mlp --epochs 4 --optimizer sgd # anything you like
sa demo mlp --fusion off # the unfused graph, for an A/B
sa demo mlp --no-fuse gemm+bias+relu # one catalog entry off; repeatable
sa demo mlp --activation gelu # the epilogue that spills its input
sa demo mlp --precision mixed # fp16 products, fp32 master weights
sa demo mlp --loss-scale 65536 # gradient loss scaling, halving on overflow
sa demo mlp --help # the full flag list
--fusion and --no-fuse are how the operator fusion pass is measured rather
than assumed: both spellings produce the same numbers and differ only in how
many kernels they launch, so the comparison has to come from one binary against
one SushiBLAS build. --no-fuse takes a form’s name exactly as the demo’s
Fusion line prints it.
--activation picks the hidden layer’s nonlinearity, and the two it offers are
the two kinds rather than a sample of the library’s four. relu — the
default — has a gradient that reads its own output, so folding it into the GEMM
epilogue destroys nothing and needs nothing extra. gelu has a gradient that
reads its input, so gemm+bias+gelu only exists because its epilogue
spills that input to a second buffer on the way past (cuBLASLt calls the shape
GELU_AUX_BIAS). Measured here on the default schedule, at --precision fp32:
| run | ops per step |
buffers | final loss |
|---|---|---|---|
--activation gelu |
25 | 112.7 KiB | 0.000343 |
--activation gelu --no-fuse gemm+bias+gelu |
27 | 112.7 KiB | 0.000343 |
--activation gelu --fusion off |
35 | 112.7 KiB | 0.000343 |
Two launches saved against the fallback and ten against the unfused graph, at
no cost in memory at all — the spill is not a new buffer, it is the
pre-activation the unfused BIAS_ADD would have written anyway, kept instead of
eliminated. The identical final loss across all three is the point of the
exercise: the spill is what makes the gradients agree, and
Regression_Fusion.AGeluMlpTrainsIdenticallyFusedAndUnfused pins it as a
byte-for-byte parameter comparison rather than as six printed decimals.
Under --precision mixed the GELU layer is not folded — the fusion line
shows gemm+bias twice — because the tanh approximation computes x*x*x,
which overflows fp16 at |x| ≈ 40.3. See
ARCHITECTURE.md §6.2 for why that is a structural
consequence of the pass order rather than a special case.
--precision is the same idea for mixed precision. fp32 — the default —
means every operation computes in the graph’s own element type and the cast
pass inserts nothing, so nothing about an existing run changes. mixed runs
the matrix products and the epilogues a fused GEMM folds into them in fp16,
keeping fp32 master weights, fp32 gradients, fp32 reductions and an fp32
optimizer update; the demo’s Precision line reports how many conversions
that cost and in which direction.
On this machine --precision mixed is a correctness switch, not a speed
one. There is no fp16 matrix unit here, so half arithmetic is emulated and
the mixed run is several times slower — see
README.md § Mixed precision
for the measured figures. What the flag is for is running the two arms from
one binary and comparing the loss curves.
--loss-scale S turns on gradient loss scaling: the backward seed is
multiplied by S, every gradient is screened for infinities and NaNs, an
overflowed parameter’s update is skipped, the scale halves, and 2000 clean
steps double it. S must be a power of two so that the unscale before the
update is exact — the demo run at --precision fp32 --loss-scale 65536
reports a loss identical to the unscaled one, which is that exactness being
visible rather than a coincidence.
It makes this demo no better, and that is the honest reading. fp16 here
neither stalls nor underflows, so the flag is for watching the machinery:
--precision mixed --loss-scale 65536 overflows once and backs off to 32768
(PyTorch’s default scale is too high for this model’s fp16 backward pass), and
--loss-scale 16777216 overflows on every step and walks the scale down eight
halvings — with compiles: 1 throughout, which is what the scale being a
device-resident value rather than a graph attribute buys.
sa demo mlp runs sushiai_demo_mlp (examples/demo_mlp.cpp) from the repo
root, so a relative data path such as the demo’s data/mnist default means
the same thing wherever you invoked sa from. It trains an MLP end to end and
reports the compile count, the per-epoch loss trajectory, the final accuracy
and the wall clock. See
README.md § The demo for real output.
Two deliberate choices:
- The demo is its own binary, not
sushiai_examplewith flags. A demo is a supported entry point whose flags and printed lines are a contract the test suite pins; an example is a teaching file people are meant to edit.examples/demo_mlp.cpp’s header comment argues it in full. - Every flag is forwarded verbatim, including
--help. The demo owns its flag surface; restating it in Typer would give two places to change and one to forget.sa demo --helpstill lists the available demos. Until the demo is built there is nothing to forward to, sosa demo mlp --helpprintssa’s own page and then says to runsa build.sa trainandsa inferencedo the same.
sushiai train and sushiai inference — config-driven runs
sushiai train examples/heavy_mlp/train.json
sushiai inference examples/heavy_mlp/model.json \
examples/heavy_mlp/heavy_mlp.ckpt \
examples/heavy_mlp/eval.json
sa and sushiai are the same program (cli/pyproject.toml binds both), so
these work under either name; they are spelled sushiai here because that is
how the design document names them.
examples/mnist_cnn/ holds a second set of the three
files, for a convolutional network on MNIST. Its README says where the IDX files go.
What the two commands are for
A model’s architecture lives in a JSON file rather than in the C++ type
system, and these two drivers consume it. The point is not convenience: it is
that apps/train.cpp contains no layer, loss, dataset or optimizer type
name and therefore never has to be edited as the library grows. Every one of
those arrives through a constant registry table keyed by a string out of the
configuration, so adding a layer type or a detection loss is a table row and
a factory function in the library.
The three documents
| File | Describes | Named by |
|---|---|---|
model.json |
the architecture alone | train.json, and inference’s first argument |
train.json |
objective, data, optimizer, epochs, checkpoint | sushiai train |
eval.json |
the data to evaluate on | inference’s third argument |
{
"sushiai_model": 1,
"input": { "features": 6 },
"layers": [
{ "type": "linear", "units": 256 },
{ "type": "gelu" },
{ "type": "linear", "units": 4 }
]
}
{
"sushiai_train": 1,
"model": "model.json",
"objective": { "type": "cross_entropy" },
"data": { "type": "rings", "batch_size": 64, "batches": 64,
"classes": 4, "seed": 3735928559 },
"optimizer": { "type": "adamw", "learning_rate": 3e-3, "weight_decay": 0.01 },
"epochs": 50,
"checkpoint": "heavy_mlp.ckpt"
}
Four things about the schema are worth knowing before you write one, and each closes a way of being silently wrong:
- Each document leads with a version key. A format change is then an error rather than a misread.
unitsis an output width only. The input width is chained frominput.features, so a dimension mismatch between adjacent layers is not expressible.- An unknown key is an error, not an ignored one. A typed
"unitss": 256must not train silently against a default, so the message lists what the layer does accept. An unknowntypelikewise lists every type the build has. - Relative paths resolve against the configuration file’s own directory,
not your working directory — so
sushiai train examples/heavy_mlp/train.jsonworks from the repository root and a configuration directory can be moved without editing what is inside it. This holds formodelandcheckpoint. Thedatablock is passed to its factory as written, so thedirectoryof anmnistdataset resolves against the working directory of the process.
Layer types
NN::LAYER_REGISTRY in include/SushiAI/nn/layer_registry.hpp holds thirteen
rows. The shapes below are per sample and leave out the batch axis. The chain
starts at [input.features], and each layer receives the shape the layer
before it reports.
type |
Keys | Input | Output |
|---|---|---|---|
linear |
units (required, positive), use_bias (default true) |
[features] |
[units] |
gelu, relu, tanh, sigmoid |
none | any shape | the same shape |
leaky_relu |
slope (default 0.01, inside the open range 0 to 1) |
any shape | the same shape |
unflatten |
height, width, channels (all required) |
[height * width * channels] |
[height, width, channels] |
conv2d |
filters, kernel (required); stride (default 1), padding (default 0), dilation (default 1), use_bias (default true) |
[H, W, C] |
[out_h, out_w, filters] |
maxpool2d, avgpool2d |
kernel (required); stride (default: the kernel), padding (default 0) |
[H, W, C] |
[out_h, out_w, C] |
global_avgpool |
none | [H, W, C] |
[C] |
flatten |
none | any shape | [element count] |
upsample |
scale (default 2, at least 1) |
[H, W, C] |
[H * scale, W * scale, C] |
kernel, stride, padding and dilation each take an integer, which sets
both axes, or a two-entry [h, w] array. A padding applies to both sides of
its axis. A convolution reads its input channel count from the shape before
it, so there is no key for it.
Images enter a model flat. A model that convolves starts with unflatten,
and a linear layer after an image-shaped layer needs a flatten before it.
The factories refuse both mistakes and name the missing layer.
Dataset types
Data::DATASET_REGISTRY in include/SushiAI/data/dataset.hpp holds two rows.
type |
Keys |
|---|---|
rings |
batch_size (default 64), batches (64), classes (4), distractors (4), ring_width (1.0), position_noise (0.12), seed; all optional |
mnist |
directory (required); split ("train" or "test", default "train"), batch_size (default 64), max_samples (default 4096; 0 loads every sample) |
mnist reads the four uncompressed IDX files under their original names and
stops with an error that names them when the split’s two files are not in
directory.
Why inference takes its data as an argument
The architecture file carries no data block, because it describes a model. So
inference needs to be told what to run on, and an optional flag that always
errored when omitted would be a lie about optionality. It traces the forward
pass only — no loss, no backward pass, no optimizer.
It reports the predicted-class distribution and top-1 agreement with the
labels the dataset supplies. The distribution alone would be useless: on
class-balanced data a perfect model and a random one both give a uniform
spread. Reading the labels names no loss and no objective — labels are part of
the Data::Dataset seam, and top-1 agreement is a property of the data and
the output rather than of what the model was trained to minimise. It is
reported only when the model’s output width matches the dataset’s class count,
because otherwise the two are not comparable and a number would be fiction.
Exit codes and output
Both exit 0 on success and non-zero on any configuration or runtime error,
printing the library’s own diagnostic — which names the file, the key and what
was found. Both print a compiles: 1 line; two CTest cases pin it, the same
way sa demo mlp’s is pinned.
sa clean and sa docs
sa clean # remove the build/ tree
sa docs # build the Doxygen API reference
sa docs bundle --release 1.2.3 # pack the documentation bundle
sa clean removes the generated API reference too, because it lives under
build/. If the tree is still there afterwards, for example because a program
holds a file in it open, the command names the tree and exits 1.
sa docs was sa doxygen. The old word still runs, prints one notice naming
the new one when it runs on a terminal, and is removed after 2027-03-25; SA_NO_DEPRECATION_NOTICE=1
silences the notice.
It runs Doxygen against the repo-root Doxyfile, which is committed — for
a while it was not, and .gitignore listed it, so the command reported
Doxyfile not found at repo root. and exited 1 in every fresh checkout.
sushiruntime commits its own for the same reason.
Output lands in build/docs/api-site/html/, and warnings — every public symbol
missing a doc comment — in build/docs/api-site/doxygen-warnings.log. Both are
gitignored; the log is where the documentation debt is read off, and
WARN_AS_ERROR is deliberately NO so a missing @param cannot fail a docs
build.
Scope is include/SushiAI minus Detail. Unlike sushiruntime, which curates a
fluent api/ subtree, SushiAI has no such split: every installed header is
surface a consumer writes against.
Doxygen itself is an external tool, not vendored — install it or point
doxygen_exe in cli/config.local.toml at an existing binary; the command
prints the right install line if it cannot find one. sa setup installs one
with the rest of the shared tools.
sa docs bundle --release X.Y.Z writes docs-bundle-X.Y.Z.tar.gz and its
.sha256 file to build/docs/bundle/; --out names another folder. The
archive is what docs.sushisystems.io reads. docs/publish.toml lists the
sections of docs/ that go into it and the pages that stay out. Because that
file sets api = true, the command runs the same Doxygen build as sa docs
and adds the XML from build/docs/api-site/xml/. It stops, naming the page
and the line, when a published page has no level-one heading, has no link from
docs/README.md, or links to a file that does not exist. Without --release
it exits 2. The release is three integers: v1.2.3 and 1.2.3-rc1 are
refused.
sa config and sa env — diagnostics
sa config # print the resolved config and where each value came from
sa env # print the environment cmake/ctest/run subprocesses run under
sa env --all # show every variable, not just the build-relevant ones (-a for short)
sa config prints both resolved sibling directories, the compiler, and the
vcpkg root, plus the source of each field: default, config.toml or
env:<VARIABLE>. A value from cli/config.local.toml is reported as
config.toml, and a value from the workspace file as default.
sa setup
sa setup --dry-run # report what is missing and what would be written
sa setup # provision it
Provisions what cli/sushistack.deps.toml lists, and what the fragments of
SushiRuntime and SushiBLAS list, into the shared root, ~/.sushisystems (or
SUSHISYSTEMS_HOME). It then writes the tool paths it found to
cli/config.local.toml and prints the doctor report. It exits 2 and names the
module when SushiRuntime or SushiBLAS has no checkout.
--dry-runreports what is missing and what would be written, and changes nothing.--yesskips the confirmation prompt before the intel/llvm download.--toolchain <name>also installs that toolchain. Repeatable.--no-gpuskips the GPU toolkit.
hub install provisions several checkouts at once; sa setup covers this one
and the two it builds on.
sa doctor
Reports whether this machine can build SushiAI, one row per check, and exits non-zero when a required check fails.
--for GROUPruns only that group’s checks and counts its optional checks as required.
sa link and sa unlink
sa link records this checkout in a SushiStack workspace’s module registry;
sa unlink removes it.
--workspace <path>names the workspace root.
Building without the CLI
The CLI is a convenience, not a requirement. See Building with CMake directly in the top-level README.md.

