Contents

namespace

SushiBLAS::Device

Declared in
include/SushiBLAS/support/device_policy.hpp

Contains

Variables

constexpr std::size_t K_WAVES_PER_COMPUTE_UNIT = 8

Holds the number of resident work-groups per compute unit a launch aims for.

See also

include/SushiBLAS/support/README.md

constexpr std::size_t K_GPU_WORK_GROUP_TARGET = 256

Work-group size a GPU launch aims for before clamping.

constexpr std::size_t K_CPU_WORK_GROUP_TARGET = 64

Work-group size a CPU launch aims for before clamping.

Small on purpose. On a CPU backend a work-group is a chunk handed to one OS thread and every barrier inside it is a real rendezvous, so wide groups buy nothing and cost synchronisation.

constexpr std::size_t K_VECTOR_TRANSACTION_BYTES = 16

Width in bytes of the widest single memory transaction to aim for.

Functions

const Profile & profile_for(const sycl::device &d)

Returns the cached Profile of a device, probing it on first sight.

Returns

A reference valid for the rest of the process.

See also

include/SushiBLAS/support/README.md

const Profile & profile_for(const sycl::queue &q)

The cached Profile for the device behind q.

std::size_t work_group_size(const Profile &p) noexcept

The work-group size to launch with, in work-items.

Clamped to what the device accepts and rounded down to a whole number of sub-groups, because a partial trailing sub-group is dead lanes on every device that has sub-groups at all.

bool hosts_local_tile(const Profile &p, const LocalTileDemand &demand) noexcept

Returns whether a device can host a work-group with demand's local memory and size.

See also

include/SushiBLAS/support/README.md

std::size_t flat_work_group_size(const Profile &p, std::size_t n) noexcept

Returns the work-group size a flat, barrier-free 1-D launch uses.

Returns

A size the device accepts; never zero.

See also

include/SushiBLAS/support/README.md

Grid1D plan_1d(const Profile &p, std::size_t n, std::size_t unroll=1) noexcept

Plans a grid-stride launch over n elements.

Parameters

unroll

Elements each work-item handles per stride step: 1 for scalar, 4 for 128-bit accesses.

Returns

A geometry whose global size is a non-zero multiple of the local size, even for n == 0.

See also

include/SushiBLAS/support/README.md

std::size_t vector_width(const Profile &p, std::size_t element_bytes) noexcept

Returns how many elements of element_bytes to load in one access.

Returns

A sycl::vec width of 1, 2, 4 or 8; always 1 on a CPU.

See also

include/SushiBLAS/support/README.md