Contents

namespace

SushiBLAS::Reduce

Declared in
include/SushiBLAS/support/reduce_1d.hpp

Contains

Typedefs

template <typename T>
using SushiBLAS::Reduce::Accumulation = Kernels::AccumulatorOf<T>

Names the type a reduction over T accumulates in: FLOAT32 for HALF, T otherwise.

See also

include/SushiBLAS/support/README.md

template <typename T>
using SushiBLAS::Reduce::AccumulationOf = Kernels::Accumulator<T>

Names Accumulation<T>::type, the accumulation type itself.

Variables

constexpr std::size_t K_MAX_GROUPS = 256

Holds the ceiling on the number of per-work-group partials.

See also

include/SushiBLAS/support/README.md

constexpr std::size_t K_CPU_TWO_STAGE_MIN_ELEMENTS = 16384

Holds the element count at which a CPU switches from one pass to two.

See also

include/SushiBLAS/support/README.md

Functions

constexpr std::size_t floor_pow2(std::size_t v) noexcept

Returns the largest power of two not exceeding v, and never zero.

See also

include/SushiBLAS/support/README.md

std::size_t group_width(const Device::Profile &p) noexcept

Returns the work-group width stage one uses on p.

Parameters

p

The target device's profile.

Returns

A power-of-two work-group size the device will accept.

std::size_t planned_partials(const Device::Profile &p, std::size_t n) noexcept

Returns how many partials a two-stage reduction over n elements produces on p.

Returns

The partial count, or 0 to request the single-pass shape.

See also

include/SushiBLAS/support/README.md

Plan plan(const Device::Profile &p, std::size_t n, std::size_t scratch_capacity) noexcept

Returns the geometry to launch with, given the scratch actually available.

Parameters

scratch_capacity

Partials the caller's scratch buffer can hold.

Returns

The plan; Plan::sequential is true when there is no stage two.

See also

include/SushiBLAS/support/README.md

template <template< typename > class Accum, typename W, typename Out, typename Read, typename Finalize>
sycl::event reduce_indexed(sycl::queue &queue, std::size_t n, Read read, Out *out, W *partials, std::size_t partial_capacity, Finalize finalize, const std::vector< sycl::event > &deps)

Reduces indices [0..n) to *out as finalize(fold(read(i)), n), in a fixed order.

Parameters

read

Callable (std::size_t i) -> W, the input of element i.

partials

Scratch for one W per work-group; may be null when partial_capacity is zero.

finalize

Callable (W value, std::size_t n), applied once; converted to Out on store.

Returns

The event for the last submitted kernel.

See also

include/SushiBLAS/support/README.md

template <template< typename > class Accum, typename T, typename A, typename Transform, typename Finalize>
sycl::event reduce(sycl::queue &queue, const T *in, A *out, std::size_t n, AccumulationOf< A > *partials, std::size_t partial_capacity, Transform transform, Finalize finalize, const std::vector< sycl::event > &deps)

Reduces in[0..n) to *out as finalize(fold(transform(in[i])), n), in a fixed order.

Parameters

partials

Scratch for one AccumulationOf per work-group; may be null when partial_capacity is zero.

transform

Callable (T) -> A, applied to each element before the fold.

finalize

Callable (AccumulationOf<A> value, std::size_t n), applied once.

Returns

The event for the last submitted kernel.

See also

include/SushiBLAS/support/README.md

AxisPlan plan_axis(const Device::Profile &p, std::size_t segments, std::size_t extent) noexcept

Picks between the two segmented shapes for segments folds of extent elements on p.

Returns

The geometry; AxisPlan::cooperative says which shape.

See also

include/SushiBLAS/support/README.md

template <template< typename > class Accum, typename T, typename Transform, typename Finalize>
sycl::event reduce_axis(sycl::queue &queue, const T *in, T *out, const AxisSpan &span, const AxisPlan &geometry, Transform transform, Finalize finalize, const std::vector< sycl::event > &deps)

Reduces each segment of span to one element of out, in a fixed order.

Parameters

geometry

Which shape to launch, normally plan_axis's answer.

transform

Callable (T) -> T, applied to each element before the fold.

finalize

Callable (T value, std::size_t n) -> T, applied once per segment.

Returns

The event for the submitted kernel.

See also

include/SushiBLAS/support/README.md