namespace
SushiBLAS::Reduce
- Declared in
include/SushiBLAS/support/reduce_1d.hpp
Contains
SushiBLAS::Reduce::AxisPlanSushiBLAS::Reduce::AxisSpanSushiBLAS::Reduce::CompensatedComplexSumSushiBLAS::Reduce::CompensatedSumSushiBLAS::Reduce::ExactSumSushiBLAS::Reduce::IndexedValueSushiBLAS::Reduce::PlanSushiBLAS::Reduce::RunningArgMaxSushiBLAS::Reduce::RunningMaxSushiBLAS::Reduce::RunningProductSushiBLAS::Reduce::ScaledSumOfSquares
Typedefs
template <typename T>
using SushiBLAS::Reduce::Accumulation = Kernels::AccumulatorOf<T>Names the type a reduction over T accumulates in: FLOAT32 for HALF, T otherwise.
See also
include/SushiBLAS/support/README.md
template <typename T>
using SushiBLAS::Reduce::AccumulationOf = Kernels::Accumulator<T>Names Accumulation<T>::type, the accumulation type itself.
Variables
constexpr std::size_t K_MAX_GROUPS = 256Holds the ceiling on the number of per-work-group partials.
See also
include/SushiBLAS/support/README.md
constexpr std::size_t K_CPU_TWO_STAGE_MIN_ELEMENTS = 16384Holds the element count at which a CPU switches from one pass to two.
See also
include/SushiBLAS/support/README.md
Functions
constexpr std::size_t floor_pow2(std::size_t v) noexceptReturns the largest power of two not exceeding v, and never zero.
See also
include/SushiBLAS/support/README.md
std::size_t group_width(const Device::Profile &p) noexceptReturns the work-group width stage one uses on p.
Parameters
pThe target device's profile.
Returns
A power-of-two work-group size the device will accept.
std::size_t planned_partials(const Device::Profile &p, std::size_t n) noexceptReturns how many partials a two-stage reduction over n elements produces on p.
Returns
The partial count, or 0 to request the single-pass shape.
See also
include/SushiBLAS/support/README.md
Plan plan(const Device::Profile &p, std::size_t n, std::size_t scratch_capacity) noexceptReturns the geometry to launch with, given the scratch actually available.
Parameters
scratch_capacityPartials the caller's scratch buffer can hold.
Returns
The plan; Plan::sequential is true when there is no stage two.
See also
include/SushiBLAS/support/README.md
template <template< typename > class Accum, typename W, typename Out, typename Read, typename Finalize>
sycl::event reduce_indexed(sycl::queue &queue, std::size_t n, Read read, Out *out, W *partials, std::size_t partial_capacity, Finalize finalize, const std::vector< sycl::event > &deps)Reduces indices [0..n) to *out as finalize(fold(read(i)), n), in a fixed order.
Parameters
readCallable
(std::size_t i) -> W, the input of element i.partialsScratch for one W per work-group; may be null when partial_capacity is zero.
finalizeCallable
(W value, std::size_t n), applied once; converted to Out on store.
Returns
The event for the last submitted kernel.
See also
include/SushiBLAS/support/README.md
template <template< typename > class Accum, typename T, typename A, typename Transform, typename Finalize>
sycl::event reduce(sycl::queue &queue, const T *in, A *out, std::size_t n, AccumulationOf< A > *partials, std::size_t partial_capacity, Transform transform, Finalize finalize, const std::vector< sycl::event > &deps)Reduces in[0..n) to *out as finalize(fold(transform(in[i])), n), in a fixed order.
Parameters
partialsScratch for one AccumulationOf per work-group; may be null when partial_capacity is zero.
transformCallable
(T) -> A, applied to each element before the fold.finalizeCallable
(AccumulationOf<A> value, std::size_t n), applied once.
Returns
The event for the last submitted kernel.
See also
include/SushiBLAS/support/README.md
AxisPlan plan_axis(const Device::Profile &p, std::size_t segments, std::size_t extent) noexceptPicks between the two segmented shapes for segments folds of extent elements on p.
Returns
The geometry; AxisPlan::cooperative says which shape.
See also
include/SushiBLAS/support/README.md
template <template< typename > class Accum, typename T, typename Transform, typename Finalize>
sycl::event reduce_axis(sycl::queue &queue, const T *in, T *out, const AxisSpan &span, const AxisPlan &geometry, Transform transform, Finalize finalize, const std::vector< sycl::event > &deps)Reduces each segment of span to one element of out, in a fixed order.
Parameters
geometryWhich shape to launch, normally plan_axis's answer.
transformCallable
(T) -> T, applied to each element before the fold.finalizeCallable
(T value, std::size_t n) -> T, applied once per segment.
Returns
The event for the submitted kernel.
See also
include/SushiBLAS/support/README.md

