simd #
SIMD vectors
The simd module provides F32x4, a fixed vector of four f32 values. Its arithmetic operators work on corresponding lanes. On the C backend, the module uses SSE on x86 when available and NEON on AArch64. TCC and C targets without these extensions use scalar operations with the same API. The non-C backend implementation is also scalar.
import simd
fn main() {
a := simd.f32x4(1, 2, 3, 4)
b := simd.broadcast_f32x4(2)
println((a * b).to_array()) // [2.0, 4.0, 6.0, 8.0]
}
load_f32x4 reads four values from a slice and store writes four values. Both return an error if the slice is too short. For the final one to three values in a loop, use load_f32x4_part, which zero-fills unused lanes, and store_part, which writes only as many lanes as the destination contains. Partial operations reject slices longer than four.
mul_add computes a multiplication followed by an addition. It does not guarantee fused rounding. The module currently covers four-lane f32 arithmetic; math.vec.Vec4 remains the general-purpose geometric vector type. Compile with -cflags -DV_SIMD_FORCE_SCALAR to use the scalar C implementation when comparing behavior or performance.
The fixed width follows the four-lane layout used by projects such as Viper. Go's portable SIMD design provides a reference for a future size-independent API; this module does not yet offer runtime vector widths or feature dispatch.
fn broadcast_f32x4 #
fn broadcast_f32x4(value f32) F32x4
broadcast_f32x4 copies value into all four lanes.
fn f32x4 #
fn f32x4(a f32, b f32, c f32, d f32) F32x4
f32x4 creates a vector from four values in lane order.
fn load_f32x4 #
fn load_f32x4(src []f32) !F32x4
load_f32x4 loads four values from the start of src.
fn load_f32x4_part #
fn load_f32x4_part(src []f32) !F32x4
load_f32x4_part loads up to four values and zero-fills the remaining lanes.
struct F32x4 #
struct F32x4 {
values [4]f32
}
F32x4 holds four f32 values for lane-wise arithmetic.
fn (F32x4) to_array #
fn (v F32x4) to_array() [4]f32
to_array returns the four lanes in order.
fn (F32x4) store #
fn (v F32x4) store(mut dst []f32) !
store writes all four lanes to the start of dst.
fn (F32x4) store_part #
fn (v F32x4) store_part(mut dst []f32) !
store_part writes one lane for each element of dst, up to four elements.
fn (F32x4) + #
fn (v F32x4) + (other F32x4) F32x4
- adds corresponding lanes.
fn (F32x4) - #
fn (v F32x4) - (other F32x4) F32x4
- subtracts corresponding lanes.
fn (F32x4) * #
fn (v F32x4) * (other F32x4) F32x4
- multiplies corresponding lanes.
fn (F32x4) / #
fn (v F32x4) / (other F32x4) F32x4
/ divides corresponding lanes.
fn (F32x4) sqrt #
fn (v F32x4) sqrt() F32x4
sqrt returns the square root of each lane.
fn (F32x4) mul_add #
fn (v F32x4) mul_add(multiplier F32x4, addend F32x4) F32x4
mul_add returns v * multiplier + addend for each lane. It does not promise fused rounding.
fn (F32x4) sum #
fn (v F32x4) sum() f32
sum adds the four lanes in lane order.