Skip to content

simd #

SIMD vectors

The simd module provides F32x4, a fixed vector of four f32 values. Its arithmetic operators work on corresponding lanes. On the C backend, the module uses SSE on x86 when available and NEON on AArch64. TCC and C targets without these extensions use scalar operations with the same API. The non-C backend implementation is also scalar.

import simd

fn main() {
    a := simd.f32x4(1, 2, 3, 4)
    b := simd.broadcast_f32x4(2)
    println((a * b).to_array()) // [2.0, 4.0, 6.0, 8.0]
}

load_f32x4 reads four values from a slice and store writes four values. Both return an error if the slice is too short. For the final one to three values in a loop, use load_f32x4_part, which zero-fills unused lanes, and store_part, which writes only as many lanes as the destination contains. Partial operations reject slices longer than four.

mul_add computes a multiplication followed by an addition. It does not guarantee fused rounding. The module currently covers four-lane f32 arithmetic; math.vec.Vec4 remains the general-purpose geometric vector type. Compile with -cflags -DV_SIMD_FORCE_SCALAR to use the scalar C implementation when comparing behavior or performance.

The fixed width follows the four-lane layout used by projects such as Viper. Go's portable SIMD design provides a reference for a future size-independent API; this module does not yet offer runtime vector widths or feature dispatch.

fn broadcast_f32x4 #

fn broadcast_f32x4(value f32) F32x4

broadcast_f32x4 copies value into all four lanes.

fn f32x4 #

fn f32x4(a f32, b f32, c f32, d f32) F32x4

f32x4 creates a vector from four values in lane order.

fn load_f32x4 #

fn load_f32x4(src []f32) !F32x4

load_f32x4 loads four values from the start of src.

fn load_f32x4_part #

fn load_f32x4_part(src []f32) !F32x4

load_f32x4_part loads up to four values and zero-fills the remaining lanes.

struct F32x4 #

struct F32x4 {
	values [4]f32
}

F32x4 holds four f32 values for lane-wise arithmetic.

fn (F32x4) to_array #

fn (v F32x4) to_array() [4]f32

to_array returns the four lanes in order.

fn (F32x4) store #

fn (v F32x4) store(mut dst []f32) !

store writes all four lanes to the start of dst.

fn (F32x4) store_part #

fn (v F32x4) store_part(mut dst []f32) !

store_part writes one lane for each element of dst, up to four elements.

fn (F32x4) + #

fn (v F32x4) + (other F32x4) F32x4
  • adds corresponding lanes.

fn (F32x4) - #

fn (v F32x4) - (other F32x4) F32x4
  • subtracts corresponding lanes.

fn (F32x4) * #

fn (v F32x4) * (other F32x4) F32x4
  • multiplies corresponding lanes.

fn (F32x4) / #

fn (v F32x4) / (other F32x4) F32x4

/ divides corresponding lanes.

fn (F32x4) sqrt #

fn (v F32x4) sqrt() F32x4

sqrt returns the square root of each lane.

fn (F32x4) mul_add #

fn (v F32x4) mul_add(multiplier F32x4, addend F32x4) F32x4

mul_add returns v * multiplier + addend for each lane. It does not promise fused rounding.

fn (F32x4) sum #

fn (v F32x4) sum() f32

sum adds the four lanes in lane order.