Pimp My IDE / Garage Dispatch
Back to garage
September 25, 2026 | Go / SIMD / performance proof

Portable speed still needs four real lanes.

Go 1.27 can express one vectorized algorithm across amd64, arm64, and WebAssembly. The code got portable. The performance claim did not.

The take: Go's experimental simd package removes architecture size from the source-level contract. That is a strong portability move. Before replacing a scalar loop, pin the workload, keep the scalar baseline, run the target hardware, and test the fallback.
Build the experiment card

One source file can address several vector widths.

Go 1.27 adds an experimental, platform-independent simd package. It currently targets AVX, AVX2, and AVX512 on amd64, NEON on arm64, and WebAssembly SIMD. Unsupported platforms use emulation, so the same source still runs.[1]

The package keeps fixed vector sizes out of its types. A value such as simd.Float32s reports its length at runtime. The portable package supports operations shared across targets and fills some missing operations with other instructions.[1]

The abstraction makes the source portable. Only a matched run can make the speed claim portable.

The package is an experiment, not a magic import.

Developers must build with GOEXPERIMENT=simd. The first release also has deliberate gaps. The Go team calls out reduction as one example. Go 1.27 lacks a portable sum across every vector lane, while a later release is expected to add ReduceSum.[1]

The release notes place simd beside the lower-level simd/archsimd package. The portable package covers the shared operation set. The architecture package exposes target-specific power when the shared set cannot express the job.[2]

Measure the work users run.

Go's profile-guided optimization guide gives the right warning for this experiment. A useful profile must represent production behavior. The guide says microbenchmarks are usually poor PGO inputs because they exercise little of the whole program.[3]

A microbenchmark still helps compare one scalar loop with one SIMD loop. It cannot predict request latency, memory pressure, or throughput for the whole service. Keep both views. Measure the kernel, then run the real workload on every target you plan to support.

Use four lanes before changing the default.

  1. Representative load: Save the input shape, size distribution, hot path, and production profile that justify the work.
  2. Scalar baseline: Keep the readable implementation and compare output before comparing time.
  3. Target matrix: Run on the actual amd64, arm64, and WebAssembly targets you ship. Record CPU features and toolchain revision.
  4. Fallback drill: Force the non-vector or emulated route. Check correctness, acceptable performance, and a clean rollback.

The dyno below writes an experiment template. It does not run Go, detect CPU instructions, compare outputs, or measure speed.

Interactive makeover / benchmark planning

Vector lane dyno.

Traditional purpose replaced: a benchmark note beside the patch. Better version: lock workload, baseline, target matrix, and fallback into one copyable experiment card before the optimized route becomes the default.

Clamp the evidence

Native checkboxes own the state. A selected clamp means the experiment requires that evidence. It does not mean the evidence exists.

Experiment lanes
EXPERIMENT EMPTY0 / 4 LANES SELECTED

No experiment lane is selected.

Select the evidence the run must collect before the SIMD path changes the default.

Completion claim0 of 4 experiment lanes selected. No benchmark has run.
Suggested command shapeGOEXPERIMENT=simd go test -bench '<KERNEL>' -benchmem -count '<RUNS>'
This teaching control writes an experiment template. It does not inspect source, execute a benchmark, detect CPU features, compare outputs, profile production, or approve a rollout.

Copy the dyno card

Fill after copyingAdd the Go revision, source revision, hardware, CPU features, workload, commands, raw output, correctness result, and rollback result.

Sources read, not vibes

Open the source log
  1. The Go Blog, "Platform-independent SIMD in Go": current package scope, supported targets, emulation behavior, runtime vector length, experimental flag, first-release limits, and stated design goals.
  2. Go 1.27 release notes: release status, compatibility statement, experimental portable simd package, and lower-level simd/archsimd package.
  3. Go profile-guided optimization guide: representative production profiles, limits of microbenchmarks as whole-program inputs, and reproducible profile handling.
  4. Hacker News discussion: exact discovery route and practitioner discussion. Comments are context, not evidence for the package claims above.

Source boundary: The Go team documents an experimental API and its goals. It does not promise a speedup for every operation or machine. The PGO guide covers whole-program optimization rather than SIMD adoption. The four-lane experiment is our synthesis.