SIMD
A data-parallel execution model in which one instruction stream applies the same operation to multiple packed data elements through vector registers.
SIMD (Single Instruction, Multiple Data) applies the same arithmetic or logical operation to multiple data elements through one instruction stream.
Processor-Level Meaning
A scalar addition operates on one pair of values. A vector instruction operates on a register containing several packed values; the effective width depends on the architecture and element type.
Examples include SSE/AVX on x86, NEON/SVE on ARM, and AltiVec/VSX on POWER systems.
What Determines Real Speedup?
SIMD does not guarantee acceleration by itself. Useful speedup depends on:
- contiguous and vectorizable data layout,
- memory bandwidth,
- load/store and alignment cost,
- branches and masking,
- scalar tail handling,
- compiler auto-vectorization quality.
If memory cannot feed the execution units fast enough, wider vectors may not increase end-to-end throughput.
Different from Thread Parallelism
SIMD is data parallelism inside one instruction stream; it is not the same as running multiple threads. Thread-level parallelism such as OpenMP and SIMD can be combined.
Applied context: IBM POWER9 AC922, Computer Architecture, and DCT/SIMD perceptual image similarity.
Related article: Hardware and Multimedia Processing
Vector Width Is Not the Same as Real Speedup
A vector instruction that processes W elements at once can offer at most roughly W-way arithmetic parallelism for an ideal kernel that is fully vectorizable and has no other bottleneck. If only a fraction p of the program benefits from that width, then even while ignoring memory and other overheads an Amdahl-style coarse upper bound is:
S <= 1 / ((1-p) + p/W)Doubling vector width therefore does not imply a twofold end-to-end speedup when a serial fraction or memory bandwidth dominates.
SIMD evaluation should record more than the instruction-set name: vectorized-loop coverage, load/store traffic, cache misses, memory bandwidth, and scalar-tail work all matter. When SIMD is combined with thread-level parallelism such as OpenMP, both layers also share the same memory subsystem.
Software-Optimization Context
SIMD speedup does not come from vector width alone. Contiguous data layout, alignment, branching, cache locality and memory bandwidth all affect the result. A vectorized loop may execute fewer instructions yet fail to scale as expected when the workload is bounded by memory bandwidth.