Back to archive
Field noteApr 09, 202642 min read

AOS vs SOA vs AoSoA: The Answer Is Not What You Think

Everyone says Structure of Arrays is faster. The real answer is "it depends" — on the number of fields, the data size, and what operation you are doing. I built an AoSoA container, benchmarked it six ways, diagnosed an MLP ceiling, cracked it with multi-stream traversal, and then ran a realistic N-body frame where SOA turned out to be 3x slower than AOS. Every isolated-benchmark intuition about layouts is a trap.

AOS vs SOA vs AoSoA: The Answer Is Not What You Think

"Use Structure of Arrays, it is faster." Every C++ performance talk. Every data-oriented design book. Every game engine post-mortem. It became received wisdom. So when I sat down to build a generic C++20 SOA container, I expected the benchmarks to confirm the folklore. Instead, I ended up with a decision table — and then I added AoSoA, the supposed "best of both worlds," watched it fail spectacularly, diagnosed why, and fixed it in a way that upended my original conclusion.

I built SOA_iterator — a variadic SOA container with a proxy iterator, benchmarked against equivalent AOS code across 240 tests. SOA is indeed faster for some operations. But for others, AOS wins by 2-3x. I then added an AoSoA container (Array of SOA blocks) under the same iterator abstraction, expecting a layout that gets SIMD and whole-object locality. The first measurements were bad: AoSoA trailed both AOS and SOA on almost everything. This was going to be a cautionary tale about how "zero-cost abstraction" is not transitive across layouts. And it was — for the first version. Then I moved iteration from operator++ to a for_each lambda, rebenchmarked, and AoSoA started beating AOS on every operation and matching SOA in cache. The article below walks through both acts honestly, because the second act only makes sense if you have seen the first.

The Setup

If you have never seen the distinction: AOS (Array of Structures) stores objects contiguously:

AOS: [x0,y0,z0] [x1,y1,z1] [x2,y2,z2] ...

Each object's fields are next to each other. Accessing one object is fast. Iterating over a single field means striding through memory — you touch every cache line but only use a fraction of each.

SOA (Structure of Arrays) transposes the layout:

SOA: [x0,x1,x2,...] [y0,y1,y2,...] [z0,z1,z2,...]

Each field has its own contiguous array. Iterating over one field is sequential, the prefetcher is happy, and SIMD vectorization works naturally. But accessing all fields of one object means touching N arrays (N cache lines in the worst case).

The conventional wisdom is "SOA is faster because it vectorizes." What the benchmarks reveal is more nuanced.

A Generic SOA in C++20

The common objection to SOA is "you lose your nice struct syntax." SOA_iterator addresses that with three abstractions: a variadic container, a proxy reference, and a forward iterator. Here is the core:

template<typename... Ts> class SOA { std::tuple<std::vector<Ts>...> arrays; // one vector per field public: template<size_t I> auto& array() { return std::get<I>(arrays); } size_t size() const { return std::get<0>(arrays).size(); } // ... };

The key expression is std::tuple<std::vector<Ts>...>. The pack expansion generates one vector per type. For SOA<int, float, double> this becomes std::tuple<std::vector<int>, std::vector<float>, std::vector<double>>.

The proxy reference is what lets you write AOS-style code against SOA data:

class Proxy { std::tuple<Ts&...> refs; // references into each array at the same index public: template<size_t I> auto& get() { return std::get<I>(refs); } };

A tuple of references bundles scattered memory locations into something that looks like a single object. And the iterator constructs one on dereference:

class Iterator { SOA* soa_; size_t index_; public: Proxy operator*() { return make_proxy(std::index_sequence_for<Ts...>{}); } private: template<size_t... Is> Proxy make_proxy(std::index_sequence<Is...>) { return Proxy{{std::get<Is>(soa_->arrays)[index_]...}}; } };

The make_proxy line is the magic: pack expansion indexes into every array simultaneously and bundles the references. With -Ofast, the compiler sees through the proxy, the iterator, and the tuple-of-references, and produces the same machine code as hand-written index-based loops. I verified this by benchmarking both the iterator-based (for (auto proxy : soa)) and raw index-based variants of every operation — their timings match within noise, confirming the abstraction really is zero cost.

The generic operations use fold expressions:

template<typename Tuple> auto sum_all_fields(const Tuple& t) { return std::apply([](const auto&... args) { return (args + ...); // unary right fold }, t); }

This works identically whether t is an AOS struct's tuple or a Proxy's tuple of references. One function, both layouts.

Field Count: The Most Dramatic Result

The first finding is the one nobody talks about explicitly: SOA's advantage scales with the number of fields you do not use.

Benchmarked on Intel Core Ultra 7 at 5.2 GHz, 1K elements, pinned to core 0, -Ofast -march=native:

FieldsAOS ReadSOA ReadSpeedup
3 (float3)426 ns382 ns1.12x
4 (float4)535 ns400 ns1.34x
8 (float8)1339 ns399 ns3.36x

The "Read" benchmark sums all fields of all elements — the friendliest case for SOA because every byte loaded is useful. Notice: the SOA time is nearly flat around 400 ns regardless of field count. The AOS time grows linearly because each cache line now holds proportionally less useful data.

At 8 fields (32 bytes per element), AOS is wasting its cache bandwidth. SOA, reading 8 separate contiguous arrays, gets SIMD vectorization and perfect cache utilization on each one.

The project's original README measured 10x on this configuration on a different machine with 24 MiB L3. My machine with 12 MiB L3 gets "only" 3.36x, but the direction of the effect is the same. The more fields you have, the more SOA wins. If your hot path only touches 1 of 8 fields, you are throwing away 87.5% of every cache line with AOS.

Scaling: The Non-Monotonic Curve

I expected the SOA advantage to diminish smoothly as arrays grow and memory becomes bandwidth-bound. That is not what happens.

For int3 reads on this machine:

SizeSOA Speedup
1K (12 KB)2.71x
4K (48 KB)2.53x
32K (384 KB)2.12x
262K (3 MB)1.34x
1M (12 MB)1.63x

The dip at 262K is not noise — it corresponds to the working set exceeding L2 (2 MiB per cluster on this CPU). Both layouts are now pulling data from L3 or beyond. The advantage of SOA comes from two independent sources:

  1. SIMD vectorization — dominant in-cache, the compiler can emit wide AVX2 loads
  2. Prefetcher-friendly access — dominant out-of-cache, hardware prefetchers love sequential streams

At the L2/L3 boundary, neither effect is at its peak. Inside L2, vectorization is the whole story. At L3 and main memory, prefetcher behavior takes over and SOA's sequential streams regain the advantage. The dip in between is where both effects are paying half their dividend.

This is the kind of finding that only shows up when you measure carefully across sizes. A single-size benchmark would miss it.

It Depends: The Operation Matters

Here is the chart that tells the real story:

At 1M elements, int3 configuration:

SOA wins:

  • Read: 1.63x
  • Compute: 1.69x (read-then-arithmetic pattern)
  • Search: 3.65x (scanning one array vs strided access)
  • ConditionalTransform: 1.54x
  • ComputePushBack: 1.76x

AOS wins:

  • FilterCopy: 2.63x (copy elements matching a predicate)
  • Merge: 1.79x (merge two sorted datasets)
  • Write: ~tie

The pattern is clear. Anything that reads fields wins with SOA. Anything that moves whole objects wins with AOS.

FilterCopy and Merge are the killer cases for SOA. When you "copy one element," AOS does a single memcpy of 12 bytes — one contiguous chunk. SOA has to do three separate push_backs into three separate vectors — three allocator touches, three cache line writes, three length updates. The overhead compounds.

Search is where the result is most surprising: 3.65x SOA advantage. I expected a modest improvement because linear search scans a single field and early-exits. But at 1M elements, the AOS layout is 12 MB — right at the edge of my 12 MiB L3. SOA only needs to read the one field being searched (4 MB), staying comfortably in L3. AOS has to pull every cache line through the hierarchy, fetching 12 bytes and using 4. At this size, the working set difference dominates everything else.

The AoSoA Detour: "Best of Both Worlds" That Is Not

Looking at the above table, the temptation is obvious. SOA wins on column access; AOS wins on whole-object moves. The textbook answer to this trade-off is AoSoA: Array of SOA Blocks — also called "tiled SOA" or "chunked SOA." You take N consecutive elements, store them as a tiny SOA tile, and then lay out an array of those tiles. Inside each tile, column access is SIMD-friendly. Across tiles, whole-object copies stay local. On paper, you get the best of both layouts.

I added an AoSoA class to the project — variadic, templated on a block size B, using the same Proxy and Iterator shape as my SOA container so the generic operations would work unchanged:

template<size_t B, typename... Ts> struct alignas(64) Block { std::tuple<std::array<Ts, B>...> data; // SOA tile of B elements }; template<size_t B, typename... Ts> class AoSoA { std::vector<Block<B, Ts...>> blocks; size_t size_ = 0; public: // Proxy: same shape as SOA::Proxy — std::tuple<Ts&...> // Iterator tracks (block_idx, offset); ++ does if (++offset==B) { offset=0; ++block_idx; } Iterator begin(); Iterator end(); Proxy operator[](size_t i); };

The goal was for for (auto p : aosoa) to work exactly like for (auto p : soa), hiding the block structure. Generic functions like sum_all_fields would not care which container they got. I benchmarked 16 configurations: 4 type packs (float3, float8, double3, int_float_double) × 4 block sizes (B = 4, 8, 16, 64). Then I stared at the numbers for a while.

At 1M float8 elements (the configuration where SOA's advantage is largest), the speedups vs AOS are:

OperationSOAAoSoA (B=8)
Read2.04x0.93x
Write0.79x0.88x
Compute1.96x0.88x
FilterCopy0.25x0.82x
Merge1.00x0.97x

AoSoA lands between AOS and SOA on every operation — and on four of the five, it is slightly worse than plain AOS. The one place it shines is FilterCopy, where it is 3.3x faster than SOA (0.82x vs 0.25x vs AOS). That one entry is the whole case for AoSoA on this workload. Everywhere else, it is mediocre.

Why? The Iterator Tax

The issue is the iterator. For SOA, operator++ is ++index_ — a single integer increment with no branches, and the compiler trivially vectorizes for (auto p : soa) into wide loads over contiguous arrays. For AoSoA, operator++ is:

Iterator& operator++() { if (++offset_ == B) { offset_ = 0; ++block_idx_; } return *this; }

That branch runs every iteration. GCC at -Ofast cannot peel and fuse it to produce a clean inner loop of B scalar operations per block followed by a block-index bump. It tries to vectorize the whole loop anyway, but ends up emitting vpermd-based permute code with register spills to stack — the classic sign that the compiler is trying too hard on data that does not fit its preferred shape. You can see it in the generated assembly:

vpermd (%rdx), %ymm13, %ymm1 vpermd (%rdx), %ymm9, %ymm3 vpermd -2016(%rdx), %ymm11, %ymm2 vblendps $240, %ymm2, %ymm1, %ymm6 vmovaps %ymm6, 1696(%rsp) ; spill vmovaps %ymm5, 160(%rsp) ; spill ...

Versus the clean SOA loop which is basically vaddps on contiguous memory. The iterator abstraction that was free for SOA (I verified this in the main section — iterator-based and index-based SOA loops benchmark identically) is not free for AoSoA, because the branch in operator++ defeats the auto-vectorizer.

I tried writing block-direct "raw" traversals that bypass the iterator and process whole blocks at a time. With per-field passes, GCC vectorizes cleanly but touches each cache line once per field — cache locality dies and performance gets dramatically worse (14 ms vs 1.6 ms at 1M). With fused per-block loops, the compiler spills registers on the nested N-field × B-element expansion. In short: to beat SOA with AoSoA under this workload, I would need to hand-write traversal code per operation, giving up exactly the generic abstraction that makes SOA_iterator worth building.

Size Scaling And Block Size

The size sweep tells the same story from another angle. At small sizes (1K elements, working set fits in L1), AoSoA does edge past AOS on float8 Read (1.5x). As the working set grows past L2 and L3, the iterator tax dominates and AoSoA trails AOS. SOA is faster than both at every size — and by an increasing margin as the compiler's SIMD advantage compounds.

Block size, surprisingly, barely matters:

B = 4, 8, 16 are within 10% of each other. Only B = 64 is noticeably worse — one block is 64×32 = 2 KB for float8, which blows past L1 line-fill granularity and breaks the compiler's unrolling heuristics. The lack of dependence on B is itself a red flag: if AoSoA were working as intended, increasing B should give the compiler bigger SIMD windows and improve Read performance up to the L1 line size. It does not, because the branch in operator++ is the bottleneck, not the block size.

Fixing It: Lambda API Instead of Iterator

The first version of this article ended right here, with the conclusion "AoSoA is a hand-written kernel primitive, not a container." I published that, then went back and tried to fix the iterator anyway. The fix turned out to be a one-liner redesign of the API.

The problem with for (auto p : aosoa) is that iteration is driven by operator++, and operator++ cannot avoid the modulo-B branch — the iterator has to know when it has hit the end of a block. But if the library drives the iteration internally, it can write a clean two-level loop that the compiler can unroll and SIMD-ize:

template<class F, size_t... Is> void for_each_impl(F&& f, std::index_sequence<Is...>) { const size_t tail = size_ % B; const size_t full = (tail == 0) ? blocks.size() : blocks.size() - 1; for (size_t bi = 0; bi < full; ++bi) { auto& blk = blocks[bi]; for (size_t i = 0; i < B; ++i) { // B is constexpr → unrolled f(std::get<Is>(blk.data)[i]...); // N refs, no tuple } } // tail block handled separately, outside the hot loop }

Notice what is different. The inner loop has a compile-time constant bound (i < B), the outer loop walks blocks sequentially with no modulo arithmetic, and the lambda f receives the field references as separate parameters (via pack expansion), so it never sees a tuple at all. GCC inlines f, sees a clean straight-line body, and emits exactly what it would for SOA — vaddps, vfmadd132ps, no permutes, no spills.

User code now looks like:

// Read: sum all fields float sum = 0; aosoa.for_each([&](auto&... xs) { sum += (xs + ...); }); // Compute: field<0>*field<1> + sum of rest aosoa.reduce(0.0f, [](float acc, auto& a, auto& b, auto&... rest) { return acc + a * b + (rest + ...); }); // FilterCopy: keep elements where f0 < f1 auto kept = aosoa.filter([](auto& x, auto& y, auto&...) { return x < y; });

Same shape as Kokkos parallel_for, same shape as Thrust for_each, same shape as std::ranges::for_each — except we control the block traversal, so we get vectorization that a general-purpose iterator cannot. I kept the old element-level Iterator in the class as a reference point; the new surface is for_each / for_each_indexed / for_each_field<I...> / reduce / filter, plus a for_each_block(f) escape hatch for power users who want to write their own block traversal.

Act 2: The Numbers

At 1M float8 elements, speedups vs AOS:

OperationSOAAoSoA (old iterator)AoSoA (for_each API)
Read2.52x0.83x1.31x
Write0.79x0.70x1.05x
Compute2.40x0.87x1.21x
FilterCopy0.27x0.82x1.33x
Merge0.99x0.99x≈

The new AoSoA is faster than AOS on every operation tested, faster than SOA on Write and FilterCopy, and within 1.5-1.9x of SOA on the SOA-wins operations (Read, Compute). The old iterator-based AoSoA lost to AOS on 4 of 5 ops; the new lambda-based AoSoA beats AOS on 4 of 5 ops.

The picture across sizes is the real confirmation:

In-cache (1K, 4K, 32K), v2_Read matches SOA_Read to within 1% — 391 ns vs 395 ns at 1K, 13541 ns vs 13750 ns at 32K, 1689 ns vs 1680 ns at 4K. The iterator tax is gone entirely. At 262K and 1M, the working set exceeds L2 and L3 respectively, the workload becomes memory-bandwidth-bound, and the SOA advantage from its multi-stream access pattern reasserts itself — but AoSoA v2 now trails SOA by ~1.6-1.9x at 1M, down from the old iterator's 3.0x.

That remaining 1.6x at 1M is not an abstraction cost. It was already there in the direct_Read handwritten benchmark I used to diagnose the problem. It is a property of the block layout versus SOA's multi-stream DRAM access pattern, and it only appears when the working set leaves L3. In-cache, the two are indistinguishable.

What this fixed in my takeaway: AoSoA behind a lambda-driven API is a real drop-in container. You do not have to hand-write block kernels. You do have to give up element-at-a-time operator++ iteration, which C++ range-based for depends on — but you get a cleaner API in exchange. aosoa.for_each([&](auto&... xs){ ... }) is arguably easier to read than for (auto p : aosoa) { ... p.get<0>() ... } anyway.

Act 3: Tuning B — What Is The Right Block Size?

The numbers above all used B=8, because that is the AVX2 register width for float and seemed like the obvious choice. But B is a compile-time parameter, and there was no reason to believe the obvious choice was the best one. So I swept B ∈ {2, 4, 8, 16, 32, 64, 128} across all 4 type configs × 4 operations × 5 sizes = 560 benchmark points, and asked three questions:

  1. Is there a single B that minimizes the average slowdown vs an oracle that picks the best B per workload?
  2. How does performance actually depend on B for each operation?
  3. How often does AoSoA v2 at its best B beat SOA?

Finding the optimal default

For each workload (cfg × op × size), I record time(B) / min(time(B')) — the penalty you pay for fixing a default B vs an omniscient oracle. Then I aggregate across all 80 workloads with geomean, median, and worst-case.

BGeomean penaltyMedianWorst case
21.55x1.364.34
41.26x1.162.28
81.12x1.062.42
161.10x1.031.79
321.12x1.081.91
641.17x1.131.82
1281.23x1.162.28

B=16 is the optimum on all three metrics. Geomean 1.097x means "on average, fixing B=16 costs you 9.7% vs the best B for that specific workload." More importantly, the worst-case is 1.79x — every other B has a ≥1.82x worst case, which is the regret you want to avoid. B=16 is the only choice that keeps the pathological-workload penalty under 2x.

Why B=16 — the per-op picture

Plotting the penalty curve per operation makes the shape clear. All four ops have a plateau optimum centered on B=16:

  • Read prefers B=16 (1.10), with B=8/32/64 within 2-3%.
  • Write prefers B=16 (1.08), with B=32 at 1.10.
  • Compute is almost flat between B=8, 16, 32 (1.079, 1.088, 1.086 — within 1%). B=32 is technically the tightest fit but the gap is noise.
  • FilterCopy prefers B=8 (1.109), with B=16 at 1.119 (1% worse).

The sweet spot is {8, 16, 32}. Outside it, B=4 and B=64 are 10-20% off, B=2 and B=128 are catastrophic. Within it, B=16 is the Pareto point — never the absolute best on every op, but never more than 1% off either, while any other single choice (like my original B=8) gives up 3-5% on Read and Write.

What the 1M picture actually looks like

The per-configuration curves at 1M are surprisingly varied. float3 Read climbs almost linearly from 0.9x at B=2 to 2.2x at B=128 — the DRAM prefetcher loves long sequential streams of a single small field, and larger blocks give it more to chew on. int+float+double Write peaks at B=16-64 around 2.2x vs AOS, then drops. float8 Compute is almost flat — this is the DRAM-bound workload where SOA dominates, and block size barely matters. FilterCopy has the clearest peak: all four configs form a bell curve centered on B=8-16, with the tail collapsing past B=64 (large blocks → more wasted pushback when the predicate is selective).

The mixed shapes are why no single B is universally best. What makes B=16 the right default is that it sits on the flat part of every curve — you never drop a cliff by picking it.

Final win rate

Across all 80 (config, op, size) cells, using the best B per cell:

  • v2 beats SOA by ≥5% on 57/80 cells (71%)
  • Tie (within ±5%) on 14/80 (18%)
  • SOA wins by ≥5% on 9/80 (11%)

The 9 SOA wins are almost entirely concentrated in the right column (1M) for Read/Compute on float8 and double3. That is the DRAM-bound regime where SOA's multi-stream access pattern matters more than in-block locality. Everywhere else — in cache, on Write, on FilterCopy, on small-field configs — AoSoA v2 with a tuned B is the fastest layout I have.

The practical recommendation: ship template<typename... Ts> using AoSoAd = AoSoA<16, Ts...> as the default alias and stop worrying about B unless you are profiling a specific hot loop. The extra 2-3% you could theoretically gain from per-workload tuning is not worth the maintenance cost of per-workload compile-time dispatch.

Act 4: Going Deeper — Where Does The Remaining 1M Gap Come From?

The 11% of cells where AoSoA v2 still loses to SOA are all read-heavy workloads at N=1M on float8/double3 — the DRAM-bound regime. Before accepting that as "layout-intrinsic" I wanted to understand exactly what was happening, and see whether any more aggressive optimization could close it. So I ran four parallel investigations — three on optimization, one on pure diagnosis.

Diagnosis — it is not codegen, it is memory-level parallelism

I compared the SOA and AoSoA v2 hot loops at float8/1M with objdump -dC. The two loops are identical in op count: 16 vaddps ymm, 2 vmovups, 14 shuffle-family ops per iteration, 128 floats processed, zero stack spills. The only visible difference is SOA keeps 8 base pointers alive in registers (r13, rbx, r11, r10, r9, r8, rdi, rcx — one per field array), while AoSoA walks 1 base pointer (rsi) through a monotonic stream of 512-byte blocks.

Measuring wall time directly on this machine (AVX2 only, 12 MiB L3, single core):

SOA_Read float8/1MAoSoA16_v2_Read float8/1M
Time793 µs1,505 µs
Bandwidth40.3 GB/s21.2 GB/s
Cycles/element4.187.82

SOA achieves 40 GB/s, which is essentially the single-core DRAM ceiling on Alder Lake / Arrow Lake. AoSoA achieves 21 GB/s. Applying Little's Law to explain the gap:

bandwidth ≈ (outstanding_lines × 64 B) / DRAM_miss_latency
         ≈ outstanding_lines × 64 / 80 ns

Intel's L2 streamer tracks ~16-32 streams and prefetches up to ~20 cache lines ahead per trained stream. SOA with 8 concurrent streams gets 8 × 20 = ~160 outstanding lines → ~128 GB/s ceiling, throttled by DRAM to 40 GB/s. AoSoA with 1 stream gets ~20 outstanding lines → ~16 GB/s ceiling, observed 21 GB/s with some help from L1 DCU adjacent-line prefetch.

The gap is pure memory-level parallelism. Not codegen. Not cache misses. Not alignment. Not the lambda protocol. The HW prefetcher is doing its job correctly on a single stream — it just cannot train enough lookahead from one stream to saturate DRAM.

This also explains a mystery from Act 3: why AoSoA wins on float3 Read (0.66x SOA — 1.52x faster). For float3, SOA has only 3 streams. 3 × 20 = 60 outstanding lines → smaller MLP advantage. AoSoA's single stream amortizes over a longer block-resident run, and the layouts cross. The "AoSoA loses to SOA on Read" story is therefore not "AoSoA is slower on Read" — it is "SOA's MLP advantage scales with its stream count, and float8 is exactly where SOA has the most streams."

Unrolling the block loop — it helps at L3 but not at DRAM

If the problem is "only 1 stream", can we present the HW multiple streams by processing multiple blocks per outer iteration? Added for_each_unrolled<UNROLL>:

for (; bi + 4 <= full; bi += 4) { // 4 back-to-back copies of the baseline element loop process(blocks[bi]); process(blocks[bi+1]); process(blocks[bi+2]); process(blocks[bi+3]); }

Written as a std::index_sequence fold with a lambda per copy — that shape lets GCC vectorize each copy independently while keeping the 4 blocks' loads interleaved enough that OoO execution sees 4 virtual streams. If the compiler merges them into one big loop, vectorization collapses to scalar vaddss; if it keeps them separate, it's fine. The lambda trick is the one that works.

Result: UNROLL=4 closes the gap at L3-resident sizes (262K) — from 0.79x SOA to 1.09x (v2 now faster than SOA). In-cache it is at parity with the baseline. At 1M it does not help — actually slightly regresses. The 1M bottleneck is not "too few virtual streams" after all; unrolling gives the DRAM subsystem 4 streams per iteration and it still does not close the gap. Either the L2 streamer does not treat 4 close addresses as 4 streams, or the bottleneck at 1M has moved to TLB / DRAM-row-switching, which unrolling cannot fix.

SIMD intrinsics — separate accumulators break a hidden dep chain

The second angle: hand-write the inner loop with AVX2 intrinsics and 8 separate __m256 accumulators (one per field). Why would this help if the codegen is already clean?

Because the lambda-driven for_each produces this pattern after inlining:

sum += (get<0>(blk)[i] + get<1>(blk)[i] + get<2>(blk)[i] + ... + get<7>(blk)[i]);

GCC turns this into vaddps reductions that sum 8 fields into a single accumulator. That's a 7-deep dependency chain per element. On ADDPS with latency 4 / throughput 0.5, the dep chain caps per-element throughput at 4/7 ≈ 0.57 cycles even when load bandwidth is plentiful. Hand-written code with 8 independent __m256 accumulators breaks that chain — each field's vector add is independent of the others, and the CPU can issue all 8 per iteration.

Added sum_all_f32_avx2() and compute_all_f32_avx2() as opt-in methods guarded by requires (is_float_only_B16) (float-only, B=16). Each one uses per-field accumulators plus software prefetch of the next block at distance 8 (4 KiB ahead).

float sum_all_f32_avx2() const requires (is_float_only_B16) { constexpr size_t N = sizeof...(Ts); __m256 acc[N]; for (size_t f = 0; f < N; ++f) acc[f] = _mm256_setzero_ps(); for (size_t bi = 0; bi < full; ++bi) { if (bi + 8 < full) { const char* next = (const char*)&blocks[bi + 8]; _mm_prefetch(next, _MM_HINT_T0); _mm_prefetch(next + 64, _MM_HINT_T0); // ... 8 cache lines of the next-ahead block } sum_block_avx2_impl<0>(blocks[bi], acc); } __m256 total = _mm256_setzero_ps(); for (size_t f = 0; f < N; ++f) total = _mm256_add_ps(total, acc[f]); return hsum256_ps(total); }

Final numbers

Float8, B=16, speedup vs SOA (higher is better):

Sizev2 baselinev2 unrolled (u4)v2 AVX2 intrinsics
1K (L1-resident)1.02x1.02x2.38x
4K (L1/L2)1.01x1.01x2.01x
32K (L2-resident)1.00x1.02x2.00x
262K (L3-resident)0.79x1.09x1.04x
1M (DRAM-bound)0.57x0.50x0.64x

Compute shows an even stronger pattern:

Sizev2 baselinev2 unrolled (u4)v2 AVX2 intrinsics
1K1.02x1.01x2.71x
32K1.01x0.99x2.03x
262K0.82x1.08x1.08x
1M0.58x0.48x0.68x

What this means

The surprise is the in-cache result. The AVX2 hand-written path is 2-2.7x faster than SOA at L1/L2 sizes. Not 2-2.7x faster than AoSoA. Faster than SOA. This is not because AoSoA is magically fast — it is because SOA's own scalar reduction is slow. The SOA Read benchmark folds 8 fields into one scalar per element, which GCC compiles to the same 7-deep ADDPS dep chain. The SOA loop saturates neither its load ports nor its ADD ports; it is latency-bound on its own reduction. Writing that reduction by hand with 8 accumulators accelerates any float8-sum workload, regardless of layout. The AoSoA container happens to be the right place to expose the opt-in method because it's where people benchmarking float8 reductions will look.

At the L3 boundary (262K), both unrolled and avx2 cross SOA and become the faster layout. The unrolled variant wins slightly (1.09x vs 1.04x) — it introduces MLP without adding uop pressure, exactly matching what the L3-resident regime needs.

At 1M the DRAM wall holds — or so it seemed. AVX2 intrinsics close ~36% of the residual gap (from 0.57x to 0.64x) but don't cross SOA. Unrolling actively makes 1M worse (the extra register pressure and instruction footprint hurt when memory is the bottleneck). At the end of Act 4, the conclusion was "no software trick tested closes the float8 1M gap" — and that conclusion held for every experiment up to that point. Act 5 will prove it wrong by identifying the one trick Act 4 did not test: splitting the block vector into K far-apart segments to create K virtual streams, which is very different from unrolling K adjacent blocks.

Prefetching was a dead end

One experiment failed instructively. I tried manual software prefetching with a sweep of distances (2, 4, 8, 16, 32 blocks ahead) and locality hints (_MM_HINT_NTA, T0, T1, T2), both single-line and all-lines per block. Every variant was 2x slower than the baseline across every size — including in-cache. The overhead of the prefetch uops themselves (8 per block when prefetching every field, 1 per block otherwise) plus the per-block conditional address computation was net negative. Prefetching only helped when embedded inside the AVX2 intrinsic loop where the prefetch instructions overlap with arithmetic — as an isolated optimization layered on top of the scalar for_each, it is useless. The "software prefetch fixes MLP" intuition is wrong on this codegen: the HW prefetcher is already at its single-stream ceiling, and adding SW prefetches can't multiply it into more streams.

Similarly, non-temporal loads (_mm256_stream_load_si256) were tested and rejected. Intel SDM Vol 3A §11.6.2 specifies that vmovntdqa on write-back memory — the default for allocator-provided storage — degrades to a regular load with a "no write-combining" hint. On this chip the NT variant is 5x slower than plain AVX2 at 1M Read. NT loads are for true streaming workloads that write to a separate region and never re-read; a benchmark reading the same container across iterations is not one.

Act 5: Multi-Stream — The 1M Wall Cracks

Act 4 ended with "no software trick tested closes the float8 1M gap." That turned out to be wrong. The trick existed — I had just not tested the right one.

The missed experiment

Every optimization in Act 4 processed blocks in adjacent order: baseline walks blocks[0], blocks[1], blocks[2]...; unrolled-4 walks blocks[0..3], blocks[4..7]...; the AVX2 intrinsic version adds prefetches but still walks sequentially. All of these produce one long stream from the hardware memory system's point of view — one monotonic address, trained by one L2 streamer tracker slot, giving ~20 cache lines of lookahead.

SOA, by contrast, walks eight field arrays stored in separate allocations. The L2 streamer sees 8 distinct monotonic streams and trains 8 trackers, giving 8 × 20 = ~160 cache lines of lookahead — enough to saturate DRAM bandwidth on this core. Act 4 diagnosed this correctly but concluded that AoSoA's single-stream access pattern was inherent to the layout.

It is not. You can fake multiple streams out of a single vector by splitting it into K segments and walking K blocks in lock-step, one from each segment. Each segment is total_blocks / K blocks apart — for 1M float8 elements with B=16, that is 62500 blocks = 32 MB, so K=2 gives streams 16 MB apart, K=4 gives 8 MB apart, K=8 gives 4 MB apart. All four are far enough that the L2 streamer cannot fuse them into one stream; it has to allocate K independent tracker slots.

Added as an opt-in method:

template<size_t K, class F> void for_each_multistream(F&& f) { const size_t seg = full / K; for (size_t bi = 0; bi < seg; ++bi) { [&]<size_t... Ks>(std::index_sequence<Ks...>) { (( [&] { auto& blk = blocks[bi + Ks * seg]; for (size_t i = 0; i < B; ++i) { f(std::get<Is>(blk.data)[i]...); } }() ), ...); }(std::make_index_sequence<K>{}); } // ... leftover blocks + tail }

The Ks fold expands into K back-to-back lambda invocations per outer iteration, each one accessing a different segment's current block. The compiler vectorizes each copy independently and the CPU issues them in parallel.

Results at 1M

float8 Read at N=1M, speedup vs SOA (higher = faster):

VariantRatio vs SOA
AoSoA v2 baseline0.57x
AoSoA v2 unrolled (u4)0.45x
AoSoA v2 AVX2 intrinsics0.67x
AoSoA v2 multi-stream k=20.68x
AoSoA v2 multi-stream k=40.75x
AoSoA v2 multi-stream k=80.76x

Multi-stream with K=8 closes the gap from 1.78x slower (baseline) to 1.32x slower (ms_k8) — the residual gap shrank by 60%. Compute shows an even stronger effect:

Variant @ 1MRatio vs SOA
AoSoA v2 baseline0.60x
AoSoA v2 AVX2 intrinsics0.72x
AoSoA v2 multi-stream k=80.84x

From 1.67x slower to 1.19x slower — a 72% gap reduction.

And at L3-resident sizes, the multi-stream variant actually beats SOA:

At N=262K (8 MB working set, L3-resident):

  • Read: baseline 0.74x → ms_k8 1.25x (AoSoA 1.25x faster than SOA)
  • Compute: baseline 0.79x → ms_k8 1.23x (AoSoA 1.23x faster than SOA)

At L3, the prefetcher has enough bandwidth headroom that each of the K=8 streams gets excellent lookahead, and the extra in-flight lines effectively hide the L3-DRAM latency. At 1M the absolute DRAM bandwidth starts to matter and the multi-stream advantage partially saturates, but it still wins.

Why this was a blind spot

The K-sweep curve is monotonic. K=1 is the baseline. K=2 closes ~30% of the gap. K=4 closes ~60%. K=8 closes ~68%. Going from 1 stream to 8 streams — every step matters. Act 4's unrolled variant for_each_unrolled<4> tried "4 adjacent blocks per iteration" and did not help at 1M; the assembly-level read (Act 4 diagnosis) suggested MLP was the bottleneck. What I missed was the difference between 4 adjacent blocks (1 stream, longer run) and 4 far-apart blocks (4 streams). The L2 streamer cares about address trajectories, not about how many cache lines you touch per iteration. Unrolling adjacent blocks does not create new streams; splitting the vector does.

Agent J's static assembly analysis (Act 4) predicted this result before the experiment: "Stream in 2 blocks at a time from interleaved bases — the L2 streamer sees 2 independent streams, which it handles better than 1 stream with embedded sub-stride. This is effectively what the new for_each_multistream<K> experiment is doing — this analysis predicts it should work, assuming K≥2 multistream and base pointers are ≥4 KiB apart." The assembly analysis was right; the prediction checked out empirically.

The residual 1.19-1.32x at 1M

Even with 8 virtual streams, AoSoA does not cross SOA at 1M. The gap shrank from ~1.7x to ~1.2x but did not close. Three plausible causes:

  1. TLB pressure. AoSoA walks 1 contiguous 32 MB region, not 8 separate 4 MB regions. With 4 KB pages, that is 8192 DTLB entries needed; the L1 DTLB holds only ~64. Huge pages (2 MB each) would shrink this to 16 entries, fitting comfortably. A separate experiment tried to enable THP via madvise(MADV_HUGEPAGE) and a custom mmap allocator with MAP_POPULATE, but AnonHugePages stayed at 0 throughout — khugepaged refused to promote due to physical-memory fragmentation and short-lived workloads. On this system, transparent huge pages for short anonymous allocations is effectively unavailable without root privilege. The TLB hypothesis cannot be formally tested from user space, and the remaining gap may well be TLB walks.

  2. DRAM bank conflicts. AoSoA's single 32 MB walk hits DRAM banks in strict linear order, while SOA's 8 separate heap allocations end up at 8 different physical base addresses that typically map to 8 different bank groups. Bank-level parallelism is inherently better with multiple disjoint allocations; multi-stream partially recovers this but not fully, because all K segments live in the same std::vector backing store.

  3. Something the L1 IP prefetcher can't learn. Agent J's disassembly noted that inside a Block, the 8 field loads come from 8 different instruction pointers all hitting the same cache line at stride-64 once, then never again — the L1 IP-based stride prefetcher has nothing to train on. SOA's layout presents each IP with a clean contiguous stream. This is a property of the Block layout itself, not of the traversal order; multi-stream does not fix it.

Any combination of these could explain the residual gap. In practice it does not matter much: AoSoA v2 ms_k8 is within 20% of SOA at 1M Read and within 19% at 1M Compute, and crosses SOA at every in-cache size. The original "AoSoA loses on Read at 1M" story from Act 1 is now "AoSoA loses on Read by less than 20% at the worst case, and wins everywhere else."

What this adds to the API

for_each_multistream<K> is a third opt-in traversal alongside for_each_unrolled<N> and sum_all_f32_avx2. Recommendation:

  • K=1 = baseline for_each (default). Use in cache or for small N.
  • K=4 or K=8 when the working set is L3-resident (256K-2M elements). This is where the multi-stream advantage is largest.
  • K=8 for DRAM-bound workloads (>2M elements). Partially closes the residual gap.
  • sum_all_f32_avx2 / compute_all_f32_avx2 still win by 2-3x in L1/L2, and are the right choice when you know the working set fits in cache.

The underlying container is unchanged; these are all view/traversal methods. You pick the shape of the loop at the call site.

A Concrete Case Study: An N-Body Simulation Frame

Every benchmark up to this point has been an isolated primitive: Read, Write, Compute, FilterCopy, Merge. Real code never runs one primitive in isolation — it runs a mix that touches the same data in different ways. The layout that wins on isolated Read is often not the layout that wins on the realistic mix.

To see what that looks like, I built a tiny particle simulator. Each particle has 8 fields — position (x, y, z), velocity (vx, vy, vz), mass, and life — so 32 bytes each, same shape as the float8 benchmarks above. A "frame" does three things every time step:

  1. Integrate — x += vx*dt; y += vy*dt; z += vz*dt; life -= dt (reads 3 velocity fields + 4 mutated fields, writes 4)
  2. Kinetic energy — total += 0.5 * mass * (vx² + vy² + vz²) (reads 4 fields into a single scalar accumulator)
  3. Cull dead — copy particles where life > 0 into a new container (whole-object filter; reads all 8 fields per particle, writes 8 per survivor)

This is the smallest interesting mix: one read-heavy-but-mutating pass, one read-only reduction, and one whole-object move. A real physics engine does much more, but these three exercise the three patterns that each layout has a different opinion about.

I implemented the frame three times: once with std::vector<Particle> (AOS), once with SOA<float × 8> using hand-written per-field loops, and once with AoSoA<16, float × 8> using for_each / reduce / filter. Same data, same dt, same cull ratio (~1% per frame).

Results — full frame timing

CPU median time per frame (µs), measured on Intel Core Ultra 7, single core:

ParticlesAOSSOAAoSoA
10,00071 µs217 µs (3.04× slower)52 µs (1.37× faster than AOS)
32,768243 µs752 µs (3.10× slower)183 µs (1.33× faster than AOS)
262,1442.73 ms6.60 ms (2.41× slower)2.15 ms (1.27× faster than AOS)
1,000,00010.11 ms26.31 ms (2.60× slower)14.00 ms (1.39× slower than AOS)

SOA is the worst layout at every size

This is the finding that flipped my mental model. In the isolated benchmarks, SOA was "2x faster on Read, 2x faster on Compute, catastrophic on FilterCopy." I expected the realistic mix to give SOA a partial win — SOA's advantage on integrate and kinetic energy would offset some of its loss on cull. That is not what happens. SOA is 3x slower than AOS on every size tested, because the cull_dead step is so slow it eats everything else alive.

What's going on: when SOA has to copy a surviving particle, it calls push_back on 8 separate vectors. That is 8 allocator checks, 8 size updates, 8 writes to different cache lines, 8 tail pointers to bump. AOS calls push_back once on a single vector of structs — 1 allocator check, 1 size update, 1 memcpy of 32 bytes into a contiguous slot. For a 1% cull rate across 1M particles, that is 10,000 surviving copies. SOA's filter step takes 17 ms out of its 26 ms frame; AOS's takes 5 ms out of 10 ms. The 2x-faster reduction and the 2x-faster integrate are irrelevant once you put them next to the 3.4x-slower filter, because the filter is the slowest step.

"SOA is faster" turns out to mean "SOA is faster at reading and computing if you never move objects around." The moment you mix in whole-object operations, SOA collapses. On a real frame that does even one filter or sort per iteration, SOA is the wrong layout.

AoSoA wins at every size except DRAM-bound 1M

At 10K, 32K, and 262K particles, AoSoA's full frame is 1.27× to 1.37× faster than AOS — which is the best layout between those two. The win comes from two directions: the integrate and kinetic steps benefit from in-block SIMD vectorization (the same for_each path that matched SOA on isolated Read), and the filter step keeps AOS-like locality because AoSoA's filter copies contiguous blocks of B=16 elements at a time into the destination.

At 1M particles the story changes. The working set is 32 MB, well beyond the 12 MiB L3, so all three layouts are DRAM-bandwidth-bound. In this regime AOS's simple single-pass copy during cull is the most bandwidth-efficient option, and AoSoA's block-indirected filter starts to pay for the extra bookkeeping per surviving particle. AoSoA at 1M is 1.39× slower than AOS — still 1.88× faster than SOA, but no longer the winner.

Multi-stream doesn't change the picture here. for_each_multistream<8> only affects the integrate step (the filter path is unchanged), and integrate is ~40% of the frame. Applying multi-stream gives a tiny speedup on integrate that does not move the overall frame time.

What this means in practice

Isolated-benchmark reasoning about layouts is a trap. The decision you actually need to make — "which container should my particle system use?" — depends on the mix of operations per frame, weighted by how much time each takes. Here is what the three layouts are actually best at when you integrate that mix honestly:

  • Pure read-heavy streaming (isolated Read at 1M, like a blur filter on a 32 MB image) → SOA wins by 2x. No filter, no sort, no move. The textbook case.
  • Mixed read + whole-object move (physics frame, particle system, game entity loop) → AoSoA wins at every size ≤ L3, AOS wins at sizes larger than L3. SOA loses everywhere.
  • Heavy filtering or sorting (CSV processing, ECS archetype swap) → AOS wins straight up, no contest.

The original "SOA vs AOS" debate assumed workloads were one of the first two patterns. Most code is actually the third, which SOA is worst at. AoSoA v2 is the layout that does not make this trade-off — it is strictly better than both SOA and AOS on realistic mixed workloads up to the DRAM boundary.

But wait — is this fair to SOA?

Good objection. The cull step uses push_back on every surviving particle, and that is specifically the operation where SOA's 8-vector design pays the cost 8 times instead of once. If I remove the filter and look at the pure streaming case — just integrate and kinetic energy, no whole-object move — SOA should get its natural win back. So I ran the same frame with the cull step removed, same data, same sizes, same layouts. Here is what the pure read+compute frame looks like:

Pure frame (integrate + kinetic energy, no filter), speedup vs AOS:

ParticlesAOSSOAAoSoAAoSoA_ms
10,0001.00x2.26x4.76x4.69x
32,7681.00x2.24x4.77x4.55x
262,1441.00x2.14x2.23x2.40x
1,000,0001.00x2.41x1.04x1.25x

Three things jump out here, and they are the opposite of the with-filter story:

  1. SOA now does what the textbooks promised. 2.14-2.41× faster than AOS at every size, including DRAM-bound 1M. This is the classic "SOA vectorizes better and streams better" win, cleanly.

  2. AoSoA in cache is absurdly fast — 4.77× AOS and 2.13× SOA at 10K and 32K particles. The per-field __m256 accumulators in for_each plus the in-block contiguity mean the CPU issues loads from 8 adjacent 64 B arrays with no dependency chain between them, and the compiler unrolls more aggressively than it does on SOA's vector-of-vectors. This is the same effect that drove the 2-2.7× cache win of sum_all_f32_avx2 in Act 4, but here it comes out of the stock for_each lambda path without any intrinsics. For small particle counts AoSoA is the clear winner.

  3. SOA reclaims the crown at 1M. This is the DRAM-bound regime from Act 5: SOA has 8 independent streams that saturate the L2 streamer's tracker slots; AoSoA has 1 (or K=8 virtual streams with multistream, which does help — AoSoA_ms is 1.25× AOS vs plain AoSoA at 1.04×). The multi-stream variant closes some of the gap but not all of it. For pure streaming at DRAM-bound sizes, SOA is the right choice — and that matches what we found back in Act 5's isolated float8 Read measurements.

Side by side: one filter step changes everything

Both halves are the same data, same particles, same integrate step, same kinetic energy reduction. The only difference is whether the frame ends with a cull_dead step that copies surviving particles into a new container. Left side: with cull. Right side: without. And the verdict about "which layout is faster" is completely different:

  • With cull (left): AoSoA wins up to 262K, AOS wins at 1M, SOA is worst everywhere. The filter step is what happens when whole-object locality matters, and SOA's per-field vector design is catastrophic there.
  • Without cull (right): SOA wins at 1M (2.4× AOS), AoSoA wins in cache by a huge margin (4.8× AOS), AOS is worst everywhere. The pure streaming regime is where SOA's multi-stream prefetching and vectorization shine, exactly as the textbooks predict.

Neither layout is "fastest" in any absolute sense. What wins depends on whether your hot loop ever moves whole objects around. A single filter step, a single sort, a single archetype reassignment, a single pass of spatial hashing — any of these flips the ranking. Most real code has at least one.

AoSoA is the only layout that is close to optimal in both regimes. It is:

  • the fastest layout at in-cache sizes regardless of whether you filter,
  • competitive with SOA at L3-resident sizes with or without filter,
  • strictly better than SOA on filter-heavy 1M workloads,
  • slightly behind SOA on pure-streaming DRAM-bound workloads.

Hold on — was SOA ever pushed the same way?

After publishing the section above, I noticed something that bothered me. Throughout this article I gave AoSoA a series of upgrades — for_each_unrolled, sum_all_f32_avx2 with hand-written intrinsics and 8 independent accumulators, for_each_multistream<8> to expose multiple streams, the L3-tuned variant. SOA, by contrast, was always a hand-written scalar for loop relying on GCC's auto-vectorizer:

for (size_t i = 0; i < n; ++i) { ke += 0.5f * M[i] * (VX[i]*VX[i] + VY[i]*VY[i] + VZ[i]*VZ[i]); }

The auto-vectorizer compiles this into a vectorized loop with one __m256 accumulator, because the source-level reduction has a single scalar variable ke. The same loop-carried dependency chain that I diagnosed inside AoSoA's reduction in Act 4 is present in SOA's reduction too — and I never fixed it for SOA. Every "AoSoA wins by 2× in cache" measurement compared hand-tuned AoSoA against compiler-defaulted SOA. That is not a fair fight.

So I wrote the SOA equivalent of sum_all_f32_avx2: explicit AVX2 intrinsics for both the integrate and the kinetic step, two independent __m256 accumulators in the reduction, manual 16-element unrolling. Same techniques AoSoA got. Same level of hand-tuning. Re-ran the pure frame:

Pure frame, all variants, speedup vs AOS:

ParticlesAOSSOA (auto)SOA (avx2)AoSoA (for_each)AoSoA (multistream<8>)
10,0001.00×2.26×8.81×4.74×4.74×
32,7681.00×2.27×8.17×4.66×4.44×
262,1441.00×1.93×2.43×2.33×2.29×
1,000,0001.00×2.29×2.49×1.05×1.27×

Pushed SOA wins everything. 1.86× faster than AoSoA's best variant at 10K. 1.75× at 32K. Roughly tied at 262K (AoSoA is within 4%). 1.96× faster at 1M. The "AoSoA crushes SOA in cache by 2-2.7×" finding from Acts 4 and 5 was not a property of the layouts at all — it was a property of how much hand-optimization each one received. As soon as SOA gets the same treatment, the cache-side advantage flips and SOA is 1.7-1.9× faster than AoSoA on pure read+compute.

This is the kind of correction that matters. The headline of the blog should not be "AoSoA is the best layout for in-cache reductions" — that turned out to be wrong. It should be: for pure streaming workloads with no whole-object moves, hand-tuned SOA is the fastest layout, regardless of working-set size. Auto-vectorized SOA leaves a factor of ~3 on the table at L1/L2 sizes because the compiler will not break the reduction dep chain on its own. Hand-write the reduction and the layout's natural advantages come back.

So when does AoSoA actually win?

The Frame benchmark (with the cull step) is where the AoSoA advantage stays real. There the layout matters because of the whole-object move, which SOA cannot optimize away with intrinsics — it would need to do 8 push_backs no matter how clever the inner loop is. The number of allocator touches per surviving particle is the bottleneck, and that is layout, not codegen.

The honest summary across the two case studies:

  1. Pure streaming workload (no filter, no sort, no whole-object move) → hand-tuned SOA is the fastest at every size. Auto-vectorized SOA leaves 3-4× on the table at L1/L2. AoSoA's "cache win" was an artifact of unfair comparison.
  2. Mixed workload with whole-object operations (real physics frames, ECS systems, anything that filters or sorts) → AoSoA wins at L1/L2/L3 because the filter step keeps AOS-like locality. AOS wins at DRAM-bound 1M because its single-vector layout has the most efficient bulk copy.
  3. SOA loses any time the workload moves objects. This is the layout's structural weakness — 8 independent vectors are great for streaming and terrible for whole-object moves.

What I was testing in Acts 4 and 5 ("can AoSoA match SOA on Read?") had the wrong premise. The real question is "can any layout do both Read and FilterCopy fast?" — and AoSoA is the only one that comes close to yes, even if hand-tuned SOA beats hand-tuned AoSoA on the Read side alone.

What I Would Recommend

Here is the practical decision table I would give someone asking "which layout?":

Use SOA when:

  • You iterate over one or a few fields at a time (hot loops rarely touch all fields)
  • You have many fields per struct (>4) — the waste in AOS compounds
  • You do read-heavy, compute-heavy, or SIMD-friendly work
  • Your working set is larger than L1 — prefetching helps
  • You do linear searches on a single field

Use AOS when:

  • You frequently copy, filter, or move whole objects between containers
  • You merge/sort data based on whole-object comparisons
  • Your hot path touches all fields of a struct together
  • Your structs are small (<16 bytes) and fit in one cache line
  • Your working set fits comfortably in L1 and you do short bursts

Consider a hybrid. Real systems often use SOA for simulation state (updated every frame) and AOS for serialization or batch operations (writing to disk, copying between subsystems). The SOA_iterator proxy lets you write both styles without code duplication.

Consider AoSoA with a lambda-driven API if your workload is mixed. The rewritten for_each/reduce/filter surface matches SOA in cache, beats AOS on every operation tested, and keeps AoSoA's FilterCopy advantage over SOA (5x faster). The only caveat is that you give up for (auto p : aosoa) syntax — the element-level iterator cannot be fixed without reintroducing the branch that broke it.

Reach for for_each_multistream<8>(f) whenever the working set is larger than L2. The multi-stream variant is the single biggest optimization I found — at L3-resident sizes it crosses SOA by 25% on float8 Read and Compute, and at DRAM-bound 1M it closes 60-72% of the residual gap. It requires zero intrinsics and is portable C++20. K=8 is the sweet spot on this hardware; K=4 is within a few percent. Use it as a drop-in replacement for for_each anywhere you would have wanted "SOA-style prefetcher behavior on an AoSoA container."

for_each_unrolled<4>(f) is subsumed by multi-stream but kept for experimentation. Unrolling adjacent blocks creates a longer run of one stream; multi-streaming creates K genuinely parallel streams. The unrolled variant is harmless but superseded by the multi-stream one at every size past L2.

For hot float-only reductions, hand-write the reduction with multiple accumulators — and do it on whichever layout you have. sum_all_f32_avx2() on AoSoA was 2-2.7× faster than auto-vectorized SOA, but the same trick applied to SOA itself (8 independent __m256 accumulators in the kinetic loop) makes SOA 1.7-1.9× faster than AoSoA's intrinsic path in cache. The win is the multi-accumulator pattern, not the layout. If you're going to reach for AVX2, write the reduction directly on raw arrays — SOA is the cleaner backing store for that pattern.

What I Took Away

  1. "SOA is faster" is not a rule. It is a rule-of-thumb with two directions. For column-oriented access it wins 2-10x. For row-oriented object movement it loses 1.5-3x. The rule depends on what your code is doing.

  2. Field count is the biggest multiplier. Going from 3 to 8 fields turned a 1.1x speedup into a 3.4x speedup (10x on machines with larger L3). If your struct has many fields and you only use a few, the savings are proportional.

  3. Cache hierarchy creates non-monotonic scaling. The SOA advantage does not decay smoothly — it dips at the L2/L3 boundary. Benchmark across sizes to understand your workload, not just at one size.

  4. Zero-cost abstraction is not transitive — but it is fixable. The proxy iterator is free for SOA because operator++ is one integer increment. The same proxy iterator applied to AoSoA is not free, because the modulo-B branch in operator++ defeats auto-vectorization. The fix is to move iteration into the library via a lambda-driven for_each — the library writes the clean two-level block loop that the iterator could not. The abstraction survives; the surface just shifts from operator++ to lambdas.

  5. Who drives the loop matters. "User calls ++it" and "library calls f(refs...)" look equivalent from outside, but the compiler sees entirely different shapes. The first gives you C++ range-based for; the second gives you vectorization. On AoSoA specifically, you cannot have both — the geometry of the layout forces a branch that any element-level iterator must pay. Kokkos, Thrust, and SIMD ISPC all made this choice for the same reason.

  6. AoSoA with the right API is the best general-purpose layout on this workload. Read, Write, Compute, and FilterCopy at 1M float8 all beat AOS; FilterCopy beats SOA by 5x; Read matches SOA in cache. At the best B per workload, AoSoA v2 beats SOA on 71% of cells across the 80-cell sweep and ties another 18%. The only losses are Read/Compute at 1M where the workload is DRAM-bound and SOA's multi-stream access pattern wins — and those are a layout/bandwidth property, not an abstraction cost.

  7. Block size B has a plateau optimum, not a peak. Sweeping B ∈ {2, 4, 8, 16, 32, 64, 128} across all 80 workloads shows a flat minimum on {8, 16, 32} and sharp cliffs outside. B=16 is the best single default — geomean penalty 1.10x vs oracle, worst case 1.79x, the only choice that keeps the worst case under 2x. Picking B=8 (AVX2 width) costs 3-5% on average; picking B=64 costs 17%. The "SIMD register width" intuition is wrong; the compiler fuses multiple registers per field quite happily up to B=32.

  8. The 1M gap is a single-stream MLP ceiling, not a codegen issue. Reading hot-loop assembly confirmed SOA and AoSoA emit identical op counts (16 vaddps, 2 vmovups, 14 shuffles, zero spills). SOA wins because it walks 8 base pointers in parallel and the L2 streamer trains ~20 lines of lookahead per stream — 8 streams × 20 lines ≈ 160 outstanding lines, enough to saturate DRAM. AoSoA has 1 stream → ~20 outstanding lines → bandwidth-capped at ~half of SOA. No software trick (prefetching, unrolling, intrinsics) closed this. The fix is more cores, not more code.

  9. Hidden reduction dep chains are worth more than layout choice. Writing sum += (x0 + x1 + ... + x7) in either SOA or AoSoA compiles to a vaddps chain with a single accumulator that caps throughput at vaddps latency, regardless of bandwidth. Rewriting the reduction with 2-8 independent __m256 accumulators makes both layouts 3-4× faster in cache. The trap I fell into was hand-tuning AoSoA's reduction (Act 4) and then comparing it to compiler-defaulted SOA — that comparison made AoSoA look 2-3× faster in cache, but it was measuring "hand-written vs auto-vectorized", not "layout vs layout". Once SOA gets the same multi-accumulator AVX2 treatment, SOA is 1.7-1.9× faster than AoSoA at L1/L2 sizes on pure read+compute. The dep-chain fix is an optimization the compiler will not do for you on either layout, and it dwarfs any layout effect.

  10. "Unrolled by N blocks" and "N far-apart streams" are not the same thing. Unrolling adjacent blocks in the outer loop gives you one long stream that the L2 streamer trains as a single tracker — no extra memory-level parallelism. Splitting the block vector into K segments and walking K blocks in lock-step gives you K independent streams that train K separate tracker slots — genuine MLP. At 1M float8, this is the difference between no improvement (unrolled) and closing 60-72% of the gap (ms_k8). The distinction looks trivial from source code but is everything to the memory subsystem.

  11. Multi-stream is the best single-core AoSoA optimization I found. At L3-resident sizes (262K) for_each_multistream<8> actually beats SOA by 25% on float8 Read and Compute. At DRAM-bound 1M sizes it shrinks the residual gap from 1.7x to 1.2x. It requires zero intrinsics, is portable C++20, and is strictly better than for_each_unrolled<N> at every size past L2. K=8 gives the best results on this hardware; K=4 is close.

  12. Measure your own workload. These numbers are for Intel Core Ultra 7 with 12 MiB L3. A machine with 24 MiB L3 gives different results (the README's float8 speedup was 10x, mine is 3.4x). The direction of the effect is universal; the magnitude is not.

The code is on GitHub: SOA_iterator.