Tiny MPI Messages: The Cost of Convenience
I like wrappers around MPI.
MPI_Send and MPI_Recv are not hard, but they are sharp in exactly the places research code gets messy: datatype bookkeeping, buffer lifetimes, counts, tags, and error handling. Boost.MPI feels like relief. mpi4py feels even better when I just want to test a communication pattern without rebuilding a C++ executable.
But the uncomfortable question is:
If I put that convenience in a hot communication path, who pays?
My first instinct was to benchmark Boost.MPI against C MPI and declare the overhead. That would have been too easy, and probably wrong. Tiny MPI messages are dominated by fixed costs, so a bad result could come from the wrapper, the compiler, the MPI implementation, Python object serialization, or simply noise.
So I made the matrix deliberately annoying.
The Matrix
I tested three APIs:
- C MPI with explicit integer buffers;
- Boost.MPI with
int,std::vector<int>, andstd::string; - mpi4py with NumPy buffers and Python objects.
I tested two MPI implementations:
- OpenMPI
5.0.10; - MPICH
4.2.3.
And for the C/C++ cases, two compiler families:
- GCC/G++
15.2.0; - Clang/Clang++
22.1.5.
The machine did not have MPI installed system-wide, so I built the test environment in disposable conda-forge prefixes under /tmp. That is not how I would benchmark a supercomputer, but it was good enough for the question I cared about: do the relative costs survive when I swap MPI implementation, compiler, and payload representation?
The benchmark source lives in scripts/mpi-wrapper-bench.
How I Tried Not To Lie
The benchmark is intentionally boring:
- exactly two ranks;
- rank 0 sends to rank 1;
- rank 1 sends the same kind of payload back;
- I report one-way latency as half the measured round trip;
- 5,000 warmup round trips;
- 50,000 measured round trips;
- 5 process-level repetitions;
- the tables show the median of the five runs.
For the larger-size sweep later in the article, I used an adaptive iteration count so that a 1 MiB payload did not dominate the whole experiment: 30,000 repetitions for tiny payloads, 10,000 at 4 KiB, 2,000 at 64 KiB, and 300 at 1 MiB. That sweep uses three process-level repetitions. The point there is shape, not nanosecond-level ranking.
All measurements are same-node measurements. This is not a network fabric benchmark. It is a tiny-message software-path benchmark.
That distinction matters. At this scale, I am mostly measuring the local MPI path, wrapper dispatch, serialization, Python overhead, and copies. On a real cluster, fabric latency would shift the baseline upward and make some of these ratios less dramatic.
C MPI: The Floor Is Not Universal
First, the raw C MPI baseline:
| MPI | compiler | 1 int | 8 ints | 64 ints |
|---|---|---|---|---|
| OpenMPI | GCC | 0.298 us | 0.335 us | 0.369 us |
| OpenMPI | Clang | 0.298 us | 0.335 us | 0.376 us |
| MPICH | GCC | 0.278 us | 0.263 us | 0.312 us |
| MPICH | Clang | 0.264 us | 0.258 us | 0.258 us |
Two things matter here.
First, GCC vs Clang is almost boring. That is good. If changing compiler had moved the C MPI floor by 2x, I would have had to be much more careful about blaming any wrapper.
Second, OpenMPI and MPICH are not identical even before wrappers enter the story. MPICH was faster on this local tiny-message setup, especially for 64 integers. That does not mean MPICH is "faster MPI" in general. It means the MPI implementation sets the floor, and the floor is already moving before Boost.MPI or Python shows up.
That is the first lesson: do not benchmark a wrapper against one MPI implementation and pretend the result is universal.
Boost.MPI: The Scalar Path Is Cheap
For a single integer, Boost.MPI was almost on the floor:
| MPI | compiler | C MPI, 1 int | Boost.MPI, int | overhead |
|---|---|---|---|---|
| OpenMPI | G++ | 0.298 us | 0.307 us | 1.03x |
| OpenMPI | Clang++ | 0.298 us | 0.311 us | 1.04x |
| MPICH | G++ | 0.278 us | 0.283 us | 1.02x |
| MPICH | Clang++ | 0.264 us | 0.279 us | 1.06x |
That is not the result I expected when I started.
For scalar payloads, Boost.MPI is not "slow MPI" here. It is a very thin layer over an already tiny operation. The compiler also barely matters: G++ and Clang++ are within noise for this test.
If your message really is a scalar or something Boost.MPI can map cleanly, the wrapper overhead is not the scary part.
The same result is easier to see as a bar chart:
The Vector Was Not The Raw Buffer
There is a subtle trap in the wording above.
When I say Boost.MPI vector[1], I do not mean "Boost.MPI sending one integer buffer." I mean this:
std::vector<int> values = {42}; world.send(1, 0, values);
Boost.MPI also lets me write the lower-level version:
world.send(1, 0, values.data(), static_cast<int>(values.size()));
Those are not the same path.
The Boost.MPI header makes this explicit. For a scalar int, send_impl(..., mpl::true_) maps directly to MPI_Send. For a pointer/count array of int, array_send_impl(..., mpl::true_) also maps directly to MPI_Send. But send(std::vector<T>) does an extra container protocol: for MPI datatypes, it first sends the vector size, then sends values.data().
So the question becomes more precise:
Can I keep Boost.MPI's communicator abstraction but avoid the
std::vectorconvenience overhead?
Yes.
For a 4-byte payload, the Boost raw-buffer path nearly collapses back to the C MPI baseline:
| MPI | C MPI buffer | Boost raw buffer | Boost vector[1] |
|---|---|---|---|
| OpenMPI | 0.298 us | 0.308 us | 0.388 us |
| MPICH | 0.249 us | 0.290 us | 0.522 us |
The difference is not mysterious. send(data(), n) is one logical payload send. send(vector) carries container shape first, so a one-element vector still pays for "this vector has size 1" before the actual integer buffer.
The same picture holds across sizes:
At small sizes, the extra vector-size exchange is visible. At large sizes, the payload movement dominates and raw/vector converge. On MPICH at 1 MiB, both raw and vector were around 31 us in the focused Boost-only run. On OpenMPI, both were around 132 us.
I also checked the generated binary. nm -C and objdump -d -C show direct references to MPI_Send and MPI_Recv for the primitive paths. The string path, in contrast, pulls in boost::mpi::packed_oarchive and calls the packed archive send path. That matches the measurement: raw buffers behave like MPI buffers; vector adds a shape message; string is the serialization path.
This is the practical answer I wanted:
- Boost.MPI as a communicator wrapper: mostly fine.
- Boost.MPI
send(data(), n): close to C MPI. - Boost.MPI
send(vector): convenient, but not free for tiny messages. - Boost.MPI
send(string)or custom objects: now you are in serialization territory.
The Container Cliff
The story changes when the payload becomes a convenient C++ object:
| MPI | compiler | vector<int>[1] | vector<int>[8] | vector<int>[64] | string[8] | string[64] |
|---|---|---|---|---|---|---|
| OpenMPI | G++ | 0.374 us | 0.412 us | 0.458 us | 0.652 us | 0.647 us |
| OpenMPI | Clang++ | 0.373 us | 0.403 us | 0.453 us | 0.653 us | 0.672 us |
| MPICH | G++ | 0.518 us | 0.520 us | 0.532 us | 0.823 us | 0.830 us |
| MPICH | Clang++ | 0.522 us | 0.522 us | 0.531 us | 0.809 us | 0.851 us |
This is the important table, but it needs the clarification above.
The cost does not explode with 64 integers. It mostly jumps when I switch from scalar/raw-ish payloads to container convenience. For std::vector<int>, that is not full Boost.Serialization; it is a size message plus a data message. For std::string, the binary contains the packed archive path, so that one really is serialization-shaped.
That is exactly the kind of trap that convenience APIs create. The dangerous thing is not that boost::mpi::communicator::send exists. The dangerous thing is that it makes it natural to send a std::vector<int> when the hot path really wanted a stable buffer and an explicit count.
In other words, Boost.MPI did not betray me. It let me write the expensive version more easily.
What Happens When The Messages Get Bigger
The first tables are intentionally tiny. That is where wrapper and serialization overhead should be most visible. But stopping at 64 integers hides the other half of the story: once the payload grows, the fixed overhead can become irrelevant, or it can expose a different copy path.
So I ran a second sweep at:
- 4 bytes;
- 32 bytes;
- 256 bytes;
- 4 KiB;
- 64 KiB;
- 1 MiB.
For this sweep I kept the plot focused: C MPI with GCC, Boost.MPI std::vector<int> with G++, mpi4py NumPy buffers, and mpi4py Python objects where the object path remained practical to measure. The GCC/Clang comparison was already boring enough in the tiny-message matrix, so I did not duplicate every compiler variant in the graph.
OpenMPI produced the most intuitive large-message shape. At 4 bytes, C MPI is around 0.30 us, Boost.MPI vector around 0.37 us, and mpi4py NumPy around 0.91 us. At 1 MiB, all three buffer-like paths land in the same rough region: C MPI at 131 us, Boost.MPI vector at 131 us, and mpi4py NumPy at 137 us.
That is the crossover I wanted to see. For tiny messages, the software path dominates. For 1 MiB messages, the payload movement dominates and the wrapper mostly disappears, at least for buffer-like payloads.
The Python object path is different. It is already 1.88 us for one integer and about 26.5 us for a 1024-integer list. I stopped plotting it after that because the lesson was no longer subtle: object serialization is not the same communication regime.
MPICH had a different shape. It was faster than OpenMPI on the small C MPI path, and very competitive through 64 KiB. At 1 MiB, the run was noisier: C MPI had a median of 32.6 us but a max run above 115 us; Boost.MPI vector landed around 101 us; mpi4py NumPy around 127 us.
I would not overclaim that 1 MiB MPICH number from this laptop setup. It is exactly the kind of result that wants a cleaner pinned run, process binding, and probably a real cluster node before turning into a claim. But it still supports the larger point: MPI implementation and transport details matter more as the payload grows.
The better conclusion is not "OpenMPI wins" or "MPICH wins." It is:
Tiny-message overhead and large-message throughput are different questions.
mpi4py Has Two Personalities
mpi4py showed the same split, louder:
| MPI | payload | 1 int | 8 ints | 64 ints |
|---|---|---|---|---|
| OpenMPI | NumPy buffer | 0.927 us | 0.910 us | 0.960 us |
| OpenMPI | Python object | 1.907 us | 2.087 us | 3.329 us |
| MPICH | NumPy buffer | 0.936 us | 0.944 us | 1.197 us |
| MPICH | Python object | 2.328 us | 2.545 us | 3.815 us |
A NumPy buffer is not C MPI, but it is still a real MPI buffer path. On this machine, it landed around 0.9-1.2 microseconds one-way. That is roughly 3x the C MPI baseline, but not absurd for Python driving MPI calls.
Python objects are a different benchmark. comm.send([1, 2, 3]) is convenient, but it is serialization. The one-way cost moved from about 0.9 us for NumPy buffers to 1.9-3.8 us for Python objects.
The 64-int Python object case was more than 10x the best C MPI floor in this test.
That does not mean "mpi4py is slow." It means mpi4py has two very different personalities:
- buffer mode: MPI with Python overhead;
- object mode: pickle and messaging.
Those are not interchangeable performance models.
What Actually Mattered
After making the matrix annoying, the ranking was clearer than the headline I originally wanted:
- MPI implementation set the floor. MPICH and OpenMPI differed before wrappers entered the picture.
- Compiler barely mattered. GCC/G++ and Clang/Clang++ were not the main variable here.
- Boost.MPI scalar sends were close to C MPI. The overhead was roughly 2-6%.
- Container convenience introduced fixed cost.
std::vector<int>sends shape plus data;std::stringgoes through the packed archive path. - mpi4py buffer mode and object mode are different worlds. NumPy buffers were moderate overhead; Python objects were serialization-heavy.
- The larger the buffer, the less the wrapper mattered. At 1 MiB, OpenMPI C MPI, Boost.MPI vector, and mpi4py NumPy were much closer than at 4 bytes.
The thing I wanted to call "wrapper overhead" split into three different costs:
- wrapper dispatch;
- payload representation;
- container shape metadata;
- serialization.
Only the first one was small enough to be boring.
The Practical Rule
My rule after this experiment:
- Use C MPI or explicit buffers in hot tiny-message paths.
- Use Boost.MPI freely for control flow, setup, metadata, and code where the message granularity is not microscopic.
- If using Boost.MPI in a hot path, be suspicious of
std::vector,std::string, and custom serialized objects. - Use mpi4py with NumPy buffers when you care about communication cost.
- Treat mpi4py Python-object sends as a convenience feature, not a latency baseline.
- Always test both the MPI implementation and the payload path before blaming the abstraction.
The shortest honest verdict is:
The wrapper was not the villain. The payload path was.
That is less satisfying than "Boost.MPI is slow", but much more useful.