False Sharing: Why Unrelated Writes Slow Each Other Down

Two CPU cores fighting over the same cache line while each writes a different variable inside it

Two threads update counter_a and counter_b respectively. The code never touches the same variable from both sides, there is no data race — and yet throughput can still drop noticeably. The problem usually is not language-level sharing; it is the granularity at which hardware maintains cache coherence.

If two frequently written variables fall on the same cache line, a write to any byte of that line by one core can invalidate the whole line held by the other core. The two cores repeatedly fight for write permission on the same line — that is false sharing.

The short version

  • Cache coherence operates at cache-line granularity, not at C/C++ field or Java object-field granularity.
  • Mainstream x86-64 and many AArch64 processors use 64-byte cache lines, but code should not treat 64 as universal without confirming the target platform.
  • The direct cost is repeated cache-line ownership migration, plus the coherence traffic and store stalls it causes. Data usually moves among core-private caches, shared caches, and the coherence interconnect — not every transfer hits DRAM.
  • Prefer designs that reduce high-frequency writes to adjacent memory by multiple threads. Padding and alignment work but enlarge objects; let performance data decide.

Why do different variables still interfere?

CPU caches store and maintain data in fixed-size lines. Assume the target processor has 64-byte cache lines; these two 8-byte counters will very likely sit on the same line:

struct Counters {
    std::atomic<std::uint64_t> a;
    std::atomic<std::uint64_t> b;
};

Thread A updates only a; thread B updates only b. From the C++ memory model’s perspective they are two independent atomic objects with no data race. But hardware cannot hand core A write permission on just the 8 bytes holding a while handing core B the 8 bytes holding b. When both fields sit on the same cache line, the coherence protocol manages the entire line.

Before writing, each core must obtain writable ownership of the line. Once core A owns it, core B’s copy is invalidated; core B’s next write must re-acquire ownership and invalidate core A’s copy. The higher the write frequency, the more frequent the ownership migrations.

MESI helps describe this, but real processors may use MESIF, MOESI, or directory-based protocols. In engineering terms, you do not need to pin every migration to two specific MESI states — just hold on to three facts:

  1. A write requires obtaining write permission on the cache line first.
  2. Copies of the same line on other cores must be invalidated or downgraded.
  3. When another core writes again, it must re-fetch the latest line and ownership through the coherence interconnect.

This overhead is often called cache-line bouncing. It does not break result correctness, but it can serialize updates that were supposed to run in parallel.

True sharing vs false sharing

  • True sharing: multiple threads read/write the same logical data, e.g. one global counter.
  • False sharing: threads access different logical data that merely happens to occupy the same cache line.

Both can appear as cache-line contention in hardware events. Diagnostic tools can first tell you “which cache line is hot”; you still need field layout and code semantics to decide which kind it is. A hot lock word, for example, is usually true sharing — the mechanism in Why Locking a Mutex Usually Doesn’t Need a System Call lives and dies on that one contended line.

How to write a benchmark the compiler cannot optimize away

If you increment a plain integer a fixed number of times and the intermediate results are never observed, an optimizing compiler may fold the whole loop into a single assignment. Such a program can print impressive speedup ratios while measuring nothing about coherence costs.

The C++17 example below uses std::atomic::fetch_add() so every read-modify-write actually happens. memory_order_relaxed provides no ordering between variables but keeps atomicity, which keeps the experiment focused on cache-line ownership contention:

#include <atomic>
#include <chrono>
#include <cstddef>
#include <cstdint>
#include <functional>
#include <iostream>
#include <thread>

constexpr std::size_t kCacheLine = 64;  // confirm on the target machine first
constexpr std::uint64_t kIterations = 100'000'000;

struct PackedCounters {
    std::atomic<std::uint64_t> a{0};
    std::atomic<std::uint64_t> b{0};
};

struct PaddedCounters {
    alignas(kCacheLine) std::atomic<std::uint64_t> a{0};
    alignas(kCacheLine) std::atomic<std::uint64_t> b{0};
};

template <class Counters>
void run(const char* name)
{
    Counters counters;
    std::atomic<int> ready{0};
    std::atomic<bool> start{false};

    auto worker = [&](std::atomic<std::uint64_t>& counter) {
        ready.fetch_add(1, std::memory_order_relaxed);
        while (!start.load(std::memory_order_acquire)) {}

        for (std::uint64_t i = 0; i < kIterations; ++i)
            counter.fetch_add(1, std::memory_order_relaxed);
    };

    std::thread t1(worker, std::ref(counters.a));
    std::thread t2(worker, std::ref(counters.b));

    while (ready.load(std::memory_order_acquire) != 2) {}
    auto begin = std::chrono::steady_clock::now();
    start.store(true, std::memory_order_release);

    t1.join();
    t2.join();
    auto end = std::chrono::steady_clock::now();

    std::chrono::duration<double> elapsed = end - begin;
    std::cout << name << ": " << elapsed.count() << " s, sum="
              << counters.a.load(std::memory_order_relaxed)
                 + counters.b.load(std::memory_order_relaxed)
              << '\n';
}

int main()
{
    run<PackedCounters>("packed");
    run<PaddedCounters>("padded");
}

Compile with optimization and thread support:

g++ -O2 -std=c++17 -pthread false_sharing.cpp -o false_sharing

Do not assume a fixed speedup factor. Results depend on CPU topology, thread migration, frequency scaling, virtualization, and background load. Pin the two threads to different physical cores and confirm they are not SMT siblings of the same physical core:

lscpu -e=CPU,CORE,SOCKET,NODE

A stricter program calls pthread_setaffinity_np() at thread entry to pin CPUs. Warm up each version, run multiple rounds, and compare medians or quantiles instead of recording a single run.

Confirm the target machine’s cache-line size from system information or the processor manual. On Linux, these files usually expose the coherence line size per cache level:

grep . /sys/devices/system/cpu/cpu0/cache/index*/{level,type,coherency_line_size}

C++17 also offers std::hardware_destructive_interference_size, expressing the minimum spacing that should separate objects to avoid destructive interference. It is an implementation-defined value; compiler support and cross-machine deployment matter, so do not mechanically substitute it for every constant without knowing your build targets.

Why is a bare char pad[56] not rigorous?

This layout appears to separate two 8-byte variables by 64 bytes:

struct Counters {
    std::uint64_t a;
    char padding[56];
    std::uint64_t b;
};

But it does not guarantee the struct’s start address is 64-byte aligned. If a begins in the middle of a cache line, b can still share a line with part of a; arrays, allocators, and enclosing structures further affect the actual addresses.

The more reliable approach is to align each hot field — or its wrapper — to the target spacing:

struct Counters {
    alignas(64) std::atomic<std::uint64_t> a{0};
    alignas(64) std::atomic<std::uint64_t> b{0};
};

Still verify the final layout with offsetof, sizeof, or a debugger, and make sure the allocation path satisfies extended alignment. Modern C++’s plain new allocates according to the type’s alignment requirement, but custom memory pools, shared-memory layouts, and cross-language FFI each need separate review.

How to locate false sharing with perf c2c

Linux perf c2c samples cache-to-cache events and aggregates results by shared data address. Use a build with debug symbols and sample under a production-like load:

perf c2c record -g -- ./your_app
perf c2c report --stdio

In the report, focus on local and remote HITM (Hit Modified) samples, the shared cache-line addresses, and the corresponding symbols and source locations. High HITM counts mean accesses hit modified lines held by another core or node — a strong hint of cache-line contention.

But high HITM alone does not prove false sharing: true contention on the same lock or shared queue produces similar signatures. After locating the address, complete two verification steps:

  1. List which fields share the cache line and which threads access each of them.
  2. Decide whether those threads operate on the same logical state, or merely write adjacent fields by coincidence.

perf c2c depends on processor PMU, kernel, and perf support. In VMs, cloud hosts, or restricted containers, hardware events may be unavailable; the sampling events also differ across architectures. When the tool is unavailable, combine perf stat, struct addresses, thread pinning, and A/B tests before and after alignment to build evidence — but do not conclude from a single latency drop.

A priority order for fixing false sharing

1. Reduce shared writes first

If multiple threads are only updating statistics, prefer per-thread or per-shard counters aggregated at read time. This separates cache lines and also reduces contention on a single atomic variable.

Java’s LongAdder takes exactly this sharding approach to lower single-point update cost under high contention. It suits statistical accumulation, but reads return an aggregated result; it should not replace AtomicLong where every update needs a single linearization point.

2. Separate high-frequency write fields

Put fields written frequently by different threads into different objects, shards, or cache lines. Hot/cold field separation also stops low-frequency metadata from being dragged along by ownership migration of the hot fields.

3. Align and pad only after confirming the bottleneck

  • C/C++ can use alignas; confirm the alignment spacing on the target platform.
  • Java can use the JDK-internal jdk.internal.vm.annotation.Contended; application code usually also needs module-access handling and -XX:-RestrictContended. It is an internal API — confirm your target JDK’s support before relying on it.
  • Go’s golang.org/x/sys/cpu.CacheLinePad comes from the external x/sys module, not the Go standard library. Platform-specific alignment wrappers also work, but check the resulting struct layout.

Padding increases object size, reduces effective cache and TLB capacity, and can amplify memory usage in arrays, queues, and systems with many small objects. Read-only fields sharing a cache line usually do not cause high-frequency invalidation, and low-frequency write fields may not be worth the space.

Conclusion

The goal of false-sharing optimization is not to give every variable its own cache line. It is to find adjacent data that different cores write frequently and that genuinely sits in a performance hotspot, then reduce ownership migration with sharding, reordering, or alignment. Confirm the access pattern before changing memory layout — that is usually safer than padding everything in sight.

Related reading: Why Locking a Mutex Usually Doesn’t Need a System Call shows the same coherence cost on the lock word itself, and malloc(100 MiB) Succeeded, So Why Is RSS Barely Higher? covers where the padded objects ultimately live in memory.

References

  1. Linux Kernel Documentation, False Sharing
  2. Linux man-pages, perf-c2c(1)
  3. C++ reference, hardware_destructive_interference_size
  4. OpenJDK, @Contended
  5. Go x/sys/cpu, CacheLinePad