Skip to content
Long Wang
Archive Search Tags About

Performance

Linux

What Happens to Your Data During a Linux read()? DMA, Page Cache, and CPU Copies

Follow a Linux buffered file read from storage to the page cache and user buffer. See what DMA handles, why the CPU …

Sep 10, 2026 · 9 min read
Linux

False Sharing: Why Unrelated Writes Slow Each Other Down

Two threads updating separate variables can still slow each other down when both land on the same cache line. See why …

Sep 2, 2026 · 8 min read
Linux

Why Locking a Mutex Usually Doesn't Need a System Call

An uncontended pthread_mutex_lock() usually runs a few user-space atomic instructions and never enters the kernel. …

Aug 27, 2026 · 8 min read
Linux

sendfile() Promises Zero-Copy. Which Copies Does It Actually Remove?

sendfile() is marketed as zero-copy, but data still moves. See which two CPU copies the read()+write() path really …

Aug 19, 2026 · 8 min read
Linux

Why Fewer Syscalls Don't Make io_uring Faster Than epoll

io_uring can reduce system calls without improving throughput or tail latency. Compare its completion model, batching, …

Jul 28, 2026 · 8 min read
Long Wang Field-tested notes on AI and backend engineering
Site ArchiveSearchTagsAbout
Follow GitHub RSS Email
Powered by Hugo and the Stacknote theme © 2026 Long Wang · All rights reserved