NUMA
A multiprocessor memory organization in which access latency and bandwidth depend on the memory node relative to the executing CPU.
Local and Remote Memory
In a NUMA system, memory can appear as one address space while access cost remains non-uniform. A CPU normally reaches memory attached to its own node with lower latency and higher bandwidth; remote-node access crosses an interconnect.
Thread and Memory Locality
Thread affinity and data placement have to be considered together. On systems with first-touch placement, the CPU node that first writes a page can influence where the physical memory is allocated.
Large heaps, database buffer pools and high-throughput native workloads can suffer when computation repeatedly accesses data on another node.
Common Misconception
NUMA is not simply a consequence of having a large amount of RAM. Socket topology, interconnect behavior and workload sharing determine the effect. Pinning a thread is not enough when its data remains remote.