Share on LinkedInBack to deep dives

Linux memory

malloc() — a deep dive

Illustration introducing malloc internals with tcache and chunks.
malloc is not just storage: glibc keeps multiple layers of allocator metadata and caches behind a simple pointer.

Introduction

Most C programs use malloc and free as if they were simple storage operations. Underneath that API, glibc is making speed-focused decisions about chunk headers, per-thread caches, bins, heap growth, and when to involve the kernel.

This walkthrough is based on experiments on a 32-bit Raspberry Pi. The exact sizes can vary by architecture and glibc version, but the mental model is useful: the allocator tries very hard to satisfy small allocations in user space before it takes the slow path into the kernel.

The hidden header

A request like malloc(8) does not reserve only 8 bytes. glibc wraps the payload in a chunk that carries allocator metadata. On the 32-bit setup used here, the chunk includes a previous-size field, a size-and-flags header, and enough payload room for future free-list pointers.

align(max(min_chunk_size, request + overhead), 8)

// 32-bit example:
// min_chunk_size = 16 bytes
// request = 8 bytes
// overhead includes allocator metadata
Diagram showing a malloc chunk with previous size, header, and payload.
malloc returns the address after the header, so the caller sees the payload while glibc keeps metadata just before it.

Reading one word before the returned pointer exposes the chunk header in this 32-bit experiment:

int *p = malloc(8);
int header = p[-1];
printf("Header: 0x%x\n", header);

The header stores the chunk size, but the lowest three bits are reused as flags because aligned chunk sizes leave those bits available.

  • P-bit: previous chunk is in use.
  • M-bit: chunk came from mmap.
  • A-bit: chunk belongs to a non-main arena.
int chunk_size = header & ~0x7;
int p_bit = header & 0x1;
int m_bit = header & 0x2;
int a_bit = header & 0x4;

Playground 1

Chunk Layout Explorer

Change the requested payload size and watch the real chunk grow with hidden metadata and alignment padding.

Header8B
Payload8B
requested: 8Busable payload: 8Bactual chunk: 16B

The returned pointer starts at payload. The allocator still owns the bytes before it.

Tcache: the secret pocket

The surprising part is what happens after a small chunk is freed. Modern glibc keeps a per-thread cache called tcache. It lets the thread reuse recently freed chunks without taking heavier locks or immediately merging neighboring free space.

int *p = malloc(8);
int *q = malloc(8);

free(p);
int header = q[-1];

Even after freeing p, the next chunk's P-bit can still say the previous chunk is in use. That is intentional: tcache keeps the freed chunk quickly reusable while avoiding immediate coalescing work.

Diagram showing adjacent malloc chunks p and q with headers and payloads.
Adjacent chunks carry headers that let glibc reason about neighboring allocation state.
Heap segment detail showing tcache per-thread structure and allocated or free chunks.
The tcache per-thread structure keeps a fast list of recently freed chunks owned by the current thread.

Playground 2

Tcache Bin Simulator

Free small chunks into tcache, then allocate again. The next matching malloc can reuse cached memory without asking the heap.

Allocated

chunk 1chunk 2

Tcache bin: 32B

empty

Heap top chunk

free top chunk
allocated: 2cached: 0/7top chunk: 7 slices

Initial heap has one top chunk ready to be sliced.

The magic number 7

To force real allocator behavior beyond tcache, the experiment fills the tcache bin. For a size class, tcache commonly holds up to seven chunks. The first seven frees can stay in tcache, so their neighbors still appear publicly busy.

A very small eighth freed chunk may still avoid merging by going through fastbins. To observe coalescing, the experiment uses chunks larger than the fastbin range and then frees enough of them to fill tcache first.

int size = 100;
void *p = malloc(size);
void *q = malloc(size);
void *r = malloc(size);
void *s = malloc(size);
void *t = malloc(size);
void *u = malloc(size);
void *v = malloc(size);
void *w = malloc(size); // 8th chunk
void *x = malloc(size);

free(p);
free(q);
free(r);
free(s);
free(t);
free(u);
free(v);

free(w);

Once tcache is full and the chunk is too large for fastbins, the freed chunk can land in the unsorted bin. At that point the allocator acknowledges the neighbor as free, and the next chunk can show P-bit = 0.

ASCII style diagram showing tcache, unsorted bin, and allocated chunks.
The eighth larger free bypasses full tcache and can move into the unsorted bin, where real coalescing logic appears.

The libc leak

A freed chunk in the unsorted bin participates in a linked list. Its forward and backward pointers can point toward allocator structures inside loaded libc. If code later reads that freed memory, it can reveal a high-memory libc address.

unsigned int *ghost_w = (unsigned int *)w;
printf("w[0] (Forward Ptr):  0x%x\n", ghost_w[0]);
printf("w[1] (Backward Ptr): 0x%x\n", ghost_w[1]);

This is why use-after-free bugs can be dangerous: a pointer that was logically freed can still contain allocator-written metadata. If an attacker can read that stale memory, it may help defeat address randomization by disclosing where libc is mapped.

Verification with perf

The allocator's two worlds become visible with perf. Repeated small allocations can be served from tcache or fastbins entirely in user space. A much larger allocation may force glibc to ask the kernel to grow the heap with brk or use mmap, depending on threshold configuration.

for (int i = 0; i < 1000; i++) {
    free(malloc(10));
}

void *p = malloc(200000);
free(p);
Perf report showing brk-related symbols during a large malloc request.
The small allocations stay quiet, but a large request shows the allocator crossing into the kernel slow path.

In the observed profile, the large request made __brk and the kernel's __se_sys_brk path visible, while the small request loop was handled without repeated system calls.

Playground 3

Fast Path vs Slow Path Flow

Pick a request type and follow the allocator path from request adjustment to returned pointer.

1Adjust request

Round the requested bytes into an allocator-friendly aligned chunk.

2Check tcache

Try the current thread's hot freed chunks first.

3Check bins

If tcache misses, look through shared allocator bins for a match.

4Slice top chunk

If no reusable chunk exists, carve space from the current heap top.

5Ask kernel

Grow the heap only when the local heap cannot satisfy the request.

6Return payload

Return the pointer after metadata, not the true start of the chunk.

A matching chunk is already in tcache, so malloc returns quickly.

malloc flow

The allocator usually starts by adjusting the request into a valid aligned chunk size. It then checks fast user-space pools such as tcache and fastbins. If those do not satisfy the request, it carves from heap space, and only asks the kernel for more memory when the local heap cannot satisfy the allocation.

Diagram showing malloc phases from request adjustment through bins, top chunk, heap growth, and returned pointer.
malloc proceeds through adjustment, bin checks, top chunk slicing, heap growth, and finally returns the payload pointer.
Cartoon flowchart showing malloc checking tcache and fastbins, then heap, then kernel.
The high-level flow: check glibc pools first, carve from the heap next, and call the kernel only when needed.

Summary

  • Chunks: returned pointers hide allocator metadata immediately before the payload.
  • Tcache: per-thread cached frees are fast and can keep chunks publicly marked as in use.
  • Bins: once tcache and fast paths no longer apply, larger freed chunks can move through unsorted-bin logic.
  • Security: freed memory may contain allocator pointers, which is why stale reads can leak process layout.
  • Kernel calls: glibc avoids brk and mmap for small hot allocations whenever it can.