Wander Lairson Costa

What every kernel programmer should know about Jump Labels

• kernel

1 Introduction

Kernel hot paths frequently evaluate conditions that rarely change: tracepoint activation, hardware mitigations, or debug logging. In a naive implementation, a standard conditional branch (if (flag)) must load the state from memory on every iteration, incurring a guaranteed cache access and wasting execution cycles.

Jump labels bypass this overhead. By dynamically modifying the executable binary, this mechanism rewrites the hot path at runtime. While a feature remains disabled, the CPU encounters a single, fast nop (or an unconditional jump around out-of-line code). Toggling the state swaps the instructions in-place to redirect execution. This architecture accepts expensive update-time coordination in exchange for zero run-time penalty in the common path.

This guide divides into two distinct sections. The first half—spanning from §2 to §4—is a cookbook and API reference for developers who need to integrate static keys into kernel modules and core code.

The second half delves into the low-level implementation details on the x86_64 architecture under Linux 7.2. It traces the machine instructions generated by the compiler, the design of the jump-table metadata, objtool integration, boot-time initialization, live code patching, the multi-processor INT3 synchronization protocol, and module loading behavior.

To prevent confusion, this tutorial strictly distinguishes between two related terms:


2 The problem jump labels solve

Kernel code is full of conditional checks that guard optional or debugging features: “is tracing enabled for this tracepoint?”, “is this security module active?”, “is this debugging feature on?”. A standard C implementation:

if (some_feature_enabled)
        do_something();

Even when some_feature_enabled remains false almost indefinitely, the CPU must still perform three steps:

  1. Load some_feature_enabled from its cache line into a register.
  2. Compare the register value against zero.
  3. Predict and branch based on the result.

Modern branch predictors handle the conditional branch in step 3 with high accuracy, hiding the misprediction penalty. However, branch prediction cannot bypass the guaranteed memory load in step 1. When this check resides in a hot path executed millions of times per second (such as the scheduler, network packet processing, or tracepoint triggers), the cumulative overhead of those memory accesses becomes a measurable performance bottleneck.

Jump labels eliminate both the memory access and the comparison for the common path by rewriting the instructions at runtime. When the feature is inactive, the instruction pipeline encounters a no-operation (nop) instruction (or an unconditional branch that bypasses the out-of-line code). When a subsystem activates the feature, the kernel traverses all call sites registered for that static key and patches the instructions in place.

The design tradeoff is stark: toggling is expensive because it requires a machine-wide CPU synchronization and precise text patching, but running the hot path is virtually free.


3 The mental model, in one diagram

When a developer guards a conditional code block using static_branch_unlikely(), the compiler generates an instruction layout where the hot path remains linear and cold blocks reside out-of-line. The entry point of this sequence is a patchable location whose instruction is determined by the runtime state of the static key:

   SOURCE CODE                     FEATURE OFF (common)         FEATURE ON
   ------------                    --------------------         ----------
   if (static_branch_unlikely      nop  (2 or 5 bytes)          jmp .Lout_of_line
       (&my_key)) {
           rare_code();            .Lout_of_line:               .Lout_of_line:
   }                                   rare_code();                 rare_code();
                                       jmp back                     jmp back
                                   (unreachable without a jmp)

In the disabled state (“FEATURE OFF”), execution flows sequentially without branching. The out-of-line block containing rare_code() is preserved in the binary, but because no active jump instruction targets the .Lout_of_line label, normal instruction execution bypasses it entirely.

The unconditional return jump (jmp back) at the end of the out-of-line block is a standard compiler optimization rather than part of the jump label framework. The compiler moves the cold block to the end of the function and inserts a jump to return execution to the statement immediately following the conditional block. This return branch remains static throughout execution; only the entry-point instruction at the top of the site is patched when toggling the key.

Key State Instruction in the Hot Path Cost when the Feature is Disabled Cost when the Feature is Enabled
Disabled (for an unlikely site) nop Negligible (linear execution, pipeline fall-through, no memory access) Unreachable (the code path is bypassed and cannot execute)
Enabled jmp .Lout_of_line Single unconditional jump Entry jump + cold path execution + return jump

Contrast this mechanism with the conventional conditional statement in §2, which always incurs memory load and comparison overhead.


4 How to use static keys (the cookbook)

Toggling a static key is not a simple memory write; it is a live code modification that patches executable instructions across every online CPU core. This operation incurs a heavy synchronization penalty. Consequently, the API offers specialized variants to manage update frequency, control caller authorization, and coordinate multiple owners.

4.1 Minimal example

A minimal static key implementation spans three phases: declaring the key, guarding the conditional branch in the hot path, and toggling the branch target from a separate control path.

#include <linux/jump_label.h>

DEFINE_STATIC_KEY_FALSE(foo_key);

void hot_path(void)
{
        /* Fast path: compiled as NOP while the key is false. */
        if (static_branch_unlikely(&foo_key))
                do_rare_thing();

        do_common_work();
}

void foo_enable(void)
{
        static_branch_enable(&foo_key);   /* slow path: patches text */
}

void foo_disable(void)
{
        static_branch_disable(&foo_key);  /* slow path: patches text */
}

Declaring the key with DEFINE_STATIC_KEY_FALSE() instantiates a struct static_key_false, which wraps an atomic_t counter in a unique wrapper type. This type-level distinction allows the compiler to differentiate the key from a struct static_key_true. Pairing this key with static_branch_unlikely() instructs the compiler to emit a nop instruction for the initial, disabled state.

4.2 Choosing TRUE vs FALSE and likely vs unlikely

Selecting an incorrect combination of key declaration and branch macro does not affect behavioral correctness, but it introduces a subtle performance penalty. If the chosen combination disagrees with the steady-state execution flow, the compiler generates a jmp instruction instead of a nop on the hot path. This layout penalty remains invisible to functional testing and cannot be corrected at runtime. The generation of either arch_static_branch or arch_static_branch_jump instruction sequences is determined entirely at compile time.

Selecting the optimal configuration depends on two design criteria:

  1. The default state at boot time: Most optional features default to a disabled state, requiring DEFINE_STATIC_KEY_FALSE. Features that remain enabled unless explicitly deactivated require DEFINE_STATIC_KEY_TRUE.
  2. The expected steady-state execution path: Developers must apply static_branch_unlikely() when the conditional block represents the rare execution path. Conversely, static_branch_likely() must guard paths where the conditional block represents the common execution path.

These declarations and branch macros can be combined arbitrarily; a false-default key is compatible with both likely and unlikely macros, as is a true-default key. The jump label subsystem coordinates these combinations to ensure that the initial default state always compiles to a cheap nop instruction.

A typical implementation for optional features uses the following pattern:

DEFINE_STATIC_KEY_FALSE(feature_key);

if (static_branch_unlikely(&feature_key))
        rare_enabled_path();

4.3 Boolean enable vs refcounted enable

The distinction between boolean and reference-counted interfaces addresses a classic coordination failure. If two independent subsystems call static_branch_enable() on a shared key, both expect the code path to remain active. If the subsystem that finishes first calls static_branch_disable(), the code path is patched off for both, leaving the second subsystem silently broken. Reference counting prevents this premature deactivation.

API Semantics When to use
static_branch_enable / static_branch_disable Force enabled count to 1 or 0 Single owner; simple on/off
static_branch_inc / static_branch_dec Refcount; patch only on 0↔1 Multiple independent users
static_branch_slow_dec_deferred Dec, but delay the 1→0 patch Userspace-driven toggles

Under the reference-counted API, a static key remains enabled as long as the counter is non-zero. The transition from zero to one triggers the initial text-patching operation to enable the branch. Subsequent increments are cheap atomic operations that bypass text patching entirely. Conversely, only the final decrement from one to zero triggers the text patch to disable the branch.

Mixing the boolean and reference-counted APIs on a single key leads to corrupt state. Although both interfaces manipulate the underlying enabled counter of struct static_key, they operate under conflicting assumptions. The boolean interface expects a binary state (strictly zero or one), whereas the reference-counted interface expects an arbitrary non-negative integer.

If a caller of static_branch_inc() has incremented the counter to two or more, invoking static_key_enable() or static_key_disable() will trigger safety assertions rather than the expected behavior. When static_key_enable() is called on a key whose value is already greater than zero, it assumes the branch is active and returns immediately. Conversely, if static_key_disable() is called when the reference count is greater than one, it detects that the count does not match the expected value of one required for a safe shutdown. In both scenarios, the kernel emits a warning via WARN_ON_ONCE and aborts the operation, leaving the instruction patch unmodified.

4.4 Reading the state without taking the branch

Every example so far, including hot_path() in §4.1, uses static_branch_likely or static_branch_unlikely as an if condition. The purpose of these macros is to compile directly into a patchable branch instruction. However, a caller occasionally requires the current boolean state of a key as an ordinary expression rather than a patchable branch. The static_key_enabled() macro provides this capability:

if (static_key_enabled(&foo_key))
        /* plain atomic read of the count — NOT the patched fast path */

Two patterns from the kernel illustrate why the API provides a dedicated read function rather than simply relying on a slow-path conditional branch.

State reporting represents the first pattern. During initialization, the kernel logs configuration decisions rather than branching on them. For example, arch/x86/kernel/cpu/bugs.c formats the active Spectre and IBPB mitigations into a message once at boot: pr_info(..., static_key_enabled(&switch_mm_always_ibpb) ? "always-on" : "conditional") Because this message is generated only during early boot, compiling a patchable fast-path instruction is unnecessary.

Control-plane guarding represents the second pattern. Before invoking the path that patches instructions across every CPU, drivers/md/dm-stats.c verifies the current state of the key: if (!static_key_enabled(&stats_enabled.key)) static_branch_enable(&stats_enabled); Querying the state of the key first avoids executing a costly text-patching sequence if the key is already active, preventing redundant calls to static_branch_enable().

Both scenarios require the current boolean state as an ordinary expression to print, compose, or evaluate during setup—operations that the patchable branch macros cannot accommodate.

On hot paths, however, callers must use the branch macros to ensure the compiler generates patchable instructions. Substituting static_key_enabled() in a hot path bypasses the jump label infrastructure entirely. This forces the CPU to pay the cost of a cache-line load on every execution—returning to the exact memory-access bottleneck §2 that static keys are designed to eliminate.

4.5 Keys must be global / static storage

A static key cannot reside on the stack or be dynamically allocated with kmalloc(). Every static branch relies on compile-time inline assembly to register the key address in the sidecar metadata section, __jump_table.

Under the hood, the inline assembly block uses the immediate operand constraint ("i") to pass the address of the key to the assembler. Because the assembler must compute a relative offset between the jump site and the key (as detailed in §7.3), the address of the key must be a link-time constant.

If a developer attempts to pass a pointer to a stack variable or a heap-allocated struct, the compiler cannot satisfy the immediate constraint and will reject the code with a compilation error. This compile-time check prevents silent runtime memory corruption that would otherwise occur when a function returns and destroys its stack-allocated key, or when a dynamically allocated key is freed.

To define static keys correctly, always place them in global or file-local static storage using DEFINE_STATIC_KEY_FALSE. The API provides several initialization macros for different scopes and patterns:

/* Global key defined in a source file (.data section) */
DEFINE_STATIC_KEY_FALSE(global_key);

/* File-local key visible only within the translation unit */
static DEFINE_STATIC_KEY_FALSE(file_local);

/* Declaration for header files to share a global key */
DECLARE_STATIC_KEY_FALSE(global_key);

For grouping multiple toggles together, define an array using DEFINE_STATIC_KEY_ARRAY_FALSE:

DEFINE_STATIC_KEY_ARRAY_FALSE(keys, 4);

if (static_branch_unlikely(&keys[i])) {
    /* ... */
}

When a key should be conditionally defined based on a Kconfig option, use DEFINE_STATIC_KEY_MAYBE paired with static_branch_maybe:

DEFINE_STATIC_KEY_MAYBE(CONFIG_FOO, foo_key);

if (static_branch_maybe(CONFIG_FOO, &foo_key)) {
    /* ... */
}

4.6 Read-only-after-init keys

While §4.5 detailed static keys designed for lifetime mutability, certain hot-path conditions require absolute immutability once configured. For instance, a hardware mitigation decided during early boot should never be toggled again. Relying solely on software-level discipline to prevent accidental toggling is fragile. Instead, the kernel provides a hardware-enforced guarantee through DEFINE_STATIC_KEY_FALSE_RO and DEFINE_STATIC_KEY_TRUE_RO:

DEFINE_STATIC_KEY_FALSE_RO(configured_once_at_boot);

These macros place the underlying static key structure in the __ro_after_init section. The kernel permits modifications via static_branch_enable() or static_branch_disable() exclusively during early boot (before the system invokes mark_rodata_ro()). Once initialization finishes, a defense-in-depth architecture locks down the key through two complementary mechanisms:

The resulting freeze provides robust security hardening. If an attacker leverages an arbitrary-write vulnerability elsewhere in the kernel to compromise system memory, they still cannot disable a hardened mitigation key. Because the enabled variable resides in write-protected memory, any modification attempt is blocked by the MMU, matching the security profile of traditional .rodata. Developers must therefore use the _RO variants for “decide once at boot, then freeze” security and performance knobs, reserving plain DEFINE_STATIC_KEY_* for variables that genuinely require dynamic runtime toggling.

4.7 Rate-limited disable (userspace-facing knobs)

When userspace can toggle a feature rapidly, patching instructions on every transition degrades system performance. Each text-patching cycle forces a full round of inter-processor interrupts (IPIs) to broadcast the instruction changes across all online CPUs. If userspace can toggle a feature rapidly—such as via a sysctl or a socket option—naive patching on every transition turns a simple state change into a machine-wide synchronization storm.

The deferred static branch API introduces deliberate asymmetry. While enabling remains immediate, disabling is deferred. static_branch_deferred_inc() is a direct alias for the standard reference-counting increment static_branch_inc()—it executes with zero delay. If a static key guards a critical tracepoint or statistic counter, deferring the enable path would cause the kernel to silently drop early events. The disable path can safely tolerate delay; keeping a feature active for a few additional milliseconds is harmless, whereas immediate text patching on high-frequency toggles is not.

The deferred decrement function, static_branch_slow_dec_deferred(), implements this asymmetry by dividing the decrement logic into two execution paths:

As a result, the feature remains fully active, fully patched, and reference-counted at one throughout the entire timeout window. This latency window enables event coalescing. If a new increment arrives before the timer expires, the reference count rises from one to two via the fast path described in §4.3. When the timer eventually fires, the delayed work handler executes a single, standard decrement. Because the count drops from two to one rather than transitioning to zero, the decrement does not trigger a text update. A rapid disable-then-enable sequence completed within the timeout window bypasses the text-patching machinery entirely.

#include <linux/jump_label_ratelimit.h>

DEFINE_STATIC_KEY_DEFERRED_FALSE(sockopt_key, HZ);

/* enable immediately */
static_branch_deferred_inc(&sockopt_key);

/* disable — may wait up to `timeout` before actually patching off */
static_branch_slow_dec_deferred(&sockopt_key);

/* force pending delayed work to finish (e.g. module exit) */
static_key_deferred_flush(&sockopt_key);

Before freeing the enclosing struct static_key_false_deferred—the structure containing the static key, the timeout interval, and the delayed_work state—the caller must invoke static_key_deferred_flush(). This function blocks until any pending deferred disable work completes. Flushing is critical during module unloading or dynamic memory reclamation. If the enclosing memory is deallocated while the timer remains active, the subsequent expiration of the timer will trigger jump_label_update_timeout() on a freed delayed_work structure, resulting in a use-after-free panic.


5 Hardware background (why this is hard)

Up to this point, we have treated nop and jmp instructions as abstract logical states of a static branch. The remainder of this guide explores how the kernel actually generates these instruction bytes and dynamically hot-swaps them on a running system. This transition from software-level branch hints to live runtime code patching relies on three fundamental hardware realities: the pipelined execution of instructions versus memory loads, the precise byte-level instruction encodings of the x86_64 architecture, and the concurrency hazards that prevent a multi-processor kernel from safely overwriting active instructions with ordinary memory writes.

5.1 What a CPU actually does with instructions

A modern x86_64 core decouples instruction execution from the instruction stream using a deeply pipelined, out-of-order execution engine. Rather than executing instructions in a strict, sequential lock-step, the hardware continuously processes instructions through four key stages:

  1. Fetch: The hardware reads raw instruction bytes from the L1 instruction cache (L1i).
  2. Decode: Decoders convert these variable-length instruction bytes—ranging from 1 to 15 bytes on the x86_64 architecture—into fixed-length internal micro-operations (uops).
  3. Execute: Execution units dispatch uops out of order to specialized execution ports. A hardware branch predictor guesses the outcomes of conditional branches to keep these pipelines fully saturated.
  4. Retire: The processor commits results back in-order using a reorder buffer (ROB) to preserve the programmer-visible illusion of sequential execution.

While a highly accurate branch predictor can mask the latency of a well-predicted conditional branch, it cannot eliminate the memory load that feeds the check. Out-of-order execution makes the branch instruction itself seem virtually free, as a correct prediction avoids pipeline flushes. However, the underlying memory load that retrieves the state of the flag (such as feature_enabled) must execute on every single pass. This load consumes a load buffer entry, an L1 data cache (L1d) read port, and an execution port. Under heavy data-cache pressure, or if a writer on another CPU core modifies the flag, cache coherence protocols invalidate the line. This invalidation forces the line to bounce across cores, turning a trivial data-cache lookup into a high-latency memory stall that halts the instruction window.

Replacing this conditional check with an unconditional nop or a direct jmp removes the memory load entirely. Because there is no condition to evaluate and no flag to read, the CPU avoids data-cache access altogether. When the static branch is disabled, the core decodes the nop at the frontend and discards it with minimal overhead, requiring no execution ports or load buffers. There is no cache line to bounce between cores and no state to track in the branch target buffer (BTB), leaving the execution engine free to focus on the surrounding instruction stream.

5.2 x86 instruction encoding: JMP and NOP

Replacing an instruction at runtime requires the original and replacement instructions to occupy the exact same number of bytes. Because x86 is a variable-length instruction set architecture (ISA), individual instructions vary from one to fifteen bytes in length. The specific instruction encodings manipulated by the jump label subsystem are:

Instruction Opcode bytes Total size Reach
INT3 (breakpoint) CC 1 byte n/a
JMP rel8 (short) EB xx 2 bytes -128..+127 bytes from the end of the instruction
JMP rel32 (near) E9 xx xx xx xx 5 bytes ±2 GiB
2-byte NOP 66 90 2 bytes —
5-byte NOP 0f 1f 44 00 00 (nopl 0x0(%rax,%rax,1)) 5 bytes —

The corresponding opcode constants are defined in arch/x86/include/asm/text-patching.h (such as JMP8_INSN_SIZE, JMP8_INSN_OPCODE, JMP32_INSN_SIZE, JMP32_INSN_OPCODE, and INT3_INSN_OPCODE) and arch/x86/include/asm/nops.h (including BYTES_NOP5).

A relative jump specifies a target offset relative to the instruction pointer of the subsequent instruction. Therefore, the relative displacement is measured from the byte immediately following the jump instruction:

displacement = destination - (instruction_address + instruction_size)

The inline helpers text_gen_insn() and __text_gen_insn() compute this exact offset when formatting the instruction buffer.

Why matching size matters: Swapping a five-byte nop with a five-byte jmp preserves the exact layout of the surrounding text. Because no instruction boundaries shift, return addresses stored on the stack, relative targets of nearby branch instructions, exception table entries, and ORC unwind metadata remain fully valid and require no relocation. The patch operates as a strictly localized, in-place byte substitution.

5.3 Why you cannot just memcpy over live code on SMP

Modifying active kernel instructions on symmetric multiprocessing systems introduces severe architectural challenges that do not exist when writing to standard data structures. These challenges stem from three distinct, concurrent properties of kernel text memory:

  1. Kernel text memory is mapped read-only after early boot under the STRICT_KERNEL_RWX configuration option in arch/Kconfig. Any direct store operation triggers a page fault, requiring a dedicated mechanism to bypass the write protection.
  2. Active instructions are fetched concurrently by other execution cores. No global pause freezes the system during a modification, meaning any core can call or continue executing the target function at any instruction boundary.
  3. Stale instruction bytes might reside within the execution pipeline or instruction cache of another core, having been fetched before the modification but not yet fully decoded or executed.

The primary danger lies in the concurrent instruction fetch described in the second point. Consider a scenario with two cores, Core A and Core B, where Core A attempts to patch a five-byte instruction while Core B executes that same instruction sequence repeatedly. A five-byte memory write is not an atomic operation on modern memory buses. The hardware executes the modification as multiple independent write cycles, each of which becomes visible to other cores at slightly different times. An instruction fetch on Core B can occur precisely in the middle of this multi-step update, reading a mixture of old and new bytes:

                     time --->

 Core A (patcher)   [ write bytes 2-4 ]   [ write bytes 0-1 ]
                                      ^
                                      |
 Core B (fetcher)             [ fetch all 5 bytes, right here ]
                                      |
                                      v
                 byte-by-byte: 0=old  1=old  2=new  3=new  4=new
                 = torn mix: 2 old bytes + 3 new bytes
                 = neither the old instruction
                   nor the new one — garbage

When Core A writes five bytes while Core B fetches the instruction, Core B can encounter a torn instruction. This mixture of old and new bytes constitutes a malformed instruction that the hardware instruction decoder cannot safely decode, triggering an invalid opcode exception or unpredictable behavior. The x86 architecture provides no guarantee that multi-byte stores to live, concurrently executing instruction areas are atomic from the perspective of an instruction fetch.

In contrast, a single-byte store is always atomic for instruction fetch operations. A CPU core can never observe a single byte in a partially modified state, as a single byte represents the minimum unit of coherent memory access. Both the Intel Software Developer Manual and the patching implementation in the Linux kernel leverage this atomic behavior to transform an unsafe multi-byte write into three safe, sequential steps:

  1. Replace the first byte of the instruction site with a single-byte breakpoint instruction, INT3 (0xCC). Because this single-byte write is atomic, any concurrent execution core either reads the original instruction or hits the breakpoint trap. A torn instruction state is impossible.
  2. Overwrite the remaining bytes of the instruction while the INT3 instruction remains at the entry boundary. No concurrent execution thread can fetch or decode these modified trailing bytes, because any execution attempt immediately traps at the preceding INT3 byte.
  3. Replace the INT3 byte with the first byte of the newly prepared instruction using another atomic, single-byte write. This single store marks the exact transition when the new instruction becomes live and executable.
  4. Execute a global synchronization across all cores between each of these steps. This synchronization, driven by an inter-processor interrupt, forces every core to execute a serializing instruction. The serialization flushes stale instruction bytes from the pipelines and instruction caches, preventing any core from executing out-of-date instruction sequences.

This synchronization protocol is implemented in smp_text_poke_batch_finish() within arch/x86/kernel/alternative.c (detailed in §10). Jump labels represent only one client of this multi-step patching mechanism; other core subsystems, including dynamic ftrace, static calls, kprobes, and alternative patching, rely on this identical atomic replacement protocol.

5.4 Writing read-only kernel text: text_poke()

Modern kernels no longer patch instruction text by clearing the write-protection (WP) bit in the %cr0 control register, performing the write, and immediately restoring the bit. While this technique was historically common, it represents a blunt instrument that compromises system integrity. Between the clearing of the WP bit and the restoration of the bit, any concurrent write from any CPU core could land on write-protected memory, rather than only the target instruction undergoing patching. Any execution thread running during this critical window could accidentally corrupt kernel memory that should remain read-only.

To eliminate this risk, __text_poke() employs a highly localized approach. Instead of modifying the permissions of the existing read-only virtual mapping of the .text section, the kernel establishes a second, transient, and private virtual mapping that points to the exact same physical page of RAM, performing the write through this alias instead.

Physical RAM operates independently of virtual memory mappings. The same physical frame can be mapped simultaneously through multiple virtual addresses, each carrying distinct page permissions. The standard virtual mapping of the .text section, which remains visible to all CPU cores throughout the lifetime of the system, is strictly read-only and executable. This permanent mapping allows any active core to fetch and execute instructions at any moment.

During a patching operation, __text_poke() temporarily configures a second virtual mapping to the same physical page. This alias is marked writable but not executable, and is restricted solely to the specific CPU core performing the patching. Tearing down this mapping within a few instructions minimizes the exposure. This design operates like a second door into a secure room: the underlying content (the raw instruction bytes) remains identical regardless of the door used to access it, but only one of the doors is ever unlocked, and then only for the patching thread.

Resolving this secondary mapping involves a sequence of safeguards implemented within the body of __text_poke(). The function manipulates the memory data to safely configure, switch, write, and tear down the temporary page(s):

static void *__text_poke(text_poke_f func, void *addr, const void *src, size_t len)
{
        bool cross_page_boundary = offset_in_page(addr) + len > PAGE_SIZE;
        struct page *pages[2] = {NULL};
        struct mm_struct *prev_mm;
        unsigned long flags;
        pte_t pte, *ptep;
        spinlock_t *ptl;
        pgprot_t pgprot;
        ...

In this signature, func represents the target copy routine, resolving to a memcpy-like function for standard text_poke() calls, or a memset-like helper for the _set variant. The parameters addr, src, and len specify the target virtual address, the source payload, and the copy size.

To determine the physical backing of the memory being patched, __text_poke() identifies the underlying physical pages. This normally requires a single page, but can require two pages if the write spans a page boundary. For core kernel text, the pages are retrieved using virt_to_page(); for text residing within a dynamic kernel module, the pages are resolved via vmalloc_to_page():

if (!core_kernel_text((unsigned long)addr)) {
        pages[0] = vmalloc_to_page(addr);
        if (cross_page_boundary)
                pages[1] = vmalloc_to_page(addr + PAGE_SIZE);
} else {
        pages[0] = virt_to_page(addr);
        if (cross_page_boundary)
                pages[1] = virt_to_page(addr + PAGE_SIZE);
}

The next phase points a pre-allocated page-table entry at the resolved physical page within a dedicated, otherwise-empty virtual address space called text_poke_mm. This address space is initialized once at boot time by poking_init(). The temporary entry is configured as writable and explicitly lacks the _PAGE_GLOBAL attribute.

Excluding the global bit keeps the overhead of the operation minimal. Because a non-global mapping is cached only within the translation lookaside buffer (TLB) of the current CPU core, dismantling the mapping requires only a local TLB invalidation via flush_tlb_mm_range(). This avoids the expensive inter-processor interrupts (IPIs) that would otherwise be required to flush the TLBs of other cores, as no other core ever loads text_poke_mm:

pgprot = __pgprot(pgprot_val(PAGE_KERNEL) & ~_PAGE_GLOBAL);
ptep = get_locked_pte(text_poke_mm, text_poke_mm_addr, &ptl);

local_irq_save(flags);

pte = mk_pte(pages[0], pgprot);
set_pte_at(text_poke_mm, text_poke_mm_addr, ptep, pte);
if (cross_page_boundary) {
        pte = mk_pte(pages[1], pgprot);
        set_pte_at(text_poke_mm, text_poke_mm_addr + PAGE_SIZE, ptep + 1, pte);
}

To perform the write, the current CPU core switches onto the private address space by calling use_temporary_mm(), saving the active mm context for later restoration. Writing to the %cr3 control register alters the active virtual-to-physical translations on this specific core. However, because modern processors employ deep, out-of-order execution pipelines, instructions situated ahead of the %cr3 write might already be fetched, decoded, or speculatively executed under the previous page translations. If left uncoordinated, execution could proceed with stale translations, resolving memory access under the wrong mapping context.

The x86 architecture mitigates this hazard by defining writes to control registers (including %cr0, %cr3, %cr4, and %dr8) as serializing instructions. The CPU core must retire all preceding instructions, discard any speculative instructions in flight, and flush non-global TLB entries before starting execution under the new register state. This serialization guarantee is an inherent property of the instruction set architecture (ISA). Consequently, loading the %cr3 register ensures that subsequent instructions see the new page-table entry before fetching memory through it, requiring no further synchronization:

prev_mm = use_temporary_mm(text_poke_mm);

This context switch must also address a secondary hazard involving hardware breakpoints and watchpoints. The debug registers (%dr0 through %dr3) are global processor state rather than thread-specific or address-space-scoped entities. Thus, any active watchpoints remain armed across the address-space switch, regardless of register serialization.

The target address text_poke_mm_addr resides in the lower, user-range half of the address space. If a user-mode debugger has registered a watchpoint that overlaps with this numeric address, the processor would trigger a debug exception mid-write, right in the middle of the code-patching sequence. This would result in a misdelivered signal to the user process or, worse, interrupt the critical patching operation.

To prevent this collision, use_temporary_mm() disables hardware breakpoints immediately after switching the address space by calling hw_breakpoint_disable(). The counterpart function hw_breakpoint_restore() restores the breakpoint state after patching is complete. This temporary disablement suppresses all breakpoints globally on the current core, including kernel-space breakpoints registered by tools such as perf.

With the writable alias in place, the core executes the copy by invoking func at the target address offset:

func((u8 *)text_poke_mm_addr + offset_in_page(addr), src, len);

For standard text_poke() invocations, func corresponds to text_poke_memcpy(), which wraps the inline copy with architectural overrides (the _set variant instead passes text_poke_memset()):

static void text_poke_memcpy(void *dst, const void *src, size_t len)
{
        lass_stac();
        __inline_memcpy(dst, src, len);
        lass_clac();
}

This design addresses strict architectural checks on user-range accesses and build-time verification rules, which dictate the use of inlined memory operations.

The transient mapping is dismantled in the reverse order of its creation. Calling pte_clear() deletes the page-table entries, while unuse_temporary_mm() switches %cr3 back to the original address space (enforcing another CPU serialization). The local TLB entry is then invalidated using flush_tlb_mm_range().

For standard text_poke() operations, the kernel validates the patch by reading back the modified bytes and performing a comparison against the source buffer. Any discrepancy triggers a BUG() panic, preventing the processor from executing corrupt or unintended instructions:

pte_clear(text_poke_mm, text_poke_mm_addr, ptep);
if (cross_page_boundary)
        pte_clear(text_poke_mm, text_poke_mm_addr + PAGE_SIZE, ptep + 1);

unuse_temporary_mm(prev_mm);
flush_tlb_mm_range(text_poke_mm, text_poke_mm_addr, text_poke_mm_addr +
                   (cross_page_boundary ? 2 : 1) * PAGE_SIZE, PAGE_SHIFT, false);

if (func == text_poke_memcpy)
        BUG_ON(memcmp(addr, src, len));

local_irq_restore(flags);

Put together, this is one physical page reached through two different virtual addresses with two different permissions:

                          physical page (the actual RAM
                          holding the instruction bytes)
                                    ^        ^
                                    |        |
              normal kernel         |        |   text_poke_mm_addr
              mapping (all CPUs,    |        |   (this CPU only,
              always present)       |        |    exists briefly)
                     |               \      /            |
                     v                \    /             v
          .text  [ RO, executable ]    \  /    [ RW, not executable ]
          (the %cr3 of each CPU         \/     (only the %cr3 of this CPU
           maps this address,                    maps this address,
           forever)                               only while patching)

The permanent mapping of .text remains read-only across all CPU cores. Only the active patching core gains transient, local access to the writable alias, and this access is restricted to the duration of __text_poke(). Once the page-table entry is cleared and the local TLB is flushed, the writable alias is completely removed, leaving only the updated read-only .text mapping. The entire text poking process is serialized by text_mutex, ensuring that concurrent threads never race to establish competing temporary mappings.

Hardware details — why SMAP, LASS, and inline copying dictate the implementation. Modern processors implement security mechanisms designed to prevent the kernel from accessing user-range virtual addresses by accident. These features guard against kernel vulnerabilities where a corrupted or attacker-controlled pointer is dereferenced within supervisor mode. Two hardware-level protections enforce these boundaries using distinct criteria:

  • SMAP (“Supervisor Mode Access Prevention”) generates a page fault if kernel code attempts to access a virtual page whose page-table entry has the _PAGE_USER bit set. This mechanism evaluates only the user permission bit of the translation entry, ignoring the numeric value of the virtual address.
  • LASS (“Linear Address Space Separation”) blocks supervisor-mode accesses to any virtual address falling under the user address space. This protection relies entirely on the numeric range of the address, regardless of whether _PAGE_USER is configured on the page.

The address text_poke_mm_addr is allocated within the lower half of the virtual address space, corresponding to the range where user processes receive mappings from mmap(). The virtual address is allocated via mm_alloc() at boot, which initializes the structure at TASK_UNMAPPED_BASE with a randomized offset. However, because the page-table entry built for this mapping has the _PAGE_USER bit cleared, the page is not user-accessible.

This configuration interacts differently with each protection mechanism. Because the _PAGE_USER bit remains clear, SMAP does not flag the access. However, because the virtual address numerically resides in the user-space range, LASS would trigger an immediate supervisor-mode page fault.

To bypass this restriction, the kernel invokes lass_stac() and lass_clac() around the copy. These functions check for the presence of X86_FEATURE_LASS on the processor; if enabled, they temporarily toggle the alignment check (AC) flag in the %rflags register, instructing the hardware to permit the access. On processors supporting only X86_FEATURE_SMAP, or neither feature, these operations resolve to no-ops or default behavior. This is conceptually identical to the user-access window established by copy_from_user().

Enabling the AC override imposes a critical constraint on the code. The kernel build-time analysis tool, objtool, enforces a strict rule prohibiting any instruction calling another function between a STAC and a CLAC instruction. The AC flag is an active CPU register state but is not automatically saved or restored during a task context switch. If the code inside the override window executes a function call that eventually yields the CPU via schedule(), the AC flag remains set. This would leak the user-access permission into the next scheduled task, compromising system security.

This build-time rule is the reason the patching copy cannot utilize the standard library implementation of memcpy(). On x86_64, memcpy() is an assembly routine defined in a separate object file, necessitating a function call. To adhere to the objtool constraint, the patching sequence employs __inline_memcpy() and __inline_memset(), which compile directly into inline rep movsb and rep stosb assembly instructions. This removes the function call entirely, satisfying the safety validation of objtool.


6 What the compiler emits (x86_64)

Every static branch call site contains either a nop or a jmp instruction of identical length. Along with this inline instruction, the compilation process generates a corresponding metadata entry in the __jump_table section. The choice between a nop and a jmp at compile time depends on the initial state of the static key and the branch hint at the call site. Two architecture-specific macros generate these instructions and populate the sidecar table.

6.1 TRUE/FALSE keys and likely/unlikely sites

The compiler must emit a concrete, valid machine instruction at every call site before jump_label_init() runs at boot. Because compilation occurs statically, the compiler cannot evaluate the live enabled reference count of a key. Instead, the generated instruction is determined by two compile-time variables:

  1. The default key state — Declared via DEFINE_STATIC_KEY_TRUE or DEFINE_STATIC_KEY_FALSE.
  2. The call-site direction hint — Specified using static_branch_likely() or static_branch_unlikely().

When the default state of the key aligns with the call-site direction hint, the compiler emits a nop instruction. When the default state and the direction hint mismatch, the compiler emits a jmp instruction. Disaligning these variables in relation to the expected steady state of the static key—the hazard described in §4.2—causes the hot path to start with a branch jump, which persists until an explicit toggle occurs at runtime.

The reference documentation at the top of include/linux/jump_label.h maps out this structural matrix:

                likely()                    unlikely()
                --------                    ----------
 key=true       ...                         ...
                NOP                         JMP L
                <br-stmts>               1: ...
            L:  ...
                                         L: <br-stmts>
                                            jmp 1b
 ------------------------------------------------------------
 key=false      ...                         ...
                JMP L                       NOP
                <br-stmts>               1: ...
            L:  ...
                                         L: <br-stmts>
                                            jmp 1b

The runtime text patcher calculates the destination state using the same Boolean logic, substituting the live state of the key for the compile-time default state. The inline function jump_label_type() implements this evaluation at runtime. An identical function, jump_label_init_type(), performs the calculation during initialization by reading the static, compile-time key type bit. The logical branch state evaluates to 1 for a likely site and 0 for an unlikely site. The resulting compile-time XOR operation translates to type ^ branch. The live implementation evaluates the active state of the branch:

static enum jump_label_type jump_label_type(struct jump_entry *entry)
{
        struct static_key *key = jump_entry_key(entry);
        bool enabled = static_key_enabled(key);
        bool branch = jump_entry_is_branch(entry);

        return enabled ^ branch;
}

6.2 How the macros pick the asm

The high-level macros static_branch_likely() and static_branch_unlikely() do not generate or manipulate instruction bytes directly. Instead, they delegate to two architecture-specific helper functions:

Both helpers communicate the path that execution took through the patched instruction. A nop instruction falls through to the next sequential instruction, causing the helper to return false. A jmp instruction transfers execution out-of-line, causing the helper to return true. The implementation details of the underlying asm goto statement are examined in §6.3.

The Boolean return value reflects the active instruction shape at the patch site. The logical association with the branch path of the caller is resolved through the negation logic implemented in the macro definitions.

The four possible combinations of key type and call-site hint map onto these two low-level helpers. The helper selection is determined by the compile-time type of the static key, while the optional negation of the return value is dictated by the direction hint:

/* CONFIG_JUMP_LABEL path in jump_label.h */
static_branch_likely(x):
  TRUE  key → !arch_static_branch(&(x)->key, true)       /* nop-default site */
  FALSE key → !arch_static_branch_jump(&(x)->key, true)  /* jmp-default site */

static_branch_unlikely(x):
  TRUE  key →  arch_static_branch_jump(&(x)->key, false) /* jmp-default site */
  FALSE key →  arch_static_branch(&(x)->key, false)      /* nop-default site */

Here, arch_static_branch() is invoked for combinations that compile to a nop, whereas arch_static_branch_jump() is utilized for combinations that compile to a jmp.

The unary negation operator (!) preceding the two likely branches maps the helper output to the expected conditional behavior of the caller. The helpers return true or false based on whether a branch jump occurred, rather than whether the static key is enabled.

For example, reading a TRUE key with likely invokes the nop-default helper. A nop instruction falls through, so arch_static_branch() returns false (indicating that no branch jump occurred), even though the conditional body must execute. Applying the unary negation operator converts this result into true, ensuring correct control flow.

The Boolean argument passed to the helper is stored directly inside the metadata of the jump table, as detailed in §6.4. The runtime text-patching engine evaluates this metadata whenever the state of the key is toggled.

The compile-time type selection uses __builtin_types_compatible_p to distinguish between struct static_key_true and struct static_key_false. Any unsupported type defaults to an unresolved call to ____wrong_branch_error(), which triggers a compilation failure.

6.3 The two asm helpers

While the macro expansion treats arch_static_branch() and arch_static_branch_jump() as opaque functions returning a Boolean value, their underlying implementations in arch/x86/include/asm/jump_label.h expose inline assembly:

static __always_inline bool arch_static_branch(struct static_key * const key,
                                               const bool branch)
{
        asm goto(ARCH_STATIC_BRANCH_ASM("%c0 + %c1", "%l[l_yes]")
                : :  "i" (key), "i" (branch) : : l_yes);

        return false;
l_yes:
        return true;
}

static __always_inline bool arch_static_branch_jump(struct static_key * const key,
                                                    const bool branch)
{
        asm goto("1:"
                "jmp %l[l_yes]\n\t"
                JUMP_TABLE_ENTRY("%c0 + %c1", "%l[l_yes]")
                : :  "i" (key), "i" (branch) : : l_yes);

        return false;
l_yes:
        return true;
}

The use of two separate inline functions, rather than a single function parameterized with a branch selection flag—such as a speculative arch_static_branch(key, branch, use_jmp)—is dictated by compile-time constraints. Because the assembly payload is evaluated during compilation, a runtime argument cannot dynamically alter the instruction bytes embedded within the function body. The compiled instructions are finalized when the translation unit is processed.

Consequently, separate functions encapsulate each instruction layout, with the logical branch path selected by the macros described in §6.2. This selection is resolved statically based on the compile-time type of the static key.

The function bodies rely on inline assembly, specifically the GCC asm goto extension. The functional components of the asm goto blocks provide several key mechanisms:

With these syntax rules defined, the structure of arch_static_branch_jump() is highly transparent, passing the following template to asm goto:

1:
jmp %l[l_yes]
<jump table entry for this site, via JUMP_TABLE_ENTRY>

The companion function arch_static_branch() is constructed analogously, but delegates the generation of the local label and instruction to the ARCH_STATIC_BRANCH_ASM macro:

#ifdef CONFIG_HAVE_JUMP_LABEL_HACK
#define ARCH_STATIC_BRANCH_ASM(key, label)              \
        "1: jmp " label " # `objtool` NOPs this \n\t"   \
        JUMP_TABLE_ENTRY(key " + 2", label)
#else /* !CONFIG_HAVE_JUMP_LABEL_HACK */
#define ARCH_STATIC_BRANCH_ASM(key, label)              \
      "1: .byte " __stringify(BYTES_NOP5) "\n\t"        \
      JUMP_TABLE_ENTRY(key, label)
#endif /* CONFIG_HAVE_JUMP_LABEL_HACK */

This preprocessor conditional is resolved at build time based on HAVE_JUMP_LABEL_HACK. Each compilation path outputs different assembly representations:

The operand declaration : : \"i\" (key), \"i\" (branch) : : l_yes); binds the C variables to the assembly template. Defining the inputs with the "i" constraint forces the compiler to resolve these parameters as compile-time constants. This constraint ensures that the metadata values remain static and available during the assembly phase, even after the compiler has inline-expanded the containing functions. The clobber list is empty, and the goto-label list specifies the branch target l_yes.

6.4 The jump table entry (sidecar metadata)

Along with the inline instruction, each branch site emits a metadata descriptor of type struct jump_entry into the __jump_table section. The JUMP_TABLE_ENTRY macro generates this metadata without outputting any executable CPU instructions:

#define JUMP_TABLE_ENTRY(key, label)                   \
        ".pushsection __jump_table,  \"aw\" \n\t"      \
        _ASM_ALIGN "\n\t"                              \
        ANNOTATE_DATA_SPECIAL "\n"                     \
        ".long 1b - . \n\t"                            \
        ".long " label " - . \n\t"                     \
        _ASM_PTR " " key " - . \n\t"                   \
        ".popsection \n\t"
Field Asm Meaning
code .long 1b - . relative offset to the patchable insn
target .long label - . relative offset to the l_yes target
key _ASM_PTR key - . relative offset to the static_key, low bits = flags

Every directive in this macro instructs the assembler to format data rather than producing executable machine instructions. Each directive is evaluated by the assembler according to specific rules:

The fields code and target utilize 32-bit displacements via .long, whereas key requires the full pointer width of _ASM_PTR. Because the patchable instruction and the target block reside within the same function body, a 32-bit offset is guaranteed to reach the target. Conversely, the target struct static_key may be located far from the call site under KASLR or within a separate kernel module, requiring a full-width relocation.

Each field is encoded as a self-relative offset, meaning the distance is measured from the address of the metadata field itself rather than from the beginning of the structure or the array. This design allows each offset to be resolved using a uniform address calculation.

For the code field:

   __jump_table[i].code lives at a memory address designated as F.

   The patchable nop/jmp instruction in .text resides at address C.

   The value written into the code field at assembly time represents the distance:

       stored value  =  C - F

   At runtime, the address of the instruction C is resolved by referencing 
   the address of the field F:

       C  =  &entry->code  +  entry->code
              ^^^^^^^^^^^^    ^^^^^^^^^^^^
              own address      the distance
              of this field    stored in it

This addition represents the implementation of jump_entry_code().

The fields for .target (resolving the branch destination) and .key (resolving the target key) are computed using the same mechanism, evaluating the offsets relative to their own field addresses. Thus, jump_entry_target() evaluates as &entry->target + entry->target, while jump_entry_key() masks out the flag bits and computes the address relative to &entry->key. Measurement relative to individual fields eliminates the need for field-specific offsets in the retrieval logic.

This relative layout also ensures that the metadata survives KASLR relocations with no boot-time overhead. If the kernel image is shifted by a constant offset at boot, both the field address F and the target address C scale by the identical offset. The relative distance C - F remains invariant, which allows jump_label_init() to read the jump table immediately without performing pointer relocation.

6.5 HAVE_JUMP_LABEL_HACK: why sites are 2 or 5 bytes

The nop-default site of arch_static_branch() is defined using a hand-coded .byte BYTES_NOP5 (representing a fixed 5-byte NOP) under certain configurations. However, this is not the instruction size that is guaranteed to land in a compiled kernel image. Depending on the properties of the call site, the NOP that resides in memory is either a compact 2-byte instruction or a full 5-byte instruction. The resolution of this size is settled before boot: the compilation process outputs a real jmp instruction, allowing the compiler to select the shortest valid encoding, and a subsequent build step converts this jmp into a NOP of identical width before the kernel image is finalized.

This optimization mechanism is gated by the HAVE_JUMP_LABEL_HACK configuration option. The architecture configuration file arch/x86/Kconfig enables this option for any build where objtool is available via select HAVE_JUMP_LABEL_HACK if HAVE_OBJTOOL (which is satisfied on all modern x86_64 builds). The preprocessor directive ARCH_STATIC_BRANCH_ASM resolves this compile-time branch.

This optimization solves a compiler limitation: generating a hard-coded .byte BYTES_NOP5 always forces a 5-byte NOP, even when a 2-byte instruction is sufficient, wasting instruction cache space at every nop-default call site. The alternative is to let the compiler emit a correctly sized jmp instruction, as the compiler can evaluate the shortest encoding needed to reach the destination target. The build process then converts this jmp into a same-sized NOP.

With the optimization hack enabled:

  1. The compiler emits a real jmp to l_yes as it would for an ordinary conditional branch. The assembler selects between two unconditional jump encodings based on the displacement distance. JMP rel8 uses a 1-byte opcode (0xEB) and a signed 1-byte relative displacement (2 bytes total), which is valid if the branch target resides within -128 to +127 bytes of the next instruction. JMP rel32 uses a 1-byte opcode (0xE9) and a 4-byte signed displacement (5 bytes total), which can reach any location within a 32-bit relative displacement.1 The terms “rel8” and “rel32” specify the bit width of the displacement field. The assembler calculates the real distance to the label and selects the 2-byte rel8 format when the target is close, falling back to the 5-byte rel32 format only when necessary. This selection ensures that the instruction size is optimal for each call site.
  2. The key expression passed to the jump table macro is defined as "%c0 + %c1 + 2", which sets bit 1 of the stored key address. This bit acts as a metadata marker that is processed during the subsequent build stage rather than being evaluated at runtime.
  3. objtool (handle_jump_alt()) processes the compiled object files during the build phase. When it detects that bit 1 of the key address is set, it overwrites the corresponding jmp instruction with a same-sized NOP and clears the associated relocation entry. This transformation is completed entirely at build time. When the kernel boots, the call sites already contain valid NOP instructions. Once this transformation is complete, jump_label_init() repurposes bit 1 of the key address to store the init-text flag via jump_entry_set_init(), as the build-time marker is no longer required.

By the time the kernel image is finalized, every nop-default call site contains a valid NOP instruction optimized to the smallest possible width: 2 bytes for sites where the original jmp utilized the rel8 format, or 5 bytes where it utilized rel32. The instruction size is fixed during the build phase and inherited at boot.

However, the struct jump_entry metadata descriptor does not record the instruction size of each site. Storing this size is impractical because the instruction width is finalized during the objtool pass, which occurs after the compiler has laid out the structure fields.

Consequently, when the patching engine toggles a site at runtime—for example, during a call to static_branch_enable() or static_branch_disable()—the patching logic must examine the instruction bytes in memory to dynamically resolve whether a 2-byte or 5-byte instruction is present. It then generates a replacement instruction of identical width to ensure that the surrounding instruction stream remains aligned. This decoding and patching pipeline—implemented via arch_jump_entry_size(), insn_decode_kernel(), and __jump_label_patch()—is detailed in §8. The critical detail is that the instruction size established by objtool is discovered and preserved during patching, preventing code relocation.

In configurations where the optimization hack is disabled (such as older toolchains lacking objtool integration), the compiler bypasses this pipeline. It compiles the #else branch of the assembly template, emitting a fixed 5-byte NOP via .byte BYTES_NOP5 regardless of target proximity. Later runtime patching always replaces this NOP with a 5-byte jump instruction. While functional, this fallback increases the instruction footprint by 3 bytes per nop-default call site compared to optimized builds.

The build-time optimization only applies to nop-default branches generated via the arch_static_branch() helper. The jmp-default helper, arch_static_branch_jump(), generates its asm goto directly, passing the key expression to JUMP_TABLE_ENTRY as a plain \"%c0 + %c1\" without the additional offset of 2.

Without this metadata flag, objtool does not modify the jump instruction, which reaches boot as a real conditional jump that is patched only when the key state is disabled. This instruction is still compiled using either the 2-byte rel8 or 5-byte rel32 layout depending on assembler-level distance calculation. The optimization hack alters whether objtool transforms the instruction post-compilation, rather than how the initial jump instruction is encoded by the assembler.

Bit 1 of the key field—the metadata segment masked off by jump_entry_key() to retrieve the underlying struct static_key pointer—serves two independent purposes across the build and boot boundary:

  build phase  -------------------------------->  boot phase  ---> execution

  bit 1 = "objtool: NOP this jmp"        jump_label_init()
  (processed and discarded by            overwrites the bit to represent:
   objtool at build time)                       bit 1 = "jump_entry_is_init"
                                         (marks the site as init-only text)

The boot-time flag is initialized using jump_entry_set_init() during the execution of jump_label_init(), and is subsequently read using jump_entry_is_init(). This dual usage is a common source of confusion when analyzing the initialization sequence.

6.6 Assembly-level picture

The integration of the C helper functions, the sidecar jump-table metadata, and the build-time instruction width optimization results in a cohesive machine-level layout. For a nop-default call site on a build configured with HAVE_JUMP_LABEL_HACK (the standard configuration on x86_64), the compiled object file contains the following representation:

.text:
        1:  0f 1f 44 00 00  ; 5-byte NOP if l_yes was far, OR 66 90 (2-byte) if
            ...             ; close (objtool rewrote a real jmp into this
                            ; at build time)

__jump_table:                   ; non-executable metadata
        .long   1b - .          ; code:   self-relative offset to the NOP instruction
        .long   L - .           ; target: self-relative offset to l_yes
        .quad   key+branch+2 - .; key: static_key address, branch bit 0,
                                ;      plus the objtool-only "+2" signal

The two potential byte sequences representing the inactive state of the instruction at the 1: label represent distinct machine instructions rather than arbitrary padding. The 5-byte sequence 0f 1f 44 00 00 represents the multi-byte NOP (nopl 0x0(%rax,%rax,1)), which incorporates an unused addressing mode to achieve a 5-byte instruction width. The 2-byte alternative 66 90 incorporates the operand-size override prefix (0x66) stacked in front of the classic single-byte NOP opcode (0x90, historically executing as xchg %ax,%ax). The prefix is added to pad the instruction to exactly 2 bytes without modifying execution behavior. Both instructions function as genuine no-ops: the CPU decodes the instruction, consumes execution cycles, and falls through to the next sequential instruction, leaving all architectural registers and flags unmodified.

Two specific aspects of this layout merit close examination:

The x86 architecture also defines HAVE_JUMP_LABEL_BATCH, which enables the batched and synchronized text-patching path. Without this batched optimization, every individual jump-table entry would require separate instruction patching and global CPU synchronization, introducing significant performance overhead.


7 Core data structures

Runtime management of dynamic patching requires cooperative interaction between active state representation and static compiler metadata. The core subsystem models this relationship through unified control structures and metadata records that catalog every patching target across the system. These components link individual branch sites back to the central keys that govern them.

7.1 struct static_key

Every static key declared via DEFINE_STATIC_KEY_TRUE or DEFINE_STATIC_KEY_FALSE boils down to a single runtime representation defined by struct static_key. Compile-time type wrappers maintain logical distinction during compilation, but they resolve to this identical underlying structure at runtime.

struct static_key {
        atomic_t enabled;
#ifdef CONFIG_JUMP_LABEL
        union {
                unsigned long type;
                struct jump_entry *entries;
                struct static_key_mod *next;
        };
#endif
};

The enabled field represents the active control state of the key. It is declared as an atomic_t reference counter where a value of zero indicates the disabled state and any positive value indicates the enabled state. This design allows static_branch_inc and static_branch_dec to coordinate multiple stacked owners on a single key. During transitions, the field can temporarily hold the value -1, indicating that an initial call to static_key_slow_inc is actively patching the instruction stream. To prevent confusing readers during this transition window, static_key_count maps -1 back to 1 before returning.

The second member of the struct is a tagged union residing in a single machine word. The low two bits of the word act as metadata flags. This optimization is possible because pointers to struct jump_entry and struct static_key_mod are at least 4-byte aligned on supported architectures. The alignment guarantees that the two least significant bits of any valid pointer address are zero, leaving them available for bitwise tagging.

Bit Macro Meaning
0 JUMP_TYPE_TRUE compile-time initial value of the key was true — this is the same type bit the type ^ branch formula from §6.1 uses
1 JUMP_TYPE_LINKED 1: the rest of the word is a next pointer (a linked list); 0: it is an entries pointer (a flat array)

This bit-packing strategy must not be confused with the tagging applied to the key field of struct jump_entry. While both employ similar bitwise operations on pointers, they serve different purposes. One tags the initial state of the key and the structure of the associated call-site list, while the other indicates the branch hint and init-section status of a specific call site.

Bit 1 handles cases where call sites are distributed across separate compilation boundaries. For keys utilized solely within the core kernel image, the linker arranges all associated call sites into a single contiguous array inside the __jump_table section of vmlinux. The entries pointer then points directly to the start of this sequence. However, kernel modules loaded at runtime carry private __jump_table sections that cannot be merged post-link. Since modules are loaded and unloaded dynamically, the active set of call sites must grow and shrink. When a module introduces new call sites for an existing key, the representation transitions. The word shifts from a direct pointer to the head of a linked list composed of struct static_key_mod nodes, where each node tracks the contribution of a specific module.

struct static_key_mod {
        struct static_key_mod *next;
        struct jump_entry *entries;
        struct module *mod;
};

Subsystems outside the core jump label implementation in kernel/jump_label.c do not interact with these raw bits directly. Accessor helpers like static_key_entries, static_key_type, static_key_linked, and static_key_set_entries mask off the tag bits before returning pointers, insulating the rest of the kernel from the underlying representation.

The type wrappers struct static_key_true and struct static_key_false allow the compiler to distinguish key polarities at compile time. Each contains a single struct static_key member:

struct static_key_true  { struct static_key key; };
struct static_key_false { struct static_key key; };

These distinct C types allow the preprocessor and compiler to select the appropriate branch behaviors. The macros static_branch_likely and static_branch_unlikely use compiler builtins to inspect the type of the passed key and dispatch execution to the correct architecture-specific code generation path.

7.2 struct jump_entry (relative form)

This structure acts as the metadata record generated for each jump site. The JUMP_TABLE_ENTRY assembler macro writes one instance of this record per call site into the __jump_table section. The x86 architecture opts into a relative variant of this metadata structure by selecting the configuration option HAVE_ARCH_JUMP_LABEL_RELATIVE. This configuration alters the layout of struct jump_entry:

struct jump_entry {
        s32 code;
        s32 target;
        long key;       /* full width: module↔vmlinux may be far under KASLR */
};

The code and target fields contain self-relative offset distances rather than absolute pointer addresses. Reconstructing the absolute address of the patchable instruction is achieved by adding the value of code to the address of the code field itself. The target destination address is reconstructed similarly. The key field, after masking off the two least significant bits, resolves to the address of the governing static_key using the same self-relative offset arithmetic.

Architectures that do not select HAVE_ARCH_JUMP_LABEL_RELATIVE employ a standard fallback structure where the fields store absolute addresses. This distinction is visible in the accessors of the fallback structure (where jump_entry_code simply performs a direct pointer return with no arithmetic). On x86, paying the minor computational cost of the addition on every lookup yields significant advantages:

  1. Size Optimization: A running kernel contains tens of thousands of individual branch sites. Storing offsets as 32-bit signed integers (s32) instead of full-width pointer values (8 bytes on x86_64) halves the size of two-thirds of each entry, substantially reducing the memory footprint of the total table.
  2. KASLR Compatibility: Because both the patchable code site and the target label reside within the boundaries of a single compiled function, a 32-bit signed offset is always sufficient to span the distance. The absolute address is resolved through relocation-free arithmetic, avoiding boot-time relocation fixups under Kernel Address Space Layout Randomization (KASLR). In contrast, the key field is kept at full pointer width (long instead of s32) because it references a static_key which can reside anywhere in the address space — including inside a different kernel module or the main vmlinux binary. Because the distance between a dynamically loaded module and the core kernel image can exceed the range of a 32-bit signed integer, this field must maintain full pointer width.

The two least significant bits of the key pointer are reserved for encoding additional metadata. This pointer tagging strategy conveys two distinct pieces of information:

Bit Meaning
0 jump_entry_is_branch: Represents the branch direction hint. A value of 1 indicates that the call site uses likely(), while 0 indicates unlikely().
1 jump_entry_is_init: Indicates whether the target instruction resides within an initialization section (__init text). Such sections are freed after boot, rendering the associated call sites unpatchable from that point forward.

The inline accessors perform direct bitwise checks to retrieve these flags:

static inline bool jump_entry_is_branch(const struct jump_entry *entry)
{
        return (unsigned long)entry->key & 1UL;
}

static inline bool jump_entry_is_init(const struct jump_entry *entry)
{
        return (unsigned long)entry->key & 2UL;
}

Bit 1 is the same marker tracked through two distinct lifetimes. During the initial build phase, it signals to objtool that a given jump instruction must be converted to a NOP. After jump_label_init processes this instruction and completes boot-time setup, the bit assumes the permanent meaning as the initialization-text flag for the remainder of the kernel uptime.

7.3 Relationship

A standard struct static_key holds a direct entries pointer (the common, non-module-linked case) referencing the associated call sites in memory:

        struct static_key
        +------------------------+
        | enabled (atomic)       |
        | type / entries / next  |   tagged union
        +-----------+------------+
                    |
                    | first jump_entry for this key
                    v
__jump_table[]  (sorted by key, then by code address)

  [ entries for key A ... ][ entries for key B ... ] ...
        |                           |
        +--> code / target / key <--+

The struct static_key structure does not store an explicit count of referencing call sites; it retains only a pointer to the first associated struct jump_entry. This single pointer is sufficient because of the ordering established in the metadata array. The function jump_label_sort_entries sorts the whole __jump_table array. Sorting is performed not on the raw bits stored in the key field, but on the decoded, absolute target address returned by jump_entry_key.

This decoding is necessary because, like the code and target fields, the key field is self-relative. Rather than storing a direct absolute address, it holds the distance to the target key, measured from the own address of the field within the table entry. The inline helper function jump_entry_key masks off the two low-order metadata flags to isolate the offset, then adds the address of the key field itself to reconstruct the absolute address of the governing key:

static inline struct static_key *jump_entry_key(const struct jump_entry *entry)
{
        long offset = entry->key & ~3L;

        return (struct static_key *)((unsigned long)&entry->key + offset);
}

Consequently, two distinct table entries referring to the same central key but residing at different offsets within the table will measure distances from different starting points. Because the base address of each field differs per slot, the raw bit patterns stored in the key fields will differ even though they resolve to the same underlying control structure:

              addr of      raw key    decode: addr + raw key
              entry->key   field      (jump_entry_key())
              ----------   --------   -----------------------
slot 0 (A):   0x1000       +0x4000    0x1000 + 0x4000 = 0x5000  ─┐
   ...                                                           ├─ same
slot 5 (B):   0x2000       +0x3000    0x2000 + 0x3000 = 0x5000  ─┘  static_key!

                                           static_key @ 0x5000

While the raw values 0x4000 and 0x3000 share no common bit pattern, factoring in the absolute address of each entry yields the identical result 0x5000. The sorting function uses this absolute address to order the table.

Once sorted, all entries associated with a given key occupy a contiguous sequence in memory. Locating every call site for a specific key requires starting at the address specified by the entries pointer of the key and scanning forward until the decoded key address changes. This layout eliminates the need to maintain an explicit count or index of associated call sites.

Within each contiguous key run, a secondary sort is performed using jump_entry_code to arrange the patchable instructions in ascending order of memory addresses. The instruction patching system requires this monotonic ordering to optimize batch patching operations and ensure predictable execution flows.

Sorting relative-offset metadata requires specialized swap logic. A standard sorting algorithm swaps array elements through a byte-for-byte copy. However, because the fields of a relative entry encode distances computed relative to the own address of the entry, moving the record to a different slot without adjusting the fields would corrupt the pointers.

The helper function jump_label_swap prevents this corruption. During a swap, the function calculates the distance delta between the source and destination slots. It then adjusts the offsets in each field by adding or subtracting delta so that each relative field, now residing at a new address, continues to resolve to the same absolute target:

static void jump_label_swap(void *a, void *b, int size)
{
        long delta = (unsigned long)a - (unsigned long)b;
        struct jump_entry *jea = a;
        struct jump_entry *jeb = b;
        struct jump_entry tmp = *jea;

        jea->code   = jeb->code - delta;
        jea->target = jeb->target - delta;
        jea->key    = jeb->key - delta;

        jeb->code   = tmp.code + delta;
        jeb->target = tmp.target + delta;
        jeb->key    = tmp.key + delta;
}

Architectures that do not use the relative form of struct jump_entry are immune to this issue. Because absolute addresses remain valid regardless of where the containing entry resides, those architectures can rely on standard byte-for-byte swaps.

7.4 Linker section

Each translation unit expanding the JUMP_TABLE_ENTRY assembler macro emits a dedicated .pushsection __jump_table … .popsection block. These metadata fragments are scattered across numerous compiled object files during compilation. To form a cohesive, contiguous table, the linker collects these fragments during the final link phase of the kernel image. The core kernel linker script directs this collation using the macro BOUNDED_SECTION_BY defined in include/asm-generic/vmlinux.lds.h:

BOUNDED_SECTION_BY(__jump_table, ___jump_table)

This macro expands to three linker directives:

__start___jump_table = .;
KEEP(*(__jump_table))
__stop___jump_table = .;

The directive *(__jump_table) instructs the linker to extract the __jump_table input section from every compiled object file and arrange them sequentially in physical memory. This process concatenates the independently generated metadata entries into a unified array, leveraging the same linker-driven section aggregation mechanism utilized for system initialization calls.

The KEEP modifier is crucial because the C code of the kernel does not directly reference individual elements within this section or invoke them like standard code symbols. Under aggressive dead-code elimination optimizations, the linker might categorize the section as unused and discard it. The KEEP instruction explicitly overrides this behavior, forcing the linker to preserve the accumulated table.

The surrounding assignments __start___jump_table and __stop___jump_table define the boundaries of the resulting array. During boot-time initialization, jump_label_init references these linker-defined symbols to locate the table in memory. Because these markers provide precise boundary addresses, the subsystem does not require an explicit compile-time count of table entries.

This linker-driven aggregation is confined to the static vmlinux binary. A kernel module loaded dynamically at runtime cannot participate in the link phase of the core kernel. Instead, each module maintains a private __jump_table section within the associated ELF object. When the module loading subsystem maps a module into memory, the loader reads this metadata and records the boundary addresses in two fields reserved inside struct module: jump_entries stores the base address of the array, and num_jump_entries stores the number of active entries. When a module shares a key with the main kernel or another module, these dynamically mapped entries are encapsulated inside struct static_key_mod nodes to integrate them into the central patching system.


8 Size of the patchable site on x86 (runtime)

At runtime, the instruction-patching subsystem on x86 must dynamically determine the size of each patchable site before performing any modification. Build-time optimizations discussed in §6.5 allow the assembler to emit either a 2-byte or a 5-byte instruction at a jump site, depending on the relative displacement to the target. However, struct jump_entry (detailed in §7.2) contains no field or metadata recording the length of the instruction that the assembler selected. When the kernel modifies a jump label on a running system, the patching engine must rediscover the instruction length by parsing the live machine bytes currently sitting in the executable memory.

The kernel delegates this discovery to arch_jump_entry_size():

/* arch/x86/kernel/jump_label.c */
int arch_jump_entry_size(struct jump_entry *entry)
{
        struct insn insn = {};

        insn_decode_kernel(&insn, (void *)jump_entry_code(entry));
        BUG_ON(insn.length != 2 && insn.length != 5);
        return insn.length;
}

The retrieval of the target address relies on jump_entry_code() (explained in §6.4). The x86 instruction decoder of the kernel, insn_decode_kernel(), parses the machine code at that location into struct insn. Rather than performing a simple byte count, the decoder fully parses the opcode, prefixes, and displacement to determine the exact boundary of the instruction, reporting the result in the length field of the structure. Whether the memory location contains the compiler-generated default nop, a branch instruction, or a previously patched instruction from an earlier state transition, the decoder resolves the true instruction boundaries.

A sanity check via the BUG_ON() macro, evaluating BUG_ON(insn.length != 2 && insn.length != 5), guards against corruption. If the decoder encounters any length other than these two supported sizes, the metadata table has desynchronized from the executable stream, making further patching unsafe.

This dynamic decoding step is unique to the x86 architecture. Architectures with fixed instruction sizes, such as arm64, define a constant JUMP_LABEL_NOP_SIZE. On those platforms, every patchable site shares a uniform, known width, which eliminates the need for runtime discovery. The generic, architecture-independent jump_entry_size() helper encapsulates this architectural difference:

static inline int jump_entry_size(struct jump_entry *entry)
{
#ifdef JUMP_LABEL_NOP_SIZE
    return JUMP_LABEL_NOP_SIZE;
#else
    return arch_jump_entry_size(entry);
#endif
}

If the architecture defines a global constant size, the compiler resolves jump_entry_size() to that constant. Otherwise, the helper falls back to the dynamic decoder.

After resolving the instruction size, the helper __jump_label_patch() prepares both candidate byte sequences (the jump instruction and the corresponding no-op sequence) before selecting the sequence to install:

size = arch_jump_entry_size(entry);
switch (size) {
case JMP8_INSN_SIZE:   /* 2 */
        code = text_gen_insn(JMP8_INSN_OPCODE, addr, dest);
        nop  = x86_nops[size];
        break;
case JMP32_INSN_SIZE:  /* 5 */
        code = text_gen_insn(JMP32_INSN_OPCODE, addr, dest);
        nop  = x86_nops[size];
        break;
}

The constants JMP8_INSN_SIZE, JMP8_INSN_OPCODE, JMP32_INSN_SIZE, and JMP32_INSN_OPCODE represent the underlying instruction lengths and raw opcodes (0xEB and 0xE9) for short and near jumps on x86. The utility text_gen_insn() synthesizes the target jump sequence, calculating the self-relative offset as dest - (addr + size). This calculation mirrors the standard x86 instruction pointer offset convention, resolved here during execution rather than at compile time.

To obtain the corresponding no-op sequence, the function indexes into the lookup table x86_nops using the decoded size. This lookup supplies a pre-calculated nop sequence matching the exact width of the site without requiring instruction generation at runtime.

Generating both candidate sequences beforehand allows the kernel to perform a pre-patching safety validation. Before modifying the instruction stream, the patching engine determines the instruction sequence expected to reside in memory at the destination. This expected sequence is the logical inverse of the target sequence being installed. If the update request specifies JUMP_LABEL_JMP, the memory location must currently hold the no-op sequence. Conversely, if the update request specifies JUMP_LABEL_NOP, the memory location must hold the active jump instruction.

The patching engine invokes memcmp() to compare the live instructions in memory against this expected sequence. A mismatch indicates that the metadata table and the instruction stream have diverged somehow, suggesting memory corruption, a race condition, or an unhandled synchronization failure. In this scenario, proceeding is unsafe. The kernel reports the divergence using pr_crit() and immediately calls BUG(), halting the processor to prevent the execution of arbitrary or corrupted instructions.


9 Life of a static key: boot, enable, disable

Everything in this section is driven by one atomic_t: key->enabled. Its value tells you both “is the feature on” and “is a patch pass currently running”. The boolean on/off cycle looks like this:

        set to -1          patch, set 1          cmpxchg(1,0), off
+----+             +----+                +----+                     +----+
|  0 | ----------> | -1 | -------------> |  1 | ------------------> |  0 |
+----+             +----+                +----+                     +----+
 off               enable                  on                        off

Read left to right, the three values in that diagram mean:

There is no symmetric -1-like transient for disable: going 1 -> 0 reads as briefly stale “on” instead (see §9.3), which is harmless.

Above value 1, a second, independent ladder exists purely for refcounting — static_branch_inc()/static_branch_dec() (§4.3) climb and descend it without ever touching text, because the key is already known to be enabled:

   1 --inc--> 2 --inc--> 3 --inc--> ...        (static_key_fast_inc_not_disabled:
   1 <--dec-- 2 <--dec-- 3 <--dec-- ...          pure atomic increment/decrement,
                                                  no jump_label_update() at all)

Only a dec that would land exactly on 1 -> 0 re-enters the boolean cycle above and triggers a real patch-off. Put differently: of all the edges across both diagrams, only the three that make up the boolean cycle itself (0 -> -1, -1 -> 1, and 1 -> 0) ever touch instruction bytes. Every step on the refcount ladder above 1 (1<->2, 2<->3, …) is a bare atomic increment or decrement, with no jump_label_update() call anywhere in it.

That is why the refcounted API in §4.3 exists: many callers can share a key without each one paying for a text-patch round-trip — the cost of patching is paid exactly once, by whichever caller happens to be the first to enable it or the last to disable it.

9.1 Boot: jump_label_init()

Before this function ever runs, the jump-label machinery is in a half-built state. The linker has already concatenated the slice of __jump_table from every translation unit into one array (§7.4), but that array is simply in link order — sites for the same key can be scattered anywhere in it, and key->entries still holds whatever its static initializer left there, which for a plain struct static_key key = STATIC_KEY_INIT_FALSE; (§7.1) is nothing useful. In other words: the raw table of call sites exists, but nothing yet knows which sites belong to which key, so static_branch_enable() or static_key_slow_inc() (§4.3, §9.2) would have nothing to walk if called this early.

arch/x86/xen/multicalls.c declares static struct static_key mc_debug __ro_after_init; and registers xen_mc_debug as an early_param() whose xen_parse_mc_debug() callback calls static_key_slow_inc(&mc_debug) directly, synchronously, while parse_early_param() is walking the command line. If that increment ran before jump_label_init() had sorted the table and pointed mc_debug at its entries, it would have nothing to patch and the key would silently stay un-patched despite the user asking for it on the command line.

The kernel guards against exactly this ordering mistake with one boolean, static_key_initialized — whose only job, per its own comment, is “to generate warnings if static_key manipulation functions are used before jump_label_init is called”; jump_label_init_ro() later even enforces it with a WARN_ON_ONCE(). That is why jump_label_init() is called very early from start_kernel() in init/main.c, strictly before parse_early_param() gets a chance to run any handler like xen_parse_mc_debug():

void __init jump_label_init(void)
{
        ...
        jump_label_sort_entries(iter_start, iter_stop);

        for (iter = iter_start; iter < iter_stop; iter++) {
                if (jump_label_type(iter) == JUMP_LABEL_NOP)
                        arch_jump_label_transform_static(iter, JUMP_LABEL_NOP);

                in_init = init_section_contains((void *)jump_entry_code(iter), 1);
                jump_entry_set_init(iter, in_init);

                iterk = jump_entry_key(iter);
                if (iterk == key)
                        continue;
                key = iterk;
                static_key_set_entries(key, iter);
        }
        static_key_initialized = true;
}

Before the loop even starts, jump_label_sort_entries() sorts the whole of __jump_table by key, then by code address — the precondition that lets the “walk all sites for this key as a linear scan” from §7.3 and the “batch must stay address-ordered” requirement from §10.2 both work later.

The loop that follows does three things per entry, in this order, trusting that the table is already sorted:

  1. Optionally run arch_jump_label_transform_static for NOP sites — a no-op on x86, since objtool and the compiler have already left the right bytes in place (§6.5); other architectures without that build-time trick do real work here.
  2. Mark __init sites, so nothing later tries to patch a call site living in memory that will be freed once init finishes.
  3. Point each key at the first entry of its contiguous run in the now-sorted table, so key->entries is ready to use the moment something calls static_branch_enable().

None of this touches instruction bytes for sites that are already correct: the compiled-in nop/jmp already matches the initial value of each key (§6.1). What this pass builds is bookkeeping — sort order, the __init flag, and the entries pointer — so that a later toggle knows exactly which sites to patch and in what order.

jump_label_init_ro() runs much later, from mark_readonly(). It walks __jump_table a second time, but this pass skips almost everything — it only acts on keys that is_kernel_ro_after_init() recognizes as living in __ro_after_init storage.2 For each matching key it calls:

static inline bool static_key_sealed(struct static_key *key)
{
        return (key->type & JUMP_TYPE_LINKED) && !(key->type & ~JUMP_TYPE_MASK);
}

static inline void static_key_seal(struct static_key *key)
{
        unsigned long type = key->type & JUMP_TYPE_TRUE;
        key->type = JUMP_TYPE_LINKED | type;
}

static_key_seal() keeps only the JUMP_TYPE_TRUE bit of the current key->type and throws the rest of the word away, replacing it with JUMP_TYPE_LINKED plus that one preserved bit. Recall from §7.1 that this word is normally a tagged pointer — the low 2 bits are a type tag, and everything above them is either the entries pointer the loop earlier in this section just installed, or a module-chain next pointer from §11. Sealing collapses it down to just the tag, with nothing left standing above bit 1:

before sealing (a live `entries` or `next` pointer, tagged):

    bit63                                        bit1  bit0
    [ real pointer value ..................... ] [ L ] [ T ]

after static_key_seal():

    bit63                                        bit1  bit0
    [ 0000000000000000000000000000000000000000 ] [ 1 ] [ T ]

T is the preserved JUMP_TYPE_TRUE bit; L is JUMP_TYPE_LINKED. static_key_sealed() is exactly the test for the “after” picture — JUMP_TYPE_LINKED set and nothing above bit 1 — so it can recognize a sealed key at a glance, regardless of which of the two pointer kinds that word used to hold.

Several call sites can share the same key, so this loop visits the same key more than once as it walks the table entry by entry. static_key_sealed() is checked before static_key_seal() runs, so the first visit seals the key and every later visit for that same key becomes a no-op.

This is safe only because __ro_after_init is a promise that the key is never toggled again after boot: once sealed, key->entries is gone, so nothing could walk it even if some later code mistakenly tried.

The payoff shows up in §11. When a module loaded afterward references a sealed key, jump_label_add_module() skips allocating and chaining a struct static_key_mod for it — the bookkeeping §11 otherwise needs so a future toggle can find and patch call sites living in other modules. Instead, it just patches the sites of that module once, immediately, to match the already-final value of the key.

The ordering against mark_rodata_ro() follows from the same fact: sealing is the last write anything makes to key->type, and that field sits inside the very __ro_after_init section mark_rodata_ro() is about to make actually read-only in the page tables. Running first just means the field has already settled into its final value before write access to it disappears.

9.2 Enabling: static_key_enable() / static_branch_enable()

jump_label_init() (§9.1) only builds bookkeeping — it never flips a key. Every site in the kernel image boots running whatever nop/jmp the compiler and objtool baked in (§6.1, §6), on or off, and stays that way until something calls static_branch_enable() or static_key_slow_inc() for the first time. This section is about that first flip, which can happen years into uptime, on a system where other CPUs may already be executing the very instructions about to be rewritten.

drivers/md/dm-stats.c shows a concrete case of this (§4.4): the first time a user asks the device-mapper stats ioctl to start recording per-region I/O counters, it flips stats_enabled, a key that has sat disabled since boot:

if (!static_key_enabled(&stats_enabled.key))
        static_branch_enable(&stats_enabled);

The guard matters as much as the call: it is what keeps a second, third, or hundredth request for the same device from re-triggering a full patch round once the key is already on (§4.4). But the first time through, this static_branch_enable(&stats_enabled) does run, and from that call onward every stats_enabled check in the I/O path — on every CPU, some of which may be running that exact code right now — has to observe the new state, and none of them may ever see a half-patched instruction. static_key_enable_cpuslocked() below is what makes that safe:

void static_key_enable_cpuslocked(struct static_key *key)
{
        ...
        jump_label_lock();
        if (atomic_read(&key->enabled) == 0) {
                atomic_set(&key->enabled, -1);      /* "enabling" */
                jump_label_update(key);             /* patch all sites */
                atomic_set_release(&key->enabled, 1);
        }
        jump_label_unlock();
}

Walking through what that function does, in order:

  1. It first checks whether the key is already enabled, and bails out if so.
  2. jump_label_mutex serializes all jump-label patching globally, so two callers enabling different keys at the same time still cannot have their text pokes race one another.
  3. It sets enabled = -1 before touching any code — this is the 0 → -1 transient from the state diagram above, and it is what makes jump_label_update() below safe to run concurrently with readers: static_key_count() and static_key_enabled() both treat -1 as enabled, so no concurrent reader ever sees a window of “disabled” while sites are only half-patched.
  4. jump_label_update() does the actual work: it patches every site for this key (§9.5).
  5. Finally, it stores 1.

static_key_enable() wraps this in cpus_read_lock() so CPUs cannot come online mid-patch.

The refcounted path, static_key_slow_inc_cpuslocked(), is not just this same function reused for inc/dec (§4.3). With multiple independent owners, more than one CPU can call it at once, and only one of them may actually be the one that flips 0 → 1:

bool static_key_slow_inc_cpuslocked(struct static_key *key)
{
        lockdep_assert_cpus_held();

        if (static_key_fast_inc_not_disabled(key))
                return true;

        guard(mutex)(&jump_label_mutex);
        if (!atomic_cmpxchg(&key->enabled, 0, -1)) {
                jump_label_update(key);
                atomic_set_release(&key->enabled, 1);
        } else {
                if (WARN_ON_ONCE(!static_key_fast_inc_not_disabled(key)))
                        return false;
        }
        return true;
}

It tries a lock-free fast path first (below), and only takes jump_label_mutex if that fails. Past that point it is the same 0 → 1 dance as static_key_enable_cpuslocked() above: exactly one caller actually performs the transition and runs jump_label_update(); anyone else who reaches the mutex finds the key already on and just falls back to the fast path to add their own count.

static_key_fast_inc_not_disabled() is what makes “further incs” in the table from §4.3 cost nothing more than an atomic — no lock, no text poke:

bool static_key_fast_inc_not_disabled(struct static_key *key)
{
        int v;

        STATIC_KEY_CHECK_USE(key);
        /*
         * Negative key->enabled has a special meaning: it sends
         * static_key_slow_inc/dec() down the slow path, and it is non-zero
         * so it counts as "enabled" in jump_label_update().
         *
         * The INT_MAX overflow condition is either used by the networking
         * code to reset or detected in the slow path of
         * static_key_slow_inc_cpuslocked().
         */
        v = atomic_read(&key->enabled);
        do {
                if (v <= 0 || v == INT_MAX)
                        return false;
        } while (!likely(atomic_try_cmpxchg(&key->enabled, &v, v + 1)));

        return true;
}

The retry loop only succeeds once it can prove, atomically, that the count is already an ordinary positive number. v <= 0 catches both a genuinely disabled key (0) and an enable already in progress (-1), sending both cases down to the slow path above instead of incrementing a value that doesn’t mean what it looks like yet. v == INT_MAX guards against overflows.

9.3 Disabling

§9.2 walked through what happens when a key crosses from off to on. Disabling is the same problem in reverse: patch every site back to its original instruction, without letting any reader ever observe a torn one. It runs on the very same rule as enabling — not every call to disable/dec needs to touch an instruction at all. jump_label_update() only has to run on the one transition that actually flips what is patched into the instruction stream: 1 → 0 for the boolean API, N → 0 for the refcounted one. Every other call just moves a plain integer and can be answered with a single atomic instruction — no lock, no text poke, because nothing about the patched code needs to change.

That is exactly the same shape as the static_key_fast_inc_not_disabled() from §9.2 on the enable side — a bare CAS loop, no lock, no patch, for every increment that doesn’t cross 0 → 1. That rule (every transition that isn’t on a patch-triggering boundary is a bare atomic op) is what both functions below are built around, on either side of the key.

void static_key_disable_cpuslocked(struct static_key *key)
{
        ...
        if (atomic_read(&key->enabled) != 1) {
                WARN_ON_ONCE(atomic_read(&key->enabled) != 0);
                return;
        }

        jump_label_lock();
        if (atomic_cmpxchg(&key->enabled, 1, 0) == 1)
                jump_label_update(key);
        jump_label_unlock();
}

static_key_disable_cpuslocked() is the mirror, in the boolean API, of the enable path from §9.2, but it skips the -1 choreography. It first checks that the key is currently exactly 1, bailing out (and warning if the value isn’t 0 either, which would mean the boolean and refcounted APIs got mixed on this key — §4.3) before doing anything else. The actual disable is then a single atomic_cmpxchg(&key->enabled, 1, 0): “if the value is currently 1, replace it with 0, and tell me whether you succeeded.” If some other CPU changed it first, the compare fails and this call does nothing further. Only on success does it call jump_label_update() to patch every site for this key back to its disabled instruction.

Compare that to enabling: there, enabled is deliberately set to -1 before any text is touched, specifically so no reader can mistake in-progress patching for “off” (point 3 of §9.2). Disabling has no matching problem to solve. During the window between the cmpxchg above and jump_label_update() finishing, enabled already reads 0 while some sites out there are still physically holding their “enabled” instruction — the opposite kind of staleness from enabling (stale “on” instead of stale “off”), but just as harmless. A reader who calls static_key_enabled() during that window is simply told “off” a few instructions before the code itself has caught up; §4.4 already covers why these state reads only ever need to be eventually correct, never instruction-exact.

The refcounted side runs through a different function, __static_key_slow_dec_cpuslocked(), which only reaches jump_label_update() on the one decrement that actually drives the count to 0 (atomic_dec_and_test() reports true exactly then, and only then). Every decrement that lands above 1 — 3 -> 2, 2 -> 1, and so on — is intercepted earlier, by static_key_dec_not_one(), which performs a plain atomic decrement and returns without ever taking the jump-label lock or looking at an instruction stream. That is the same “every transition that isn’t on a patch-triggering boundary is a bare atomic op” rule we discussed earlier, now seen from the decrement side:

static bool static_key_dec_not_one(struct static_key *key)
{
        int v;

        v = atomic_read(&key->enabled);
        do {
                WARN_ON_ONCE(v < 0);

                if (WARN_ON_ONCE(v == 0))
                        return true;
                if (v <= 1)
                        return false;
        } while (!likely(atomic_try_cmpxchg(&key->enabled, &v, v - 1)));

        return true;
}

The CAS loop is what makes this safe against concurrent decrements: if another CPU wins the race and changes key->enabled between the read of this CPU and its cmpxchg, v is refreshed and the loop just re-checks the same two conditions against the new value rather than clobbering it.

9.4 Deferred / rate-limited dec

This is the internal counterpart to the rate-limited disable in §4.7, seen here from the function-by-function angle of this section rather than the call-site angle of the cookbook. §4.7 already covers why a delayed decrement exists and how the coalescing works when toggles arrive faster than the timeout; this is only the piece that was left out there — where the state-machine logic actually lives.

__static_key_slow_dec_deferred() opens with the exact same static_key_dec_not_one() check §9.3 just introduced: if this decrement would not bring the count down to the 1 -> 0 boundary, it is a plain atomic decrement and the function returns immediately, no different from the non-deferred path.

The two paths only diverge on the one decrement that would disable the key. Where the __static_key_slow_dec_cpuslocked() from §9.3 reaches straight for jump_label_update() at that point, this function instead calls schedule_delayed_work() — a standard kernel workqueue primitive that runs a callback once a given delay has elapsed, rather than right away — and returns without touching a single instruction. The count is deliberately left at 1 rather than dropped to 0; only the plan to disable has been recorded, in the timer. In full:

void __static_key_slow_dec_deferred(struct static_key *key,
                    struct delayed_work *work,
                    unsigned long timeout)
{
        if (static_key_dec_not_one(key))
                return;

        schedule_delayed_work(work, timeout);
}

When that timer eventually fires, jump_label_update_timeout() runs the ordinary, undeferred decrement path from §9.3. If nothing else touched the key in the meantime, the count is still exactly 1, the decrement finally lands on the 1 -> 0 boundary, and jump_label_update() runs for real. If instead another caller incremented the key again while the timer was pending, the count is no longer 1 by the time the timer fires — so static_key_dec_not_one() intercepts that decrement too, as an ordinary atomic op, and jump_label_update() is never reached. Nothing needed patching back, because nothing was ever patched in the first place.

9.5 jump_label_update() → __jump_label_update()

Every path in §9.2 through §9.4 eventually funnels into this function — it is the one that actually walks the sites of a key and asks the architecture layer to patch each one. Reading the outer function first:

static void jump_label_update(struct static_key *key)
{
        ...
        if (static_key_linked(key)) {
                __jump_label_mod_update(key);   /* walk module list */
                return;
        }
        entry = static_key_entries(key);
        if (entry)
                __jump_label_update(key, entry, stop, init);
}

The first branch is the module case from §7.1: if bit 1 of key->type is set (LINKED), the sites of this key are not one contiguous run inside __jump_table, but scattered across a linked list of per-module entry tables (struct static_key_mod), so a separate helper has to walk that list instead of a flat array — §11 covers __jump_label_mod_update() in full. Otherwise, static_key_entries(key) recovers the pointer §9.1 stored during boot — the first entry of the contiguous run for this key inside the sorted, vmlinux-only table — and the real patching happens in __jump_label_update().

On x86, which defines HAVE_JUMP_LABEL_BATCH, that function looks like this:

for (; entry < stop && jump_entry_key(entry) == key; entry++) {
        if (!jump_label_can_update(entry, init))
                continue;
        if (!arch_jump_label_transform_queue(entry, jump_label_type(entry))) {
                arch_jump_label_transform_apply();
                BUG_ON(!arch_jump_label_transform_queue(...));
        }
}
arch_jump_label_transform_apply();

The loop condition — advance while entry < stop and jump_entry_key(entry) == key — only works because of the sort from §9.1: every site belonging to this key is guaranteed to sit in one unbroken run starting at entry, so the loop can walk forward blindly and stop the instant it reaches a site belonging to some other key, with no need to search the rest of the table.

For each site still in range, jump_label_type(entry) computes enabled ^ branch — the exact static/dynamic formula from the table in §6.1, but now evaluated against the current live state of the key rather than its compile-time initial value — to decide whether this specific site should end up holding a JUMP_LABEL_NOP or a JUMP_LABEL_JMP.

Before acting on that answer, jump_label_can_update() filters out two kinds of site that must not be touched at all: one still living in __init text after boot has finished (that memory may already have been freed, which is exactly what the jump_entry_is_init flag from §9.1 was recorded to detect), and one that kernel_text_address() does not even recognize as live kernel text — built-in code that is __exit-only and therefore can never run, so patching it would be pointless even though it is technically still present:

static bool jump_label_can_update(struct jump_entry *entry, bool init)
{
        if (!init && jump_entry_is_init(entry))
                return false;

        if (!kernel_text_address(jump_entry_code(entry))) {
                WARN_ONCE(!jump_entry_is_init(entry),
                          "can't patch jump_label at %pS",
                          (void *)jump_entry_code(entry));
                return false;
        }

        return true;
}

Rather than patching each surviving site immediately, one at a time, the loop splits the work into two separate jobs: arch_jump_label_transform_queue() builds up a batch, and arch_jump_label_transform_apply() executes it. Splitting them apart is what lets every site belonging to one key ride through a single INT3-synchronized patch round (§10.2) instead of paying for one round per site.

Queuing a site. Each call to arch_jump_label_transform_queue() computes the replacement bytes for that one site and hands (address, new bytes, length) to smp_text_poke_batch_add(). That function appends the request to a pending array; nothing is written to memory yet. The one exception is early boot: only one CPU is running, so there is no concurrent fetcher to synchronize against and nothing worth batching — the function calls the non-batching arch_jump_label_transform() directly instead.

Applying the batch. arch_jump_label_transform_apply() executes everything queued so far. It calls smp_text_poke_batch_finish() which runs the three-step INT3 dance once for the whole batch instead of once per site. __jump_label_update() calls it once, after its loop ends, to flush whatever is still pending.


10 x86 text patching: the gory details

This is where the torn-write argument from §5.3 turns into working code.

10.1 Early boot vs live SMP

Very early in boot, only the boot CPU is running and .text is still writable. __jump_label_transform(), the function every x86 patch eventually funnels through, checks for exactly that window and takes the cheap route whenever it still holds:

/* arch/x86/kernel/jump_label.c: __jump_label_transform() */
if (init || system_state == SYSTEM_BOOTING) {
        text_poke_early(...); /* IRQ-disabled + `memcpy()` + sync_core() */
        return;
}
smp_text_poke_single(...);      /* or batch_add during queueing (§9.2) */

system_state == SYSTEM_BOOTING is a proxy for one fact: only the boot CPU exists so far. That alone is reason enough to skip the whole IPI-synchronized protocol of §10.3 — with nobody else around to race the write, text_poke_early() can just do a plain IRQ-disabled + memcpy() + sync_core() and be done. .text also happens to still be writable at this point, so the alias trick from §5.4 isn’t needed either — but that is a bonus the check gets for free, not something it verifies directly: smp_init() wakes every other CPU well before mark_rodata_ro() ever runs, so there is a real stretch of boot where other CPUs are already up while .text is still writable. Jump labels take the full protocol for that entire stretch anyway, because the only thing this check ever verifies is whether this CPU is still provably alone.

The leading init in that condition is not the per-site __init-text flag from §9.1 (jump_entry_is_init()). It is a separate, whole-system flag threaded down from init = system_state < SYSTEM_RUNNING inside jump_label_update() itself (§9.5), true a little longer than SYSTEM_BOOTING alone. On the batching path of x86, though, that value never actually reaches here: arch_jump_label_transform_queue() only calls this function through its own system_state == SYSTEM_BOOTING fallback, passing a hardcoded 0 for init every time it does. So on x86 this condition is, in practice, exactly system_state == SYSTEM_BOOTING — the init || half of it only ever matters on architectures that call this function directly, without going through batching at all.

10.2 Batching API used by jump labels

§10.1 settled how a single site gets patched once the decision to patch it is made; this section is about when jump labels actually pull that trigger. A busy tracepoint can have thousands of call sites sharing one key, and paying the full IPI-synchronized protocol (§10.3) separately for each one would be needless — the sites can be collected first and the expensive part paid once for the whole group. Two arch-level hooks make that possible: one that queues the newly computed bytes for a site without touching hardware yet, and one that flushes everything queued so far in a single synchronized pass:

bool arch_jump_label_transform_queue(...)
{
        if (system_state == SYSTEM_BOOTING) {
                arch_jump_label_transform(entry, type);
                return true;
        }
        mutex_lock(&text_mutex);
        jlp = __jump_label_patch(entry, type);
        smp_text_poke_batch_add(addr, jlp.code, jlp.size, NULL);
        mutex_unlock(&text_mutex);
        return true;
}

void arch_jump_label_transform_apply(void)
{
        mutex_lock(&text_mutex);
        smp_text_poke_batch_finish();
        mutex_unlock(&text_mutex);
}

The queue collects many sites (a busy tracepoint may have thousands); one batch_finish() amortizes the IPI syncs. The queue must stay address-sorted; if a new address would break order, or the page-sized array is full, smp_text_poke_batch_add() flushes early (text_poke_addr_ordered() in alternative.c). That is why jump_label_cmp sorts by code address within each key.

How many fit in one batch? The pending patches live in a single statically-allocated page, struct smp_text_poke_loc:

struct smp_text_poke_loc {
        s32 rel_addr;   /* addr := _stext + rel_addr           */
        s32 disp;       /* branch displacement, for emulation  */
        u8  len;        /* 1, 2, 5, or 6 bytes                 */
        u8  opcode;     /* first opcode byte, for emulation    */
        u8  text[5];    /* the new instruction bytes           */
        u8  old;        /* byte that used to be there, for perf/PT tracing */
};                       /* 16 bytes, naturally aligned         */

#define TEXT_POKE_ARRAY_MAX (PAGE_SIZE / sizeof(struct smp_text_poke_loc))
/* 4096 / 16 = 256 entries per flush on a 4K-page x86_64 build */

rel_addr is relative to _stext3, not to the entry itself like jump_entry — cheap, because every patch site is, by definition, in kernel text. A tracepoint or jump-label key with more than 256 call sites needs more than one batch_finish() round (i.e. more than 3 IPI rounds, §10.3) to fully enable/disable.

__jump_label_update() (§9.5) does contain a generic “queue full → apply → retry” branch, and it looks like the natural place to expect this 256-site overflow to be handled.

On x86 the “queue full → apply → retry” branch never runs, though: it only fires when arch_jump_label_transform_queue() itself returns false, and the implementation of that function on this architecture never returns false. So the branch is dead code here.

for (; (entry < stop) && (jump_entry_key(entry) == key); entry++) {

        if (!jump_label_can_update(entry, init))
                continue;

        if (!arch_jump_label_transform_queue(entry, jump_label_type(entry))) {
                /*
                 * Queue is full: Apply the current queue and try again.
                 */
                arch_jump_label_transform_apply();
                BUG_ON(!arch_jump_label_transform_queue(entry, jump_label_type(entry)));
        }
}
arch_jump_label_transform_apply();

The overflow is actually caught one layer further down, entirely inside smp_text_poke_batch_add() — the entire guard is these three lines:

void smp_text_poke_batch_add(void *addr, const void *opcode, size_t len, const void *emulate)
{
        if (text_poke_array.nr_entries == TEXT_POKE_ARRAY_MAX || !text_poke_addr_ordered(addr))
                smp_text_poke_batch_finish();
        __smp_text_poke_batch_add(addr, opcode, len, emulate);
}

Every iteration of that for loop does one thing: it calls arch_jump_label_transform_queue(), which computes the new bytes for one site and appends one smp_text_poke_loc to the array — the cheap, per-site step this section has been describing. The if (!arch_jump_label_transform_queue(...)) branch is the “queue full → apply → retry” path discussed above; on x86 it never runs, since that function never returns false. Nothing expensive happens inside the loop.

arch_jump_label_transform_apply() sits outside the loop, called exactly once after it exits, for every site the loop just queued. That single call is what finally triggers batch_finish(), the three-phase IPI-synchronized protocol from §10.3.

So the entire set of call sites for a key — whether it has one or close to 256 — rides through on that one shared batch_finish(), and only a key with more than 256 sites forces a second round.

10.3 The INT3 SMP algorithm (smp_text_poke_batch_finish)

This is the payoff of everything §5.3 through §10.2 built toward. The key fact from §5.3 was that only a single-byte store is atomic with respect to instruction fetch — nothing wider is. The protocol below never trusts a multi-byte write to be safe on its own; instead it uses one atomic single-byte store to plant a trap on top of the site, uses that trap to absorb any CPU unlucky enough to fetch through mid-update, and only then fills in the rest. Three writes, three synchronizations, one site at a time across the whole batch. Documented at the top of the function in alternative.c:

For each site in the vector:
  (1) Write INT3 (0xCC) over the first byte
      → IPI sync all CPUs   (serialize pipelines / I-caches)

  (2) Write bytes 1..N-1 of the new instruction
      → IPI sync again      unnecessary, according to Intel,
                            but better safe than sorry 

  (3) Write byte 0 of the new instruction (replaces INT3)
      → IPI sync again

Here are the actual bytes of one site across the three phases, patching a 5-byte NOP (0f 1f 44 00 00) into a 5-byte JMP rel32 (e9 + 4-byte displacement, shown as ?? ?? ?? ??):

 start (before)      0f 1f 44 00 00      any fetch: executes the NOP

 phase 1 (INT3 in)  cc 1f 44 00 00      any fetch: #BP -> handler emulates
                    ^^                  the NEW instruction (jumps to l_yes)
                    trap byte

                    -------- IPI sync --------

 phase 2 (tail in)  cc ?? ?? ?? ??      same as phase 1: byte 0 is still
                    ^^                  INT3, so any fetch still traps and
                    still traps         gets emulated — the real tail bytes
                                        underneath are now correct, but
                                        nothing reads them yet

                    -------- IPI sync --------

 phase 3 (done)     e9 ?? ?? ?? ??      any fetch: executes the real JMP
                    ^^ real opcode      directly, no trap needed anymore

                    -------- IPI sync --------

The subtle point this diagram is here to make: the observable behavior of the site flips the instant the sync in phase 1 completes, not at phase 3. From phase 1 onward, any CPU landing on this address — whether by falling into it in a hot loop or by literally executing byte 0 — gets the effect of the new instruction, because the #BP handler always emulates the pending new instruction (it was computed and stashed in the queue back at smp_text_poke_batch_add() time, long before phase 1 starts). A CPU that hits the address mid-update and one that hits it after phase 3 land on the same outcome — the only difference is whether it got there by trapping into the handler or by executing the finished bytes directly. Phases 2 and 3 exist to make that direct path available, so steady-state execution stops paying the #BP tax.

The writing side. All of the above is driven by smp_text_poke_batch_finish() (it early-returns immediately if text_poke_array.nr_entries is 0 — nothing queued, nothing to do). Trimmed of the cond_resched() softlockup guard, the perf/Intel-PT tracing hook, and a 6-byte-opcode edge case, it opens by arming the refcount of every CPU — the release side of the release/acquire pairing the “Why INT3?” sidebar below explains — then runs the three phases in order:

for_each_possible_cpu(i)
        atomic_set_release(per_cpu_ptr(&text_poke_array_refs, i), 1);
smp_wmb();

Phase 1 writes INT3 over the first byte of every site, saving the byte it replaces (for the perf/PT tracing hook trimmed out above), then syncs once for the whole batch:

for (i = 0; i < text_poke_array.nr_entries; i++) {
        text_poke_array.vec[i].old = *(u8 *)text_poke_addr(&text_poke_array.vec[i]);
        text_poke(text_poke_addr(&text_poke_array.vec[i]), &int3, INT3_INSN_SIZE);
}
smp_text_poke_sync_each_cpu();

Phase 2 writes bytes 1..len-1 of the new instruction for every site — safe, since byte 0 is still INT3 — and syncs again, but only if some site actually has tail bytes to write (a single-byte patch has none):

for (do_sync = 0, i = 0; i < text_poke_array.nr_entries; i++) {
        int len = text_poke_array.vec[i].len;

        if (len - INT3_INSN_SIZE > 0) {
                text_poke(text_poke_addr(&text_poke_array.vec[i]) + INT3_INSN_SIZE,
                          text_poke_array.vec[i].text + INT3_INSN_SIZE,
                          len - INT3_INSN_SIZE);
                do_sync++;
        }
}
if (do_sync)
        smp_text_poke_sync_each_cpu();

Phase 3 writes byte 0, replacing the INT3 — skipped for the corner case (discussed next) where the new opcode is itself 0xCC — and syncs again if anything actually changed:

for (do_sync = 0, i = 0; i < text_poke_array.nr_entries; i++) {
        u8 byte = text_poke_array.vec[i].text[0];

        if (byte == INT3_INSN_OPCODE)
                continue;
        text_poke(text_poke_addr(&text_poke_array.vec[i]), &byte, INT3_INSN_SIZE);
        do_sync++;
}
if (do_sync)
        smp_text_poke_sync_each_cpu();

Draining after phase 3. After the last sync, the writer cannot yet assume every CPU has left the handler — a CPU could be inside smp_text_poke_int3_handler() right up until that sync completes (it entered before the sync, is still emulating). So smp_text_poke_batch_finish() ends with this drain and the final reset:

for_each_possible_cpu(i) {
        atomic_t *refs = per_cpu_ptr(&text_poke_array_refs, i);

        if (unlikely(!atomic_dec_and_test(refs)))
                atomic_cond_read_acquire(refs, !VAL);
}
text_poke_array.nr_entries = 0;

atomic_dec_and_test() decrements the refcount of every CPU; for any that don’t immediately hit zero, atomic_cond_read_acquire(refs, !VAL) spin-waits — i.e. for stragglers still mid-handler to finish and drop their own reference. Only then does the function reset text_poke_array.nr_entries = 0, making the buffer safe to reuse for the next batch.

In the common case — jump labels, static calls, ftrace, none of which ever replace a site with a literal 0xCC — phase 3 already wrote every final byte, and smp_text_poke_sync_each_cpu() already fenced out any in-flight handler. So this drain loop observes zero immediately, and the comment in the source calls it out explicitly: “unless the replacement instruction is INT3, this case goes unused.”

It exists for the corner case (other clients, not jump labels) where the new opcode of a site is 0xCC itself. Byte 0 is therefore left alone in phase 3 (writing 0xCC over an existing 0xCC would be a no-op anyway, so that per-entry sync is skipped), and the only thing standing between “batch done” and “safe to reuse the array” is the atomic_cond_read_acquire() spin-wait itself.

Why INT3? The whole point of writing INT3 first (phase 1 of the three-phase protocol above) is that it gives the kernel a way to intercept any CPU that would otherwise have executed half-written bytes, and make it run the finished instruction instead. While the site is mid-update, any CPU that hits it takes #BP (a fault, vector 3), and that fault always lands in smp_text_poke_int3_handler(), which is wired up as the #BP handler in traps.c. Its locals are just tpl (the matched struct smp_text_poke_loc *), ret, and ip; here is what it actually does, one check at a time.

1. Bail on user mode. This handler only ever concerns itself with traps hit while executing kernel text:

if (user_mode(regs))
        return 0;

2. Confirm a batch is actually in flight, via a per-CPU refcount, text_poke_array_refs:

smp_rmb();

if (!try_get_text_poke_array())
        return 0;

try_get_text_poke_array() is just an atomic increment-if-nonzero:

static __always_inline bool try_get_text_poke_array(void)
{
        atomic_t *refs = this_cpu_ptr(&text_poke_array_refs);
        return raw_atomic_inc_not_zero(refs);   /* 0 → fails: no batch active */
}

Before touching any site, smp_text_poke_batch_finish() arms the refcount of every CPU to 1 with atomic_set_release(), then smp_wmb()s, and only then writes the INT3 bytes. The smp_rmb() above is the mirror-image acquire. That release/acquire pairing is what guarantees: if the #BP of a CPU fires (meaning it must have fetched the INT3 the writer stored), that same CPU is also guaranteed to see a fully-populated, non-zero-refcount text_poke_array — never a half-written vector.

3. Find which site trapped, binary-searching text_poke_array.vec for this regs->ip - 1 (skipping straight to a direct compare when there is exactly one entry):

ip = (void *) regs->ip - INT3_INSN_SIZE;

if (unlikely(text_poke_array.nr_entries > 1)) {
        tpl = __inline_bsearch(ip, text_poke_array.vec, text_poke_array.nr_entries,
                              sizeof(struct smp_text_poke_loc),
                              patch_cmp);
        if (!tpl)
                goto out_put;
} else {
        tpl = text_poke_array.vec;
        if (text_poke_addr(tpl) != ip)
                goto out_put;
}

ip += tpl->len;

4. Emulate the new instruction by editing regs and returning — the return address of the #BP handler becomes the effect of the emulated instruction, so iret resumes execution as if the new bytes had actually executed:

switch (tpl->opcode) {
case INT3_INSN_OPCODE:
        goto out_put;             /* explicit INT3, not ours; do not consume */

case RET_INSN_OPCODE:
        int3_emulate_ret(regs);
        break;

case CALL_INSN_OPCODE:
        int3_emulate_call(regs, (long)ip + tpl->disp);
        break;

case JMP32_INSN_OPCODE:
case JMP8_INSN_OPCODE:
        int3_emulate_jmp(regs, (long)ip + tpl->disp);
        break;

case 0x70 ... 0x7f: /* Jcc */
        int3_emulate_jcc(regs, tpl->opcode & 0xf, (long)ip, tpl->disp);
        break;

default:
        BUG();
}

ret = 1;
  • JMP rel8/rel32 (also a pending NOP, encoded as JMP with disp == 0, i.e. “jump to the next instruction”): int3_emulate_jmp() just overwrites regs->ip.
  • CALL/RET/Jcc: same idea (push a fake return address / pop one / conditionally add the displacement) — jump labels never generate these, but static calls and ftrace share this exact engine and do.
  • If the replacement opcode is itself 0xCC (some other text-poke client is intentionally installing a breakpoint, not passing through this emulator): the handler does not consume the trap — it falls through so that debugging infrastructure (kgdb, kprobes) gets its #BP.

Jump labels only ever exercise the JMP32/JMP8 case — the RET/CALL/Jcc arms exist because static calls and ftrace share this exact handler.

5. Release the refcount and report the trap as handled:

out_put:
        put_text_poke_array();
        return ret;

So no CPU ever executes a torn multi-byte instruction: control either sees the old bytes (before the sync in phase 1 completes everywhere), takes the INT3+emulation path (the entire window from phase 1 to phase 3), or sees the finished new bytes (after the sync in phase 3).

10.4 What “sync” means

The three-phase protocol from §10.3 names “IPI sync” as a step three times over, without ever saying what happens when that IPI lands. This section answers that — not what the patching CPU does (§10.3 already covered that), but what every other CPU is forced to do in response.

smp_text_poke_sync_each_cpu() is the function invoked at each of those three points, and its job is deceptively narrow: send an IPI to every CPU other than the one doing the patching, have each of them run sync_core(), and block until every last one has reported back. That blocking is not incidental — the whole structure of §10.3 depends on each phase being globally visible before the next one begins. If the patching CPU raced ahead and wrote the tail bytes of phase 2 while some other core was still mid-fetch on the phase-1 view of the site, the entire point of planting the INT3 trap first would be defeated.

What actually happens on the receiving end is where the CPU model from §5.1 finally pays for itself. A modern core does not execute an instruction the moment it sees its bytes: it fetches ahead of where it is currently retiring, decodes into microcodes, and may be holding several instructions of that pipeline in flight at once. A CPU that fetched the old bytes moments before the patch landed can still be sitting on a stale decode of them — and no ordinary memory write, however carefully sequenced, undoes that. Something has to reach into the pipeline itself and discard the stale work. That is exactly what a serializing instruction is architecturally defined to do: retire everything already in flight, drop any speculative or partially-decoded work, and guarantee that the very next fetch goes out fresh.

Two ways to get that guarantee are available, and which one runs depends on the CPU generation. X86_FEATURE_SERIALIZE, present on most CPUs manufactured since around 2020, provides a dedicated instruction whose only job is this flush — cheap and direct:

static __always_inline void serialize(void)
{
        /* Instruction opcode for SERIALIZE; supported in binutils >= 2.35. */
        asm volatile(".byte 0xf, 0x1, 0xe8" ::: "memory");
}

That comment is the real reason this is written as raw .byte values rather than a mnemonic the assembler recognizes: SERIALIZE was only added to binutils in 2.35, so a kernel built with an older assembler still needs to be able to emit the three raw opcode bytes of the instruction (0F 01 E8) by hand. The "memory" clobber tells the compiler this call is a full optimization barrier — it must not reorder ordinary memory accesses across it.

On older hardware, the kernel falls back to iret_to_self(). An interrupt frame is just the handful of words — return SS, RSP, RFLAGS, CS, and RIP — that the CPU itself pushes onto the stack whenever a real interrupt or exception fires, and which iret later pops to hand control back to whatever was running. Normally software never builds one by hand; the CPU builds it automatically at the moment of a real trap, the same way it did for the #BP frame the INT3 handler above edited before its own iret. iret_to_self() does the half of that job the CPU would do: it pushes those same five words by hand, with no real interrupt behind them, sets the saved RIP field to the very next instruction after the iret, and then executes iret against that fabricated frame. As far as the CPU can tell, this is a completely ordinary return from an interrupt, so it does everything architecture requires of one — including the serializing flush this function exists to get. But because the fabricated return address is simply “keep going from here,” nothing about the actual control flow changes: execution lands right back where it would have anyway, one instruction later.

Here is the actual body, from arch/x86/include/asm/sync_core.h (the 64-bit variant — a 32-bit build takes a shorter path that skips the SS/RSP pushes, since a same-privilege 32-bit iret doesn’t pop them):

static __always_inline void iret_to_self(void)
{
        unsigned int tmp;

        asm volatile (
                "mov %%ss, %0\n\t"      /* SS is a segment register; it can't be   */
                                         /* pushed directly in this form, so copy   */
                                         /* it into a GPR first                     */
                "pushq %q0\n\t"          /* frame field 1 (bottom): return SS       */
                "pushq %%rsp\n\t"        /* frame field 2: return RSP — but this    */
                                         /* captures RSP *after* the SS push above  */
                                         /* already moved it down by 8              */
                "addq $8, (%%rsp)\n\t"   /* ...so correct the just-pushed copy back */
                                         /* up by 8, to the RSP value from before   */
                                         /* this function started pushing anything  */
                "pushfq\n\t"             /* frame field 3: return RFLAGS            */
                "mov %%cs, %0\n\t"       /* same GPR trick as SS, for CS this time  */
                "pushq %q0\n\t"          /* frame field 4: return CS                */
                "pushq $1f\n\t"          /* frame field 5 (top): return RIP — the   */
                                         /* address of local label "1:" below, i.e. */
                                         /* the instruction right after this one    */
                "iretq\n\t"              /* pop all five fields and "return" — the  */
                                         /* CPU treats this exactly like returning  */
                                         /* from a genuine interrupt                */
                "1:"                     /* execution resumes here, indistinguishable */
                                         /* from simply falling through to this point */
                : "=&r" (tmp), ASM_CALL_CONSTRAINT : : "cc", "memory");
}

The iret is architecturally required to be a serializing event on every CPU, which is precisely the guarantee this fallback needs — it works identically at any privilege level (so it survives under paravirtualization) and never exits to a hypervisor, both properties this code cannot give up. The price is that it measures a bit more than twice as slow as the dedicated instruction, and it unconditionally unmasks NMIs, which SERIALIZE does not.

One more candidate is conspicuously missing. CPUID also serializes, and on paper looks like the most portable option of all. The kernel avoids it here for a practical reason, not a correctness one: under virtualization, CPUID commonly traps out to the hypervisor, and this is exactly the kind of hot, latency-sensitive path — run on every online CPU, on every single key toggle — that cannot tolerate an unpredictable VM exit in the middle of it.

The choice is a single feature check, decided once per call:

static __always_inline void sync_core(void)
{
        if (static_cpu_has(X86_FEATURE_SERIALIZE)) {
                serialize();
                return;
        }
        iret_to_self();
}

10.5 Writing through RO mappings

Every store §10.3 walked through — the INT3, the tail bytes, the final first byte — is not a raw write to .text. Each one is a full call to text_poke(), meaning each one pays the entire §5.4 dance in full: build the temporary alias in text_poke_mm, switch %cr3 onto it, copy through STAC/CLAC, switch back, tear the alias down. Nothing about these writes being unusually small — as small as the single INT3 byte in phase 1 — or unusually frequent lets any of them skip a step.

All of that happens under text_mutex, and this is the piece §5.4 leaned on without yet saying where it comes from: text_mutex is what stops two unrelated patchers — jump labels, static calls, ftrace, kprobes, the alternatives machinery — from ever building two competing temporary mappings onto the same text_poke_mm address space at once. Jump labels layer a second lock, jump_label_mutex, on top of that, but the two are not protecting the same thing: text_mutex serializes individual pokes at the hardware level, while jump_label_mutex serializes the higher-level operation of enabling or disabling one whole key — the enabled counter update and the walk over the jump_entry run for that key, not just the bytes it eventually writes.

The two are also held for deliberately different spans. On the queueing path (§10.2), text_mutex is acquired and released once per call to arch_jump_label_transform_queue() — bracketing only the __jump_label_patch() computation for that one site and its single smp_text_poke_batch_add() append. It is free again in between sites, so some unrelated text_poke() caller elsewhere in the kernel is free to interleave its own single-site poke while jump labels are still accumulating theirs for this key. Only once the whole batch is ready does arch_jump_label_transform_apply() take text_mutex back and hold it continuously across the entire three-phase smp_text_poke_batch_finish() from §10.3 — that phase genuinely cannot tolerate a second patcher walking in mid-batch, since the correctness argument of the INT3 protocol assumes the batch array it iterates is exactly the one it built. jump_label_mutex, by contrast, stays held across all of that from the first line of jump_label_update() onward — there is no benefit to releasing it early, since doing so would only let a second, unrelated static_branch_enable() call start interleaving its own bookkeeping with that of this one, not let this one finish any faster.

10.6 End-to-end timeline for one enable

§10.1 through §10.5 examined the machinery one piece at a time — which patch path early boot takes, how sites get batched, the three INT3 phases, what “synchronize” actually does on the wire, and how each of those writes reaches memory that is nominally read-only. Laid end to end, a single static_branch_enable() call looks like this:

static_branch_enable(&key)
  cpus_read_lock()
  jump_label_mutex
  enabled = -1
  jump_label_update(key)
    for each jump_entry of key:
      __jump_label_patch()           # compute nop↔jmp bytes, sanity memcmp
      smp_text_poke_batch_add()      # append to vector (may flush if full)
    smp_text_poke_batch_finish()
      arm text_poke_array_refs = 1 on every CPU (smp_wmb)
      text_poke INT3 × N
      IPI sync                          # step 1
      text_poke tails (bytes 1..N-1) × N
      IPI sync                          # step 2 ("paranoid")
      text_poke first bytes × N
      IPI sync                          # step 3
      drain: wait for text_poke_array_refs == 0 on every CPU
      text_poke_array.nr_entries = 0    # buffer free for next batch
  enabled = 1 (release)
  unlock…

After this, every previously-nop site for that key is a jmp (or vice versa), and the hot path behavior has flipped — without any flag load.

Notice what does, and does not, scale with the number of call sites. A key can have one jump_entry or a few hundred, but the number of IPI rounds is fixed at three,4 because smp_text_poke_batch_finish() runs its three phases once over the entire batched vector, not once per site (§10.3). The only thing that grows that fixed cost is the §10.2 overflow case: past 256 queued sites, a second full batch_finish() round is required, so a tracepoint with, say, 300 call sites pays six IPI rounds total, not 300 x 3.


11 Modules: the trickiest part

Static keys often live in vmlinux (or module A) while call sites live in module B — a tracepoint defined in the core kernel, say, with trace_*() call sites scattered across several drivers loaded as modules. The picture from §7.1 of key->entries as one pointer into one contiguous run of jump_entry records only holds while every call site for a key lives in a single object.

Modules break that assumption in the least convenient way possible: they load and unload independently of vmlinux and of each other, in an order nothing can predict ahead of time, so the representation has to be able to grow and shrink at runtime instead of being settled once at boot the way §9.1 settles it for the vmlinux-only case. That is what makes this section “the trickiest part” — every step has to stay correct no matter how many objects currently contribute to a key, or in what order they arrived. Over its lifetime, that one word/pointer union from §7.1 can end up in exactly three states:

/*
               next                  next                  next
 key (LINKED) -----> static_key_mod -----> static_key_mod -----> NULL
                     (mod A / vmlinux)     (mod B)
                         |                    |
                      entries              entries
                         |                    |
                         v                    v
                    [jump_entry…]        [jump_entry…]
*/
struct static_key_mod {
        struct static_key_mod *next;
        struct jump_entry *entries;
        struct module *mod;
};

Concretely: say the static_key of a tracepoint is declared in vmlinux, but nothing in the core kernel itself calls it — only the driver module A does, loaded first. Module A becomes the sole contributor for the key, so key->entries points straight at the own run of module A, no list involved (the one-home-only case above, even though the struct of the key lives in vmlinux while its only call site lives in a module — those are two independent facts, and only the second one matters here). Module B now loads and also calls the same tracepoint. A second object just started contributing, so the key flips into linked mode: one static_key_mod node is built to wrap what key->entries already pointed at (the run of module A), a second node is built for module B, and key->next now heads that two-node list — exactly the diagram above.

Exactly two functions drive every transition between those three states: jump_label_add_module() when a module loads, and jump_label_del_module() when one unloads. Neither is called directly — both run from a module notifier5 (jump_label_module_notify(), registered with .priority = 1) that hooks MODULE_STATE_COMING/ MODULE_STATE_GOING — the two transitions the module loader fires while a module is being mapped in and while it is being torn down.

The priority value is not an arbitrary tie-breaker — it enforces a strict order. Notifier chains run higher-priority callbacks first, and while jump labels register at .priority = 1, tracepoints register their own, separate module notifier at .priority = 0 (lower) — so the jump-label notifier always runs first. That ordering matters because tracepoints are themselves built on static keys. When a new module loads (MODULE_STATE_COMING), its notifier wants to start touching the static keys behind its own tracepoints — but those keys are only ready to be touched once jump_label_add_module() has registered (and patched where needed), the jump entries of that module. Running the jump-label notifier first guarantees that the setup is already done by the time the tracepoint notifier runs.

11.1 Loading: jump_label_add_module()

This is the function that actually produces every transition described above, once per key contributed by the newly-loaded module. Trimmed of its early-return-if-empty check, the __init-flag bookkeeping §9.1 already covered, and its -ENOMEM error paths, here it is, one piece at a time.

The signature, its local state, and the first thing it does:

static int jump_label_add_module(struct module *mod)
{
        struct jump_entry *iter_start = mod->jump_entries;
        struct jump_entry *iter_stop = iter_start + mod->num_jump_entries;
        struct jump_entry *iter;
        struct static_key *key = NULL;
        struct static_key_mod *jlm, *jlm2;

        jump_label_sort_entries(iter_start, iter_stop);

jump_label_sort_entries() sorts the jump table of this module — the same routine §9.1 used on the vmlinux-wide table, for the same reason. After sorting, all entries that reference the same static_key sit next to each other in the array.

That adjacency matters because the work this function does — allocating static_key_mod nodes, wiring them into the list, deciding whether to patch — is per-key, not per-call-site. A module that calls trace_sched_switch() at ten different places still needs just one list node for that tracepoint key, not ten. With entries sorted by key, the loop can handle each distinct key exactly once: run the per-key body when the first entry for a key appears, then skip all subsequent entries for the same key with a single pointer comparison.

That skip pattern is what the top of the loop implements:

        for (iter = iter_start; iter < iter_stop; iter++) {
                struct static_key *iterk = jump_entry_key(iter);

                if (iterk == key)
                        continue;
                key = iterk;

The for loop advances iter through every entry in the table, one at a time. On each iteration, jump_entry_key() reads which static_key that entry belongs to. If it is the same key the loop just finished processing, continue skips the entry — all per-key work was already done when the first entry of that group was reached. Only when iterk differs from key does execution fall through into the per-key body below.

Two consequences follow from that design. First, the per-key body runs once per distinct key in this module, not once per call site, and each of the three states from the introduction of §11 corresponds to exactly one path through it. Second, iter at the point where the body runs always points at the first entry of a key group in the sorted table — the start of a contiguous run. That matters at the end of the function, where iter is passed as a range start to __jump_label_update(), which walks forward from there to patch every call site for that key in this module.

Once inside the per-key body, the first check asks whether this module even needs to consider the linked-list machinery at all — the one-home-only case:

                if (within_module((unsigned long)key, mod)) {
                        static_key_set_entries(key, iter);   /* one home only */
                        continue;
                }

within_module() asks whether the static_key struct itself — not a call site, not a jump entry, but the struct that holds enabled and entries — lives inside memory owned by mod.

If it does, then mod is, by construction, the very first and only object that has ever contributed call sites for this key. The reason is physical: before this module loaded, the memory backing that static_key was not even mapped. No other module or vmlinux could have built a jump entry referencing an address that did not yet exist. The one-home-only direct-pointer form is therefore not just adequate here, it is the only form this key has ever needed — which is why this branch continues past all the linked-list machinery that follows.

A key that fails that check has call sites outside this module, so before building any list node the sealed case has to be checked:

                if (static_key_sealed(key))
                        goto do_poke;   /* sealed: patch once, keep no link */

A sealed key has already forgotten its entries/next union for good (§9.1), so there is nothing left to link this module into. Instead the goto skips straight past all the list-building below, to a single comparison shared with the ordinary case — reached down in the text.

An ordinary, unsealed key that reaches this point is the linked case. This is where the list from the introduction of §11 actually gets built or grown, in two steps.

The first step handles a one-time transition. Up to this point the key might still be in direct-pointer form — one pointer, one contributing object, no list. Before anything can be prepended to it, that existing pointer has to be wrapped in a list node so there is a next field to link through:

                jlm = kzalloc(sizeof(struct static_key_mod), GFP_KERNEL);
                if (!static_key_linked(key)) {
                        jlm2 = kzalloc(sizeof(struct static_key_mod), GFP_KERNEL);
                        scoped_guard(rcu)
                                jlm2->mod = __module_address((unsigned long)key);
                        jlm2->entries = static_key_entries(key);
                        static_key_set_mod(key, jlm2);
                        static_key_set_linked(key);
                }

static_key_linked() returns false only when the key is still in direct-pointer form, meaning exactly one object has contributed to it so far. The code allocates jlm2 to retroactively wrap whatever key->entries already pointed at — the run belonging to that first object — discovers which module owns that object via __module_address(), and installs jlm2 as the head of a new one-node list. Once static_key_set_linked() flips the linked bit, this wrapping step never runs again for this key — every later module finds static_key_linked() already true and skips straight into the second step.

The second step prepends a node for the newly-arriving module onto the now-guaranteed-to-exist list:

                jlm->mod = mod;
                jlm->entries = iter;
                jlm->next = static_key_mod(key);
                static_key_set_mod(key, jlm);
                static_key_set_linked(key);

jlm->entries is set to iter — the first entry of this key group in the sorted table, the same pointer the one-home-only case would have stored directly into key->entries. static_key_set_mod() makes jlm the new head of the list, with the previous head linked behind it via jlm->next.

Both the sealed-key goto and the ordinary path above fall into the same final check, once per key:

do_poke:
                if (jump_label_type(iter) != jump_label_init_type(iter))
                        __jump_label_update(key, iter, iter_stop, true);
        }
        return 0;
}

Every call site a module ships with is compiled around one fixed assumption: the default value the key had at compile time — the same type ^ branch formula from §6.1, exposed here as jump_label_init_type() — which decided whether the assembler macro emitted this site as a nop or a jmp in the first place (§6.2), frozen into the module from then on. But the live jump label type value can have moved away from that compiled-in default already, if something else called static_branch_enable()/disable() on this key before the module ever loaded. do_poke catches exactly that mismatch and patches the new sites immediately, before the code of the module has a chance to run and observe them in the wrong state — whether it arrived there via the sealed-key goto above, or by falling through normally after linking the key into the list.

11.2 Toggling an already-linked key

When someone calls static_branch_enable()/static_branch_disable()/static_branch_inc()/static_branch_dec() on a key whose call sites span more than one object, the flat array scan from §9.5 is not enough — the sites are scattered across separate per-object tables, each with its own bounds, and the patcher has to visit every one of them.

jump_label_update() detects that case: if static_key_linked() returns true, it hands off to __jump_label_mod_update() instead of doing the walk itself. That function walks the linked list, calling __jump_label_update() once per node:

static void __jump_label_mod_update(struct static_key *key)
{
        struct static_key_mod *mod;

        for (mod = static_key_mod(key); mod; mod = mod->next) {
                struct jump_entry *stop;
                struct module *m;

                if (!mod->entries)
                        continue;

                m = mod->mod;
                if (!m)
                        stop = __stop___jump_table;
                else
                        stop = m->jump_entries + m->num_jump_entries;
                __jump_label_update(key, mod->entries, stop,
                                    m && m->state == MODULE_STATE_COMING);
        }
}

The loop visits each static_key_mod node in the list and passes that entries pointer and matching stop bound to __jump_label_update(). Three details in that loop deserve explanation.

First, stop cannot be a single kernel-wide constant the way it is in the flat walk from §9.5. Each module has its own private __jump_table section (§7.4), bounded by the jump_entries/num_jump_entries fields of that module. Without the correct per-node bound, __jump_label_update() would walk forward past the end of a module table into unrelated memory. The loop has to look up that bound for each node individually.

Second, m — the mod field of the list node, not the function parameter — can be NULL. That is not an error: it is the node that represents vmlinux itself. The NULL originates in §11.1: the first-time linking step calls __module_address() to discover which module owns the key, and __module_address() returns NULL for addresses inside the core kernel. For that node the correct bound is the global __stop___jump_table, which marks the end of the vmlinux-wide table.

Third, mod->entries can genuinely be NULL. That happens when an object defines a key but contains no call sites for it — the linking code from §11.1 still creates a list node to represent that object (it is a contributor), but there is nothing there to patch. Concretely: a module exports a DEFINE_STATIC_KEY_FALSE, and call sites in other modules are the only consumers. Skipping the node is correct.

The final argument, m && m->state == MODULE_STATE_COMING, tells __jump_label_update() whether to also patch entries living in the __init section of the module. A module still in MODULE_STATE_COMING has not finished running its init function yet, so its init section is still mapped and its call sites there are reachable. A module already in MODULE_STATE_LIVE has discarded that section — patching into freed memory would be a use-after-free, not a harmless no-op.

11.3 Unloading: jump_label_del_module()

jump_label_del_module() is the mirror image of jump_label_add_module(), run on MODULE_STATE_GOING: for each distinct key this module contributes to, it finds and removes the static_key_mod node that §11.1 created. The text-patching machinery does not need to undo anything here — the .text section of the module is about to be unmapped entirely, so the patchable sites in it simply cease to exist. What does need cleaning up is the linked-list bookkeeping that still references them.

The function uses the same sorted-table, skip-by-key loop structure as loading. Three of the four skip cases mirror §11.1 directly:

        for (iter = iter_start; iter < iter_stop; iter++) {
                if (jump_entry_key(iter) == key)
                        continue;

                key = jump_entry_key(iter);

                if (within_module((unsigned long)key, mod))
                        continue;

                /* No @jlm allocated because key was sealed at init. */
                if (static_key_sealed(key))
                        continue;

                /* No memory during module load */
                if (WARN_ON(!static_key_linked(key)))
                        continue;

The first three are familiar from §11.1: skip duplicate entries for the same key (the sorted-table grouping from the sort step), skip keys whose struct lives inside this module (the struct vanishes with the module, so there is no list to update), and skip sealed keys (no node was ever created for them). The fourth is a defensive check: if the key is not in the linked state at this point, something went wrong during loading — likely an allocation failure that jump_label_add_module() could not recover from. The WARN_ON flags the inconsistency without crashing6, and the continue skips the key rather than dereferencing a pointer that was never set up.

For every key that passes all four checks, the function walks the linked list to find and splice out the node belonging to this module:

                prev = &key->next;
                jlm = static_key_mod(key);

                while (jlm && jlm->mod != mod) {
                        prev = &jlm->next;
                        jlm = jlm->next;
                }

                /* No memory during module load */
                if (WARN_ON(!jlm))
                        continue;

                if (prev == &key->next)
                        static_key_set_mod(key, jlm->next);
                else
                        *prev = jlm->next;

                kfree(jlm);

The while loop advances through the list until it finds the node whose mod field matches the departing module, keeping prev pointed at the next pointer of the preceding node so the splice has something to patch. If no matching node is found — again a sign that something went wrong during loading — a second WARN_ON fires and the key is skipped. Otherwise, the standard singly-linked-list splice removes the node: if it was the head of the list (prev == &key->next), the next pointer of the key itself is updated via static_key_set_mod(); if it was somewhere in the middle, the next pointer of the predecessor is patched directly.

After the splice, one more step checks whether the list can be eliminated entirely:

                jlm = static_key_mod(key);
                /* if only one etry is left, fold it back into the static_key */
                if (jlm->next == NULL) {
                        static_key_set_entries(key, jlm->entries);
                        static_key_clear_linked(key);
                        kfree(jlm);
                }

If exactly one node remains, the key no longer needs the list form — only one object still contributes to it. The code folds the entries pointer of that last node back into key->entries directly, clears the LINKED bit, and frees the node. This is the linked-state transition from §11.1 running in reverse: a key that needed the list form only because two objects happened to overlap returns to the one-home-only direct-pointer form the moment that overlap ends.


12 Fallback: CONFIG_JUMP_LABEL=n

JUMP_LABEL is optional. Most distro kernels end up with it on — arm64 selects it outright, and on x86 PREEMPT_DYNAMIC pulls it in — but a minimal config can legitimately leave it off. Everything from §5 onward assumed the option was enabled. This section covers what happens when it is not.

Without CONFIG_JUMP_LABEL, the struct shrinks to a bare counter. The entries/next/type union from §7 is compiled out entirely:

struct static_key {
        atomic_t enabled;
};

jump_label_init() sets static_key_initialized to true and returns. There is no jump table to sort, no entries to pre-patch.

The call-site macros turn into ordinary branch-hinted conditionals. static_branch_likely() and static_branch_unlikely() reduce to:

#define static_branch_likely(x)   likely_notrace(static_key_enabled(&(x)->key))
#define static_branch_unlikely(x) unlikely_notrace(static_key_enabled(&(x)->key))

No asm goto, no jump table, no patching — just a likely_notrace()/unlikely_notrace() hint around a read of enabled.

static_key_count() is a plain raw_atomic_read():

static __always_inline int static_key_count(struct static_key *key)
{
        return raw_atomic_read(&key->enabled);
}

Compare this with the CONFIG_JUMP_LABEL=y version in kernel/jump_label.c, which clamps negative values (n >= 0 ? n : 1). That clamp exists because static_key_enable() temporarily sets enabled to -1 while the patching pass runs (see §9.2). Without patching, enabled never goes negative, so the clamp is unnecessary.

static_key_enable() and static_key_disable() are simpler for the same reason. Each one checks whether enabled already holds the target value and returns early if so. If it holds something unexpected (neither 0 nor 1), a WARN_ON_ONCE fires. Otherwise, a plain atomic_set() writes the new value. No cmpxchg, no intermediate -1, no jump_label_update() call.

jump_label_lock()/jump_label_unlock() are empty stubs — there is no patch pass to serialize.

The net effect on every hot path is exactly the cost §2 opened with: a memory load of enabled, a compare, and a conditional branch. The likely()/unlikely() hint steers the branch predictor the same way a compiled-in NOP or JMP would, but it cannot eliminate the branch itself — that is the optimization that CONFIG_JUMP_LABEL=y adds.

Jump labels are an optimization, not a correctness feature: behavior matches; only the mechanism changes.


13 Worked micro-example (bytes on the wire)

Every mechanism described so far — the asm helpers, the jump-table entry, the objtool hack, the size-discovery check, the enabled state machine, and the INT3 protocol — touches one concrete call site at some point in its life. This section traces a single, minimal site through that entire life, from compile time through one enable and one disable, close enough to see the actual instruction bytes change rather than just the names of the steps that change them. Suppose:

DEFINE_STATIC_KEY_FALSE(k);

void f(void)
{
        if (static_branch_unlikely(&k))
                printk("on\n");
        something();
}

Two small assumptions turn this from a symbolic description into an actual trace, and neither changes anything about how the mechanism works, only which specific numbers show up:

  1. f() is called after boot, on a live, multi-CPU system — the interesting §10 INT3 path, not the single-CPU §10.1 boot shortcut (a boot-time toggle of this same key would use text_poke_early() instead, with no INT3 involved at all).
  2. The compiler places l_yes — the out-of-line block containing the printk() call — 80 bytes past the end of the patch site. 80 fits in a signed byte, so the build-time trick from §6.5 picks the 2-byte JMP rel8 encoding rather than the 5-byte rel32 form. A farther l_yes would just mean 5 bytes instead of 2 everywhere below (§5.2); nothing else about the trace would change.

From the source code above, the compiler, linker, and objtool produce the bytes that sit in vmlinux at the patch site. For orientation, here is where the three pieces end up — the patch site in .text, the out-of-line target, and the sidecar entry in __jump_table:

 .text (function f)                  __jump_table (one entry)
 ──────────────────────────         ──────────────────────────
       ...                           code:   delta to 1:
  1:   [  2 bytes  ]  patch site     target: delta to l_yes
       ...                           key:    delta to &k.key + 2
       call something                         bit 0 = 0 (branch)
       ret                                    bit 1 = 1 (objtool)
       ...
  l_yes:     80 bytes past 1:+2
       call printk
       jmp back ------> (after 1:+2)

The six steps below trace how those bytes arrive at their final state:

  1. k is a FALSE key read with static_branch_unlikely(). Per the table in §6.2, that combination calls the nop-default helper arch_static_branch(&k.key, false), not the jmp-default arch_static_branch_jump(). That is the type ^ branch formula from §6.1 at work: type = 0 (FALSE), branch = 0 (unlikely), so type ^ branch = 0 — the hint agrees with the default, and the nop-default path gets chosen.

  2. The asm goto inside arch_static_branch() emits 1: jmp l_yes at the patch site, plus one raw row in __jump_table via JUMP_TABLE_ENTRY() (§6.4): code is the self-relative distance to 1:, target is the self-relative distance to l_yes, and key is the self-relative distance to &k.key + 0 + 2.

    That + 2 puts a 1 in bit 1 of the stored key address. Bit 0 is branch — here 0, matching the unlikely() hint. Bit 1 is the build-time signal to objtool: “NOP this jmp” (§6.5); after boot, jump_entry_set_init() repurposes this same bit as the __init-text flag. At runtime, jump_entry_key() masks both bits off to recover the real address of k.

  3. The assembler picks the encoding by the real distance to l_yes — here, 80 bytes forward (assumption 2 above), well inside the -128..+127 reach of a signed byte. It emits the 2-byte JMP rel8 form: opcode EB, followed by disp = dest - (addr + insn_size) = 80 = 0x50 (the disp formula from §5.2). The two bytes actually sitting at 1: right after assembly, before objtool ever runs, are EB 50.

  4. During the build, objtool calls handle_jump_alt() (§6.5 step 3), which sees bit 1 set in the stored key operand and rewrites those exact two bytes, in place, from the jmp (EB 50) to the 2-byte NOP (66 90, the NOP encoding from §5.2) — same size, so nothing around the site shifts. By the time vmlinux is linked, the live bytes at 1: are already 66 90, and objtool itself is long gone.

  5. At boot, jump_label_init() sorts __jump_table by key, then loops over every entry (§9.1). For the one entry belonging to k, the loop body does the following:

     jump_label_init()
     ├── jump_label_sort_entries()                sort __jump_table by key
     └── for each entry:                          (k has exactly one)
         ├── jump_label_type() = NOP              enabled(0) ^ branch(0)
         │   └── arch_jump_label_transform_static()  no-op on x86
         ├── jump_entry_set_init(entry, false)    code not in __init
         └── static_key_set_entries(&k, entry)    wires k.entries → entry
    

    jump_label_type() computes enabled(0) ^ branch(0) = NOP — the expected state matches the live bytes, which are already 66 90. So arch_jump_label_transform_static() is a genuine no-op: x86 never overrides the generic fallback, whose entire body is one comment, /* nothing to do on most architectures */. No instruction bytes get rewritten at this site.

  6. The same loop iteration does two pieces of bookkeeping that wire k into the runtime data structures. First, jump_entry_set_init() checks whether the code at 1: lives in __init text — it does not, so bit 1 of the stored key address (the same bit objtool used in step 2) gets cleared to 0. Second, static_key_set_entries() points k.entries at this entry — the pointer that jump_label_update() will follow when static_branch_enable(&k) runs later. With that in place, the hot path in f() from the very first time it runs is: decode 66 90 (falls through, no load of k, §5.1), then call something.

In summary, the same two bytes at 1: passed through three stages before the kernel ever ran a line of f():

 Stage                Bytes at 1:    Why
 ─────────────────    ───────────    ──────────────────────────────────
 After assembly       EB 50 (JMP)   assembler picks rel8 for +80 distance
 After objtool        66 90 (NOP)   handle_jump_alt() sees bit 1 in key
 At boot              66 90 (NOP)   jump_label_init(): NOP expected, NOP found

13.2 Runtime static_branch_enable(&k)

§10.6 already lays out the full call chain a batch of sites goes through on enable; k has exactly one entry, so this is that same chain with the actual bytes for every step filled in. Both directions of the public API are one-line macros in include/linux/jump_label.h:

#define static_branch_enable(x) static_key_enable(&(x)->key)
#define static_branch_disable(x)    static_key_disable(&(x)->key)

The full call chain for this one enable, with the concrete byte values for k filled in at each level:

 static_branch_enable(&k)
 └─ static_key_enable(&k.key)
    ├── cpus_read_lock / jump_label_lock
    ├── enabled: 0 --> -1              callers already see "on"
    ├── jump_label_update(key)
    │   └─ __jump_label_update()
    │      ├── type = enabled(true) ^ branch(0) = JMP
    │      ├── arch_jump_label_transform_queue()
    │      │   └─ __jump_label_patch(entry, JMP)
    │      │      ├── size  = 2        (live-decode of 66 90)
    │      │      ├── code  = EB 50    (text_gen_insn)
    │      │      ├── nop   = 66 90    (x86_nops[2])
    │      │      ├── memcmp(addr, nop) -- pre-flight OK
    │      │      └── smp_text_poke_batch_add(addr, EB 50, 2)
    │      └── arch_jump_label_transform_apply()
    │          └─ smp_text_poke_batch_finish()
    │             └── INT3 three-phase: 66 90 --> EB 50
    ├── enabled: -1 --> 1              atomic_set_release
    └── jump_label_unlock / cpus_read_unlock

The numbered steps below walk through this chain in detail:

  1. static_branch_enable(&k) expands to static_key_enable(&k.key), which acquires jump_label_lock() and walks the 0 → -1 → (patch) → 1 state machine from §9.2. enabled starts at 0, gets set to -1 (“enabling in progress”), and static_key_count() already reports -1 as “on” — so no caller sees a false “off” window while patching runs. Then jump_label_update(&k.key) does the actual patching, and only after it returns does enabled get published as 1 with release ordering.
  2. jump_label_update() finds the one entry of k and computes jump_label_type(entry) = enabled ^ branch. enabled is the transient -1 from step 1, which static_key_enabled() reports as true; branch is the stored hint bit, 0. true ^ false = JMP — the live nop becomes a jmp.
  3. arch_jump_label_transform_queue() calls __jump_label_patch(), which re-derives the size by decoding the live bytes at 1: — arch_jump_entry_size() returns 2, matching what objtool left behind (§8). The function then builds both sequences: nop = 66 90 (from x86_nops[2]) and code = EB 50 (from text_gen_insn() — the same bytes computed in §13.1 step 3, since addr and dest have not moved). Because this is a nop→jmp transition, it calls memcmp() to verify the live bytes are currently 66 90; a mismatch would be a BUG(). They match, so the patch 66 90 → EB 50 gets queued via smp_text_poke_batch_add() (§10.2).
  4. The loop of jump_label_update() over the entries of k ends here — there is only the one — so arch_jump_label_transform_apply() immediately calls smp_text_poke_batch_finish(), which runs the three-phase protocol from §10.3 on this one queued site, now with the real two bytes instead of a placeholder:
   start (before)     66 90              any fetch: executes the 2-byte NOP

   phase 1 (INT3 in)  cc 90              any fetch: #BP -> handler emulates
                      ^^                 the pending JMP (jumps to l_yes)
                      trap byte

                      -------- IPI sync --------

   phase 2 (tail in)  cc 50              same as phase 1: byte 0 still
                      ^^                 traps and gets emulated — byte 1
                      still traps        is now its final value underneath

                      -------- IPI sync --------

   phase 3 (done)     eb 50              any fetch: executes the real
                      ^^ real opcode     JMP rel8 directly, no trap needed

                      -------- IPI sync --------
  1. From the instant the sync in phase 1 completes, any CPU landing on this address already gets the effect of the jump via emulation (§10.3); phases 2-3 only retire the trap-and-emulate path in favor of the real bytes. Once phase 3 lands, every subsequent call to f() decodes EB 50, jumps 80 bytes forward into l_yes, runs printk("on\n"), then hits the compiler-emitted jmp back (§3) and falls into something().

13.3 static_branch_disable(&k): same protocol, asymmetric math

Disabling runs the mirror call, static_branch_disable(&k) → static_key_disable(&k.key) — a single atomic_cmpxchg(&key->enabled, 1, 0) instead of the enable state machine, because disable has no in-progress state to protect (§9.3: a reader who sees stale “on” for a few more instructions is exactly the harmless case, unlike stale “off”). If that cmpxchg succeeds, jump_label_update() runs again, this time computing jump_label_type() = false ^ false = NOP (enabled now 0, branch still 0).

__jump_label_patch() rebuilds the same two sequences as before — nop = 66 90, code = EB 50, both unchanged, since addr, dest, and size haven’t moved — but this time expects the live bytes to be the jmp and installs the nop. That is the one genuine asymmetry in this whole worked example: enabling always needs a fresh, target-specific displacement computed by text_gen_insn(); disabling never does, because the encoding of a NOP does not depend on where the branch would have gone — the exact same fixed bytes go back every time a site of this size is disabled, no matter what it was jumping to.

The same three-phase protocol (§10.3) runs again, in the opposite byte direction:

 start (before)     eb 50              executes the JMP rel8

 phase 1 (INT3 in)  cc 50              #BP -> handler now emulates the
                    ^^                 pending NOP (a JMP with disp == 0
                    trap byte          — not a special case)

                    -------- IPI sync --------

 phase 2 (tail in)  cc 90              byte 1 now its final value; byte 0
                    ^^                 still traps
                    still traps

                    -------- IPI sync --------

 phase 3 (done)     66 90              executes the real NOP directly
                    ^^ real opcode

                    -------- IPI sync --------

Put together, the build-time 66 90 of this one site, the EB 50 of the first enable, and the 66 90 of this disable again are the entire lifecycle that §5-§10 spent this whole tutorial describing in the abstract — the same two bytes, chosen and re-derived by a different mechanism at each stage, but never touched by anything other than the three sanctioned writers: the assembler once at compile time, objtool once at build time, and __jump_label_patch() through the protocol from §10.3 every time after that.

Stage Live bytes at 1: Who wrote them
After assembly, before objtool EB 50 Compiler/assembler (§6.3, §5.2)
After objtool, at boot 66 90 handle_jump_alt() (§6.5)
After static_branch_enable(&k) EB 50 __jump_label_patch() via INT3 protocol (§10.3)
After static_branch_disable(&k) 66 90 Same, reverse direction

14 Further reading in-tree


  1. A 4-byte signed displacement covers a range of ±2 GiB. Because x86_64 kernels are compiled using the -mcmodel=kernel option, the entire kernel image is restricted to the upper 2 GiB of the address space. This model guarantees that any relative offset within the kernel binary falls within the reach of a rel32 displacement. ↩

  2. __ro_after_init is a section attribute (include/linux/cache.h, __section(".data..ro_after_init")) for data that is written during boot but never again afterward — unlike const, which the compiler must be able to enforce at compile time, this is a promise the author makes about runtime behavior. The kernel makes that promise real at mark_rodata_ro() time, when the whole .data..ro_after_init section is remapped read-only in the page tables, so any later write attempt — a bug, or an author breaking their own promise — faults instead of silently corrupting state. ↩

  3. _stext is a linker-defined symbol, not a C variable — it marks the address where the .text section of the kernel begins, set directly in arch/x86/kernel/vmlinux.lds.S. Every function in core kernel text lives at some fixed offset from it, which is what lets smp_text_poke_loc.rel_addr (a plain s32) address any patch site with 4 bytes instead of the full 8-byte pointer jump_entry.key needs (§6.4) for its potentially-far-away static_key. ↩

  4. Strictly, phases 2 and 3 each sync only if at least one site in the batch actually needed a write in that phase, tracked by a do_sync counter inside smp_text_poke_batch_finish(). The phase 3 sync, for instance, is skipped if the final first byte of every site already happens to equal INT3. An ordinary nop-to-jmp toggle always writes something in both phases, so three rounds is what actually happens in practice; the fixed count just isn’t unconditional at the code level. ↩

  5. A module notifier is a callback registered against the module-load notifier chain (module_notify_list) via register_module_notifier(). The module loader walks that chain with blocking_notifier_call_chain() at each state transition (MODULE_STATE_COMING, MODULE_STATE_GOING, etc.), invoking every registered struct notifier_block in priority order — this is the generic mechanism subsystems use to react to modules loading/unloading, not something jump labels invented. ↩

  6. Unless panic_on_warn = 1 ↩

comments powered by Disqus