What every kernel programmer should know about Jump Labels
• kernel
- 1 Introduction
- 2 The problem jump labels solve
- 3 The mental model, in one diagram
- 4 How to use static keys (the cookbook)
- 5 Hardware background (why this is hard)
- 6 What the compiler emits (x86_64)
- 7 Core data structures
- 8 Size of the patchable site on x86 (runtime)
- 9 Life of a static key: boot, enable, disable
- 10 x86 text patching: the gory details
- 11 Modules: the trickiest part
- 12 Fallback:
CONFIG_JUMP_LABEL=n - 13 Worked micro-example (bytes on the wire)
- 14 Further reading in-tree
1 Introduction
Kernel hot paths frequently evaluate conditions that rarely change:
tracepoint activation, hardware mitigations, or debug logging. In a
naive implementation, a standard conditional branch (if (flag)) must
load the state from memory on every iteration, incurring a guaranteed
cache access and wasting execution cycles.
Jump labels bypass this overhead. By dynamically modifying the
executable binary, this mechanism rewrites the hot path at runtime.
While a feature remains disabled, the CPU encounters a single, fast
nop (or an unconditional jump around out-of-line code). Toggling the
state swaps the instructions in-place to redirect execution. This
architecture accepts expensive update-time coordination in exchange for
zero run-time penalty in the common path.
This guide divides into two distinct sections. The first half—spanning from §2 to §4—is a cookbook and API reference for developers who need to integrate static keys into kernel modules and core code.
The second half delves into the low-level implementation details on the
x86_64 architecture under Linux 7.2. It traces the machine
instructions generated by the compiler, the design of the jump-table
metadata, objtool integration, boot-time initialization, live code
patching, the multi-processor INT3 synchronization protocol, and
module loading behavior.
To prevent confusion, this tutorial strictly distinguishes between two related terms:
- Jump label: The low-level architecture-specific mechanism,
consisting of a patchable instruction site in
.textand a tracking entry in the metadata section. - Static key: The developer-facing, high-level API
(
DEFINE_STATIC_KEY_FALSE(),static_branch_unlikely()) used to express conditional branches in C.
2 The problem jump labels solve
Kernel code is full of conditional checks that guard optional or debugging features: “is tracing enabled for this tracepoint?”, “is this security module active?”, “is this debugging feature on?”. A standard C implementation:
if (some_feature_enabled)
do_something();
Even when some_feature_enabled remains false almost indefinitely, the
CPU must still perform three steps:
- Load
some_feature_enabledfrom its cache line into a register. - Compare the register value against zero.
- Predict and branch based on the result.
Modern branch predictors handle the conditional branch in step 3 with high accuracy, hiding the misprediction penalty. However, branch prediction cannot bypass the guaranteed memory load in step 1. When this check resides in a hot path executed millions of times per second (such as the scheduler, network packet processing, or tracepoint triggers), the cumulative overhead of those memory accesses becomes a measurable performance bottleneck.
Jump labels eliminate both the memory access and the comparison for
the common path by rewriting the instructions at runtime. When the
feature is inactive, the instruction pipeline encounters a no-operation
(nop) instruction (or an unconditional branch that bypasses the
out-of-line code). When a subsystem activates the feature, the kernel
traverses all call sites registered for that static key and patches the
instructions in place.
The design tradeoff is stark: toggling is expensive because it requires a machine-wide CPU synchronization and precise text patching, but running the hot path is virtually free.
3 The mental model, in one diagram
When a developer guards a conditional code block using
static_branch_unlikely(),
the compiler generates an instruction layout where the hot path remains
linear and cold blocks reside out-of-line. The entry point of this
sequence is a patchable location whose instruction is determined by the
runtime state of the static key:
SOURCE CODE FEATURE OFF (common) FEATURE ON
------------ -------------------- ----------
if (static_branch_unlikely nop (2 or 5 bytes) jmp .Lout_of_line
(&my_key)) {
rare_code(); .Lout_of_line: .Lout_of_line:
} rare_code(); rare_code();
jmp back jmp back
(unreachable without a jmp)
In the disabled state (“FEATURE OFF”), execution flows sequentially
without branching. The out-of-line block containing rare_code() is
preserved in the binary, but because no active jump instruction targets
the .Lout_of_line label, normal instruction execution bypasses it
entirely.
The unconditional return jump (jmp back) at the end of the out-of-line
block is a standard compiler optimization rather than part of the jump
label framework. The compiler moves the cold block to the end of the
function and inserts a jump to return execution to the statement
immediately following the conditional block. This return branch remains
static throughout execution; only the entry-point instruction at the top
of the site is patched when toggling the key.
| Key State | Instruction in the Hot Path | Cost when the Feature is Disabled | Cost when the Feature is Enabled |
|---|---|---|---|
| Disabled (for an unlikely site) | nop |
Negligible (linear execution, pipeline fall-through, no memory access) | Unreachable (the code path is bypassed and cannot execute) |
| Enabled | jmp .Lout_of_line |
Single unconditional jump | Entry jump + cold path execution + return jump |
Contrast this mechanism with the conventional conditional statement in §2, which always incurs memory load and comparison overhead.
4 How to use static keys (the cookbook)
Toggling a static key is not a simple memory write; it is a live code modification that patches executable instructions across every online CPU core. This operation incurs a heavy synchronization penalty. Consequently, the API offers specialized variants to manage update frequency, control caller authorization, and coordinate multiple owners.
4.1 Minimal example
A minimal static key implementation spans three phases: declaring the key, guarding the conditional branch in the hot path, and toggling the branch target from a separate control path.
#include <linux/jump_label.h>
DEFINE_STATIC_KEY_FALSE(foo_key);
void hot_path(void)
{
/* Fast path: compiled as NOP while the key is false. */
if (static_branch_unlikely(&foo_key))
do_rare_thing();
do_common_work();
}
void foo_enable(void)
{
static_branch_enable(&foo_key); /* slow path: patches text */
}
void foo_disable(void)
{
static_branch_disable(&foo_key); /* slow path: patches text */
}
Declaring the key with
DEFINE_STATIC_KEY_FALSE()
instantiates a
struct static_key_false,
which wraps an
atomic_t
counter in a unique wrapper type. This type-level distinction allows the
compiler to differentiate the key from a
struct static_key_true.
Pairing this key with
static_branch_unlikely()
instructs the compiler to emit a nop instruction for the initial,
disabled state.
4.2 Choosing TRUE vs FALSE and likely vs unlikely
Selecting an incorrect combination of key declaration and branch macro
does not affect behavioral correctness, but it introduces a subtle
performance penalty. If the chosen combination disagrees with the
steady-state execution flow, the compiler generates a jmp instruction
instead of a nop on the hot path. This layout penalty remains
invisible to functional testing and cannot be corrected at runtime. The
generation of either
arch_static_branch
or
arch_static_branch_jump
instruction sequences is determined entirely at compile time.
Selecting the optimal configuration depends on two design criteria:
- The default state at boot time: Most optional features default
to a disabled state, requiring
DEFINE_STATIC_KEY_FALSE. Features that remain enabled unless explicitly deactivated requireDEFINE_STATIC_KEY_TRUE. - The expected steady-state execution path: Developers must apply
static_branch_unlikely()when the conditional block represents the rare execution path. Conversely,static_branch_likely()must guard paths where the conditional block represents the common execution path.
These declarations and branch macros can be combined arbitrarily; a
false-default key is compatible with both likely and unlikely macros, as
is a true-default key. The jump label subsystem coordinates these
combinations to ensure that the initial default state always compiles to
a cheap nop instruction.
A typical implementation for optional features uses the following pattern:
DEFINE_STATIC_KEY_FALSE(feature_key);
if (static_branch_unlikely(&feature_key))
rare_enabled_path();
4.3 Boolean enable vs refcounted enable
The distinction between boolean and reference-counted interfaces
addresses a classic coordination failure. If two independent subsystems
call
static_branch_enable()
on a shared key, both expect the code path to remain active. If the
subsystem that finishes first calls
static_branch_disable(),
the code path is patched off for both, leaving the second subsystem
silently broken. Reference counting prevents this premature
deactivation.
| API | Semantics | When to use |
|---|---|---|
static_branch_enable / static_branch_disable |
Force enabled count to 1 or 0 | Single owner; simple on/off |
static_branch_inc / static_branch_dec |
Refcount; patch only on 0↔1 | Multiple independent users |
static_branch_slow_dec_deferred |
Dec, but delay the 1→0 patch | Userspace-driven toggles |
Under the reference-counted API, a static key remains enabled as long as the counter is non-zero. The transition from zero to one triggers the initial text-patching operation to enable the branch. Subsequent increments are cheap atomic operations that bypass text patching entirely. Conversely, only the final decrement from one to zero triggers the text patch to disable the branch.
Mixing the boolean and reference-counted APIs on a single key leads to
corrupt state. Although both interfaces manipulate the underlying
enabled counter of
struct static_key,
they operate under conflicting assumptions. The boolean interface
expects a binary state (strictly zero or one), whereas the
reference-counted interface expects an arbitrary non-negative integer.
If a caller of static_branch_inc() has incremented the counter to two
or more, invoking
static_key_enable()
or
static_key_disable()
will trigger safety assertions rather than the expected behavior. When
static_key_enable() is called on a key whose value is already greater
than zero, it assumes the branch is active and returns immediately.
Conversely, if static_key_disable() is called when the reference count
is greater than one, it detects that the count does not match the
expected value of one required for a safe shutdown. In both scenarios,
the kernel emits a warning via
WARN_ON_ONCE
and aborts the operation, leaving the instruction patch unmodified.
4.4 Reading the state without taking the branch
Every example so far, including hot_path() in
§4.1, uses
static_branch_likely
or
static_branch_unlikely
as an if condition. The purpose of these macros is to compile directly
into a patchable branch instruction. However, a caller occasionally
requires the current boolean state of a key as an ordinary expression
rather than a patchable branch. The
static_key_enabled()
macro provides this capability:
if (static_key_enabled(&foo_key))
/* plain atomic read of the count — NOT the patched fast path */
Two patterns from the kernel illustrate why the API provides a dedicated read function rather than simply relying on a slow-path conditional branch.
State reporting represents the first pattern. During initialization, the
kernel logs configuration decisions rather than branching on them. For
example,
arch/x86/kernel/cpu/bugs.c
formats the active Spectre and IBPB mitigations into a message once at
boot:
pr_info(..., static_key_enabled(&switch_mm_always_ibpb) ? "always-on" : "conditional")
Because this message is generated only during early boot, compiling a
patchable fast-path instruction is unnecessary.
Control-plane guarding represents the second pattern. Before invoking
the path that patches instructions across every CPU,
drivers/md/dm-stats.c
verifies the current state of the key:
if (!static_key_enabled(&stats_enabled.key)) static_branch_enable(&stats_enabled);
Querying the state of the key first avoids executing a costly
text-patching sequence if the key is already active, preventing
redundant calls to
static_branch_enable().
Both scenarios require the current boolean state as an ordinary expression to print, compose, or evaluate during setup—operations that the patchable branch macros cannot accommodate.
On hot paths, however, callers must use the branch macros to ensure the
compiler generates patchable instructions. Substituting
static_key_enabled() in a hot path bypasses the jump label
infrastructure entirely. This forces the CPU to pay the cost of a
cache-line load on every execution—returning to the exact memory-access
bottleneck §2 that static keys are
designed to eliminate.
4.5 Keys must be global / static storage
A static key cannot reside on the stack or be dynamically allocated with
kmalloc().
Every static branch relies on compile-time inline assembly to register
the key address in the sidecar metadata section,
__jump_table.
Under the hood, the inline assembly block uses the immediate operand
constraint ("i") to pass the address of the key to the assembler.
Because the assembler must compute a relative offset between the jump
site and the key (as detailed in §7.3), the address of
the key must be a link-time constant.
If a developer attempts to pass a pointer to a stack variable or a heap-allocated struct, the compiler cannot satisfy the immediate constraint and will reject the code with a compilation error. This compile-time check prevents silent runtime memory corruption that would otherwise occur when a function returns and destroys its stack-allocated key, or when a dynamically allocated key is freed.
To define static keys correctly, always place them in global or
file-local static storage using
DEFINE_STATIC_KEY_FALSE.
The API provides several initialization macros for different scopes and
patterns:
/* Global key defined in a source file (.data section) */
DEFINE_STATIC_KEY_FALSE(global_key);
/* File-local key visible only within the translation unit */
static DEFINE_STATIC_KEY_FALSE(file_local);
/* Declaration for header files to share a global key */
DECLARE_STATIC_KEY_FALSE(global_key);
For grouping multiple toggles together, define an array using
DEFINE_STATIC_KEY_ARRAY_FALSE:
DEFINE_STATIC_KEY_ARRAY_FALSE(keys, 4);
if (static_branch_unlikely(&keys[i])) {
/* ... */
}
When a key should be conditionally defined based on a Kconfig option,
use
DEFINE_STATIC_KEY_MAYBE
paired with
static_branch_maybe:
DEFINE_STATIC_KEY_MAYBE(CONFIG_FOO, foo_key);
if (static_branch_maybe(CONFIG_FOO, &foo_key)) {
/* ... */
}
4.6 Read-only-after-init keys
While §4.5 detailed static keys
designed for lifetime mutability, certain hot-path conditions require
absolute immutability once configured. For instance, a hardware
mitigation decided during early boot should never be toggled again.
Relying solely on software-level discipline to prevent accidental
toggling is fragile. Instead, the kernel provides a hardware-enforced
guarantee through DEFINE_STATIC_KEY_FALSE_RO and
DEFINE_STATIC_KEY_TRUE_RO:
DEFINE_STATIC_KEY_FALSE_RO(configured_once_at_boot);
These macros place the underlying static key structure in the
__ro_after_init
section. The kernel permits modifications via
static_branch_enable()
or
static_branch_disable()
exclusively during early boot (before the system invokes
mark_rodata_ro()).
Once initialization finishes, a defense-in-depth architecture locks down
the key through two complementary mechanisms:
- Physical write protection: The memory page tables mapping the key
structure—specifically the
enabledcounter—are remapped to read-only. Any subsequent write attempt via the standard decrement or increment APIs triggers a hardware-level page fault. - Metadata sealing: The function
jump_label_init_ro()permanently seals the key (as discussed in §11). It zeroes out theentriespointer of the key and sets theJUMP_TYPE_LINKEDbit. Consequently, even if an attacker manages to bypass the page-table protection to overwriteenabled,jump_label_update()will find no registered call sites to patch.
The resulting freeze provides robust security hardening. If an attacker
leverages an arbitrary-write vulnerability elsewhere in the kernel to
compromise system memory, they still cannot disable a hardened
mitigation key. Because the enabled variable resides in
write-protected memory, any modification attempt is blocked by the MMU,
matching the security profile of traditional .rodata. Developers must
therefore use the _RO variants for “decide once at boot, then freeze”
security and performance knobs, reserving plain DEFINE_STATIC_KEY_*
for variables that genuinely require dynamic runtime toggling.
4.7 Rate-limited disable (userspace-facing knobs)
When userspace can toggle a feature rapidly, patching instructions on every transition degrades system performance. Each text-patching cycle forces a full round of inter-processor interrupts (IPIs) to broadcast the instruction changes across all online CPUs. If userspace can toggle a feature rapidly—such as via a sysctl or a socket option—naive patching on every transition turns a simple state change into a machine-wide synchronization storm.
The deferred static branch API introduces deliberate asymmetry. While
enabling remains immediate, disabling is deferred.
static_branch_deferred_inc()
is a direct alias for the standard reference-counting increment
static_branch_inc()—it
executes with zero delay. If a static key guards a critical tracepoint
or statistic counter, deferring the enable path would cause the kernel
to silently drop early events. The disable path can safely tolerate
delay; keeping a feature active for a few additional milliseconds is
harmless, whereas immediate text patching on high-frequency toggles is
not.
The deferred decrement function,
static_branch_slow_dec_deferred(),
implements this asymmetry by dividing the decrement logic into two
execution paths:
- Active reference remaining: If the reference count remains greater than one after the decrement, the function executes a plain, low-overhead atomic decrement. This path matches the behavior of the standard decrement in §4.3.
- Final reference transition: If the decrement would reduce the
reference count to zero, the function intercepts the operation.
static_key_dec_not_one()identifies this state and leaves the count intact at one.__static_branch_slow_dec_deferred()then invokesschedule_delayed_work()to initialize a timer fortimeoutjiffies. The physical decrement and subsequent text patching are deferred until the timer expires and invokesjump_label_update_timeout().
As a result, the feature remains fully active, fully patched, and
reference-counted at one throughout the entire timeout window. This
latency window enables event coalescing. If a new increment arrives
before the timer expires, the reference count rises from one to two via
the fast path described in §4.3.
When the timer eventually fires, the delayed work handler executes a
single, standard decrement. Because the count drops from two to one
rather than transitioning to zero, the decrement does not trigger a text
update. A rapid disable-then-enable sequence completed within the
timeout window bypasses the text-patching machinery entirely.
#include <linux/jump_label_ratelimit.h>
DEFINE_STATIC_KEY_DEFERRED_FALSE(sockopt_key, HZ);
/* enable immediately */
static_branch_deferred_inc(&sockopt_key);
/* disable — may wait up to `timeout` before actually patching off */
static_branch_slow_dec_deferred(&sockopt_key);
/* force pending delayed work to finish (e.g. module exit) */
static_key_deferred_flush(&sockopt_key);
Before freeing the enclosing
struct static_key_false_deferred—the
structure containing the static key, the timeout interval, and the
delayed_work state—the caller must invoke
static_key_deferred_flush().
This function blocks until any pending deferred disable work completes.
Flushing is critical during module unloading or dynamic memory
reclamation. If the enclosing memory is deallocated while the timer
remains active, the subsequent expiration of the timer will trigger
jump_label_update_timeout() on a freed delayed_work structure,
resulting in a use-after-free panic.
5 Hardware background (why this is hard)
Up to this point, we have treated nop and jmp instructions as
abstract logical states of a static branch. The remainder of this guide
explores how the kernel actually generates these instruction bytes and
dynamically hot-swaps them on a running system. This transition from
software-level branch hints to live runtime code patching relies on
three fundamental hardware realities: the pipelined execution of
instructions versus memory loads, the precise byte-level instruction
encodings of the x86_64 architecture, and the concurrency hazards that
prevent a multi-processor kernel from safely overwriting active
instructions with ordinary memory writes.
5.1 What a CPU actually does with instructions
A modern x86_64 core decouples instruction execution from the instruction stream using a deeply pipelined, out-of-order execution engine. Rather than executing instructions in a strict, sequential lock-step, the hardware continuously processes instructions through four key stages:
- Fetch: The hardware reads raw instruction bytes from the L1 instruction cache (L1i).
- Decode: Decoders convert these variable-length instruction bytes—ranging from 1 to 15 bytes on the x86_64 architecture—into fixed-length internal micro-operations (uops).
- Execute: Execution units dispatch uops out of order to specialized execution ports. A hardware branch predictor guesses the outcomes of conditional branches to keep these pipelines fully saturated.
- Retire: The processor commits results back in-order using a reorder buffer (ROB) to preserve the programmer-visible illusion of sequential execution.
While a highly accurate branch predictor can mask the latency of a
well-predicted conditional branch, it cannot eliminate the memory load
that feeds the check. Out-of-order execution makes the branch
instruction itself seem virtually free, as a correct prediction avoids
pipeline flushes. However, the underlying memory load that retrieves the
state of the flag (such as feature_enabled) must execute on every
single pass. This load consumes a load buffer entry, an L1 data cache
(L1d) read port, and an execution port. Under heavy data-cache pressure,
or if a writer on another CPU core modifies the flag, cache coherence
protocols invalidate the line. This invalidation forces the line to
bounce across cores, turning a trivial data-cache lookup into a
high-latency memory stall that halts the instruction window.
Replacing this conditional check with an unconditional nop or a direct
jmp removes the memory load entirely. Because there is no condition to
evaluate and no flag to read, the CPU avoids data-cache access
altogether. When the static branch is disabled, the core decodes the
nop at the frontend and discards it with minimal overhead, requiring
no execution ports or load buffers. There is no cache line to bounce
between cores and no state to track in the branch target buffer (BTB),
leaving the execution engine free to focus on the surrounding
instruction stream.
5.2 x86 instruction encoding: JMP and NOP
Replacing an instruction at runtime requires the original and replacement instructions to occupy the exact same number of bytes. Because x86 is a variable-length instruction set architecture (ISA), individual instructions vary from one to fifteen bytes in length. The specific instruction encodings manipulated by the jump label subsystem are:
| Instruction | Opcode bytes | Total size | Reach |
|---|---|---|---|
INT3 (breakpoint) |
CC |
1 byte | n/a |
JMP rel8 (short) |
EB xx |
2 bytes | -128..+127 bytes from the end of the instruction |
JMP rel32 (near) |
E9 xx xx xx xx |
5 bytes | ±2 GiB |
| 2-byte NOP | 66 90 |
2 bytes | — |
| 5-byte NOP | 0f 1f 44 00 00 (nopl 0x0(%rax,%rax,1)) |
5 bytes | — |
The corresponding opcode constants are defined in
arch/x86/include/asm/text-patching.h
(such as JMP8_INSN_SIZE, JMP8_INSN_OPCODE, JMP32_INSN_SIZE,
JMP32_INSN_OPCODE, and INT3_INSN_OPCODE) and
arch/x86/include/asm/nops.h
(including
BYTES_NOP5).
A relative jump specifies a target offset relative to the instruction pointer of the subsequent instruction. Therefore, the relative displacement is measured from the byte immediately following the jump instruction:
displacement = destination - (instruction_address + instruction_size)
The inline helpers
text_gen_insn()
and
__text_gen_insn()
compute this exact offset when formatting the instruction buffer.
Why matching size matters: Swapping a five-byte nop with a
five-byte jmp preserves the exact layout of the surrounding text.
Because no instruction boundaries shift, return addresses stored on the
stack, relative targets of nearby branch instructions, exception table
entries, and ORC unwind metadata remain fully valid and require no
relocation. The patch operates as a strictly localized, in-place byte
substitution.
5.3 Why you cannot just memcpy over live code on SMP
Modifying active kernel instructions on symmetric multiprocessing systems introduces severe architectural challenges that do not exist when writing to standard data structures. These challenges stem from three distinct, concurrent properties of kernel text memory:
- Kernel text memory is mapped read-only after early boot under the STRICT_KERNEL_RWX configuration option in arch/Kconfig. Any direct store operation triggers a page fault, requiring a dedicated mechanism to bypass the write protection.
- Active instructions are fetched concurrently by other execution cores. No global pause freezes the system during a modification, meaning any core can call or continue executing the target function at any instruction boundary.
- Stale instruction bytes might reside within the execution pipeline or instruction cache of another core, having been fetched before the modification but not yet fully decoded or executed.
The primary danger lies in the concurrent instruction fetch described in the second point. Consider a scenario with two cores, Core A and Core B, where Core A attempts to patch a five-byte instruction while Core B executes that same instruction sequence repeatedly. A five-byte memory write is not an atomic operation on modern memory buses. The hardware executes the modification as multiple independent write cycles, each of which becomes visible to other cores at slightly different times. An instruction fetch on Core B can occur precisely in the middle of this multi-step update, reading a mixture of old and new bytes:
time --->
Core A (patcher) [ write bytes 2-4 ] [ write bytes 0-1 ]
^
|
Core B (fetcher) [ fetch all 5 bytes, right here ]
|
v
byte-by-byte: 0=old 1=old 2=new 3=new 4=new
= torn mix: 2 old bytes + 3 new bytes
= neither the old instruction
nor the new one — garbage
When Core A writes five bytes while Core B fetches the instruction, Core B can encounter a torn instruction. This mixture of old and new bytes constitutes a malformed instruction that the hardware instruction decoder cannot safely decode, triggering an invalid opcode exception or unpredictable behavior. The x86 architecture provides no guarantee that multi-byte stores to live, concurrently executing instruction areas are atomic from the perspective of an instruction fetch.
In contrast, a single-byte store is always atomic for instruction fetch operations. A CPU core can never observe a single byte in a partially modified state, as a single byte represents the minimum unit of coherent memory access. Both the Intel Software Developer Manual and the patching implementation in the Linux kernel leverage this atomic behavior to transform an unsafe multi-byte write into three safe, sequential steps:
- Replace the first byte of the instruction site with a single-byte
breakpoint instruction,
INT3(0xCC). Because this single-byte write is atomic, any concurrent execution core either reads the original instruction or hits the breakpoint trap. A torn instruction state is impossible. - Overwrite the remaining bytes of the instruction while the
INT3instruction remains at the entry boundary. No concurrent execution thread can fetch or decode these modified trailing bytes, because any execution attempt immediately traps at the precedingINT3byte. - Replace the
INT3byte with the first byte of the newly prepared instruction using another atomic, single-byte write. This single store marks the exact transition when the new instruction becomes live and executable. - Execute a global synchronization across all cores between each of these steps. This synchronization, driven by an inter-processor interrupt, forces every core to execute a serializing instruction. The serialization flushes stale instruction bytes from the pipelines and instruction caches, preventing any core from executing out-of-date instruction sequences.
This synchronization protocol is implemented in smp_text_poke_batch_finish() within arch/x86/kernel/alternative.c (detailed in §10). Jump labels represent only one client of this multi-step patching mechanism; other core subsystems, including dynamic ftrace, static calls, kprobes, and alternative patching, rely on this identical atomic replacement protocol.
5.4 Writing read-only kernel text: text_poke()
Modern kernels no longer patch instruction text by clearing the
write-protection (WP) bit in the %cr0 control register, performing the
write, and immediately restoring the bit. While this technique was
historically common, it represents a blunt instrument that compromises
system integrity. Between the clearing of the WP bit and the restoration
of the bit, any concurrent write from any CPU core could land on
write-protected memory, rather than only the target instruction
undergoing patching. Any execution thread running during this critical
window could accidentally corrupt kernel memory that should remain
read-only.
To eliminate this risk,
__text_poke()
employs a highly localized approach. Instead of modifying the
permissions of the existing read-only virtual mapping of the .text
section, the kernel establishes a second, transient, and private virtual
mapping that points to the exact same physical page of RAM, performing
the write through this alias instead.
Physical RAM operates independently of virtual memory mappings. The same
physical frame can be mapped simultaneously through multiple virtual
addresses, each carrying distinct page permissions. The standard virtual
mapping of the .text section, which remains visible to all CPU cores
throughout the lifetime of the system, is strictly read-only and
executable. This permanent mapping allows any active core to fetch and
execute instructions at any moment.
During a patching operation, __text_poke() temporarily configures a
second virtual mapping to the same physical page. This alias is marked
writable but not executable, and is restricted solely to the specific
CPU core performing the patching. Tearing down this mapping within a few
instructions minimizes the exposure. This design operates like a second
door into a secure room: the underlying content (the raw instruction
bytes) remains identical regardless of the door used to access it, but
only one of the doors is ever unlocked, and then only for the patching
thread.
Resolving this secondary mapping involves a sequence of safeguards
implemented within the body of __text_poke(). The function manipulates
the memory data to safely configure, switch, write, and tear down the
temporary page(s):
static void *__text_poke(text_poke_f func, void *addr, const void *src, size_t len)
{
bool cross_page_boundary = offset_in_page(addr) + len > PAGE_SIZE;
struct page *pages[2] = {NULL};
struct mm_struct *prev_mm;
unsigned long flags;
pte_t pte, *ptep;
spinlock_t *ptl;
pgprot_t pgprot;
...
In this signature, func represents the target copy routine, resolving
to a memcpy-like function for standard text_poke() calls, or a
memset-like helper for the _set variant. The parameters addr,
src, and len specify the target virtual address, the source payload,
and the copy size.
To determine the physical backing of the memory being patched,
__text_poke() identifies the underlying physical pages. This normally
requires a single page, but can require two pages if the write spans a
page boundary. For core kernel text, the pages are retrieved using
virt_to_page();
for text residing within a dynamic kernel module, the pages are resolved
via
vmalloc_to_page():
if (!core_kernel_text((unsigned long)addr)) {
pages[0] = vmalloc_to_page(addr);
if (cross_page_boundary)
pages[1] = vmalloc_to_page(addr + PAGE_SIZE);
} else {
pages[0] = virt_to_page(addr);
if (cross_page_boundary)
pages[1] = virt_to_page(addr + PAGE_SIZE);
}
The next phase points a pre-allocated page-table entry at the resolved
physical page within a dedicated, otherwise-empty virtual address space
called
text_poke_mm.
This address space is initialized once at boot time by
poking_init().
The temporary entry is configured as writable and explicitly lacks the
_PAGE_GLOBAL
attribute.
Excluding the global bit keeps the overhead of the operation minimal.
Because a non-global mapping is cached only within the translation
lookaside buffer (TLB) of the current CPU core, dismantling the mapping
requires only a local TLB invalidation via
flush_tlb_mm_range().
This avoids the expensive inter-processor interrupts (IPIs) that would
otherwise be required to flush the TLBs of other cores, as no other core
ever loads text_poke_mm:
pgprot = __pgprot(pgprot_val(PAGE_KERNEL) & ~_PAGE_GLOBAL);
ptep = get_locked_pte(text_poke_mm, text_poke_mm_addr, &ptl);
local_irq_save(flags);
pte = mk_pte(pages[0], pgprot);
set_pte_at(text_poke_mm, text_poke_mm_addr, ptep, pte);
if (cross_page_boundary) {
pte = mk_pte(pages[1], pgprot);
set_pte_at(text_poke_mm, text_poke_mm_addr + PAGE_SIZE, ptep + 1, pte);
}
To perform the write, the current CPU core switches onto the private
address space by calling
use_temporary_mm(),
saving the active mm context for later restoration. Writing to the
%cr3 control register alters the active virtual-to-physical
translations on this specific core. However, because modern processors
employ deep, out-of-order execution pipelines, instructions situated
ahead of the %cr3 write might already be fetched, decoded, or
speculatively executed under the previous page translations. If left
uncoordinated, execution could proceed with stale translations,
resolving memory access under the wrong mapping context.
The x86 architecture mitigates this hazard by defining writes to control
registers (including %cr0, %cr3, %cr4, and %dr8) as serializing
instructions. The CPU core must retire all preceding instructions,
discard any speculative instructions in flight, and flush non-global TLB
entries before starting execution under the new register state. This
serialization guarantee is an inherent property of the instruction set
architecture (ISA). Consequently, loading the %cr3 register ensures
that subsequent instructions see the new page-table entry before
fetching memory through it, requiring no further synchronization:
prev_mm = use_temporary_mm(text_poke_mm);
This context switch must also address a secondary hazard involving
hardware breakpoints and watchpoints. The debug registers (%dr0
through %dr3) are global processor state rather than thread-specific
or address-space-scoped entities. Thus, any active watchpoints remain
armed across the address-space switch, regardless of register
serialization.
The target address
text_poke_mm_addr
resides in the lower, user-range half of the address space. If a
user-mode debugger has registered a watchpoint that overlaps with this
numeric address, the processor would trigger a debug exception
mid-write, right in the middle of the code-patching sequence. This would
result in a misdelivered signal to the user process or, worse, interrupt
the critical patching operation.
To prevent this collision, use_temporary_mm() disables hardware
breakpoints immediately after switching the address space by calling
hw_breakpoint_disable().
The counterpart function
hw_breakpoint_restore()
restores the breakpoint state after patching is complete. This temporary
disablement suppresses all breakpoints globally on the current core,
including kernel-space breakpoints registered by tools such as perf.
With the writable alias in place, the core executes the copy by invoking
func at the target address offset:
func((u8 *)text_poke_mm_addr + offset_in_page(addr), src, len);
For standard text_poke() invocations, func corresponds to
text_poke_memcpy(),
which wraps the inline copy with architectural overrides (the _set
variant instead passes
text_poke_memset()):
static void text_poke_memcpy(void *dst, const void *src, size_t len)
{
lass_stac();
__inline_memcpy(dst, src, len);
lass_clac();
}
This design addresses strict architectural checks on user-range accesses and build-time verification rules, which dictate the use of inlined memory operations.
The transient mapping is dismantled in the reverse order of its
creation. Calling
pte_clear()
deletes the page-table entries, while
unuse_temporary_mm()
switches %cr3 back to the original address space (enforcing another
CPU serialization). The local TLB entry is then invalidated using
flush_tlb_mm_range().
For standard text_poke() operations, the kernel validates the patch by
reading back the modified bytes and performing a comparison against the
source buffer. Any discrepancy triggers a
BUG()
panic, preventing the processor from executing corrupt or unintended
instructions:
pte_clear(text_poke_mm, text_poke_mm_addr, ptep);
if (cross_page_boundary)
pte_clear(text_poke_mm, text_poke_mm_addr + PAGE_SIZE, ptep + 1);
unuse_temporary_mm(prev_mm);
flush_tlb_mm_range(text_poke_mm, text_poke_mm_addr, text_poke_mm_addr +
(cross_page_boundary ? 2 : 1) * PAGE_SIZE, PAGE_SHIFT, false);
if (func == text_poke_memcpy)
BUG_ON(memcmp(addr, src, len));
local_irq_restore(flags);
Put together, this is one physical page reached through two different virtual addresses with two different permissions:
physical page (the actual RAM
holding the instruction bytes)
^ ^
| |
normal kernel | | text_poke_mm_addr
mapping (all CPUs, | | (this CPU only,
always present) | | exists briefly)
| \ / |
v \ / v
.text [ RO, executable ] \ / [ RW, not executable ]
(the %cr3 of each CPU \/ (only the %cr3 of this CPU
maps this address, maps this address,
forever) only while patching)
The permanent mapping of .text remains read-only across all CPU cores.
Only the active patching core gains transient, local access to the
writable alias, and this access is restricted to the duration of
__text_poke(). Once the page-table entry is cleared and the local TLB
is flushed, the writable alias is completely removed, leaving only the
updated read-only .text mapping. The entire text poking process is
serialized by
text_mutex,
ensuring that concurrent threads never race to establish competing
temporary mappings.
Hardware details — why SMAP, LASS, and inline copying dictate the implementation. Modern processors implement security mechanisms designed to prevent the kernel from accessing user-range virtual addresses by accident. These features guard against kernel vulnerabilities where a corrupted or attacker-controlled pointer is dereferenced within supervisor mode. Two hardware-level protections enforce these boundaries using distinct criteria:
- SMAP (“Supervisor Mode Access Prevention”) generates a page fault if kernel code attempts to access a virtual page whose page-table entry has the
_PAGE_USERbit set. This mechanism evaluates only the user permission bit of the translation entry, ignoring the numeric value of the virtual address.- LASS (“Linear Address Space Separation”) blocks supervisor-mode accesses to any virtual address falling under the user address space. This protection relies entirely on the numeric range of the address, regardless of whether
_PAGE_USERis configured on the page.The address
text_poke_mm_addris allocated within the lower half of the virtual address space, corresponding to the range where user processes receive mappings frommmap(). The virtual address is allocated viamm_alloc()at boot, which initializes the structure atTASK_UNMAPPED_BASEwith a randomized offset. However, because the page-table entry built for this mapping has the_PAGE_USERbit cleared, the page is not user-accessible.This configuration interacts differently with each protection mechanism. Because the
_PAGE_USERbit remains clear, SMAP does not flag the access. However, because the virtual address numerically resides in the user-space range, LASS would trigger an immediate supervisor-mode page fault.To bypass this restriction, the kernel invokes
lass_stac()andlass_clac()around the copy. These functions check for the presence ofX86_FEATURE_LASSon the processor; if enabled, they temporarily toggle the alignment check (AC) flag in the%rflagsregister, instructing the hardware to permit the access. On processors supporting onlyX86_FEATURE_SMAP, or neither feature, these operations resolve to no-ops or default behavior. This is conceptually identical to the user-access window established bycopy_from_user().Enabling the AC override imposes a critical constraint on the code. The kernel build-time analysis tool, objtool, enforces a strict rule prohibiting any instruction calling another function between a STAC and a CLAC instruction. The AC flag is an active CPU register state but is not automatically saved or restored during a task context switch. If the code inside the override window executes a function call that eventually yields the CPU via
schedule(), the AC flag remains set. This would leak the user-access permission into the next scheduled task, compromising system security.This build-time rule is the reason the patching copy cannot utilize the standard library implementation of
memcpy(). On x86_64,memcpy()is an assembly routine defined in a separate object file, necessitating a function call. To adhere to the objtool constraint, the patching sequence employs__inline_memcpy()and__inline_memset(), which compile directly into inlinerep movsbandrep stosbassembly instructions. This removes the function call entirely, satisfying the safety validation of objtool.
6 What the compiler emits (x86_64)
Every static branch call site contains either a nop or a jmp
instruction of identical length. Along with this inline instruction, the
compilation process generates a corresponding metadata entry in the
__jump_table
section. The choice between a nop and a jmp at compile time depends
on the initial state of the static key and the branch hint at the call
site. Two architecture-specific macros generate these instructions and
populate the sidecar table.
6.1 TRUE/FALSE keys and likely/unlikely sites
The compiler must emit a concrete, valid machine instruction at every
call site before
jump_label_init()
runs at boot. Because compilation occurs statically, the compiler cannot
evaluate the live
enabled
reference count of a key. Instead, the generated instruction is
determined by two compile-time variables:
- The default key state — Declared via
DEFINE_STATIC_KEY_TRUEorDEFINE_STATIC_KEY_FALSE. - The call-site direction hint — Specified using
static_branch_likely()orstatic_branch_unlikely().
When the default state of the key aligns with the call-site direction
hint, the compiler emits a nop instruction. When the default state and
the direction hint mismatch, the compiler emits a jmp instruction.
Disaligning these variables in relation to the expected steady state of
the static key—the hazard described in
§4.2—causes the hot
path to start with a branch jump, which persists until an explicit
toggle occurs at runtime.
The reference documentation at the top of
include/linux/jump_label.h
maps out this structural matrix:
likely() unlikely()
-------- ----------
key=true ... ...
NOP JMP L
<br-stmts> 1: ...
L: ...
L: <br-stmts>
jmp 1b
------------------------------------------------------------
key=false ... ...
JMP L NOP
<br-stmts> 1: ...
L: ...
L: <br-stmts>
jmp 1b
The runtime text patcher calculates the destination state using the same
Boolean logic, substituting the live state of the key for the
compile-time default state. The inline function
jump_label_type()
implements this evaluation at runtime. An identical function,
jump_label_init_type(),
performs the calculation during initialization by reading the static,
compile-time key type bit. The logical branch state evaluates to 1 for
a likely site and 0 for an unlikely site. The resulting
compile-time XOR operation translates to type ^ branch. The live
implementation evaluates the active state of the branch:
static enum jump_label_type jump_label_type(struct jump_entry *entry)
{
struct static_key *key = jump_entry_key(entry);
bool enabled = static_key_enabled(key);
bool branch = jump_entry_is_branch(entry);
return enabled ^ branch;
}
6.2 How the macros pick the asm
The high-level macros
static_branch_likely()
and
static_branch_unlikely()
do not generate or manipulate instruction bytes directly. Instead, they
delegate to two architecture-specific helper functions:
arch_static_branch()operates as the nop-by-default helper. The generated patch site initially contains a standardnopinstruction. (As detailed in the optimization mechanism of §6.5, the compiler may emit ajmpthatobjtoolrewrites into a same-sizednopbefore boot).arch_static_branch_jump()operates as the jmp-by-default helper. It places a concrete branch jump instruction at the call site when the object file is generated.
Both helpers communicate the path that execution took through the
patched instruction. A nop instruction falls through to the next
sequential instruction, causing the helper to return false. A jmp
instruction transfers execution out-of-line, causing the helper to
return true. The implementation details of the underlying asm goto
statement are examined in §6.3.
The Boolean return value reflects the active instruction shape at the patch site. The logical association with the branch path of the caller is resolved through the negation logic implemented in the macro definitions.
The four possible combinations of key type and call-site hint map onto these two low-level helpers. The helper selection is determined by the compile-time type of the static key, while the optional negation of the return value is dictated by the direction hint:
/* CONFIG_JUMP_LABEL path in jump_label.h */
static_branch_likely(x):
TRUE key → !arch_static_branch(&(x)->key, true) /* nop-default site */
FALSE key → !arch_static_branch_jump(&(x)->key, true) /* jmp-default site */
static_branch_unlikely(x):
TRUE key → arch_static_branch_jump(&(x)->key, false) /* jmp-default site */
FALSE key → arch_static_branch(&(x)->key, false) /* nop-default site */
Here, arch_static_branch() is invoked for combinations that compile to
a nop, whereas arch_static_branch_jump() is utilized for
combinations that compile to a jmp.
The unary negation operator (!) preceding the two likely branches
maps the helper output to the expected conditional behavior of the
caller. The helpers return true or false based on whether a branch
jump occurred, rather than whether the static key is enabled.
For example, reading a TRUE key with likely invokes the nop-default
helper. A nop instruction falls through, so arch_static_branch()
returns false (indicating that no branch jump occurred), even though
the conditional body must execute. Applying the unary negation operator
converts this result into true, ensuring correct control flow.
The Boolean argument passed to the helper is stored directly inside the metadata of the jump table, as detailed in §6.4. The runtime text-patching engine evaluates this metadata whenever the state of the key is toggled.
The compile-time type selection uses
__builtin_types_compatible_p
to distinguish between
struct static_key_true
and
struct static_key_false.
Any unsupported type defaults to an unresolved call to
____wrong_branch_error(),
which triggers a compilation failure.
6.3 The two asm helpers
While the macro expansion treats
arch_static_branch()
and
arch_static_branch_jump()
as opaque functions returning a Boolean value, their underlying
implementations in
arch/x86/include/asm/jump_label.h
expose inline assembly:
static __always_inline bool arch_static_branch(struct static_key * const key,
const bool branch)
{
asm goto(ARCH_STATIC_BRANCH_ASM("%c0 + %c1", "%l[l_yes]")
: : "i" (key), "i" (branch) : : l_yes);
return false;
l_yes:
return true;
}
static __always_inline bool arch_static_branch_jump(struct static_key * const key,
const bool branch)
{
asm goto("1:"
"jmp %l[l_yes]\n\t"
JUMP_TABLE_ENTRY("%c0 + %c1", "%l[l_yes]")
: : "i" (key), "i" (branch) : : l_yes);
return false;
l_yes:
return true;
}
The use of two separate inline functions, rather than a single function
parameterized with a branch selection flag—such as a speculative
arch_static_branch(key, branch, use_jmp)—is dictated by compile-time
constraints. Because the assembly payload is evaluated during
compilation, a runtime argument cannot dynamically alter the instruction
bytes embedded within the function body. The compiled instructions are
finalized when the translation unit is processed.
Consequently, separate functions encapsulate each instruction layout, with the logical branch path selected by the macros described in §6.2. This selection is resolved statically based on the compile-time type of the static key.
The function bodies rely on inline assembly, specifically the GCC
asm goto extension. The functional components of the asm goto blocks
provide several key mechanisms:
asm goto(template : : inputs : : goto-labels)— This extension to inline assembly allows control flow to branch from the assembly block directly to a C label. Unlike a standard inline assembly block, which falls through to the subsequent C statement,asm gotoregisters a list of destination C labels. Falling off the end of the instruction template acts as a logical fall-through, executingreturn false;. Jumping to the designated assembly label transfers control directly to the C block marked byl_yes, executingreturn true;. This mechanism allows the compiler to treat the assembly block as a standard conditional branch.%0,%1, … — These placeholders represent the operands listed after the colons. The modifiers%c0and%c1instruct the compiler to format the first operand (key) and the second operand (branch) as bare constants. This format strips decoration characters, such as the$prefix used for immediate values in x86 syntax. This representation is made possible by specifying the"i"constraint, which is detailed at the end of this subsection.%l[l_yes]— This operand specifier outputs the assembler-generated label corresponding to the target C labell_yes, providing a direct branch target for jump instructions.- Numeric local labels — Temporary labels, such as
1:, are reusable local identifiers. The suffixfrefers to the next occurrence of the label forward in the file, whilebrefers to the nearest occurrence backward. This local labeling allows macros likeJUMP_TABLE_ENTRYto expand repeatedly without generating naming conflicts. .(current location counter) — This symbol represents the memory address of the byte currently being emitted by the assembler. Expressions such as1b - .compute the relative distance between a target label and the current instruction, which allows offsets to be calculated statically.- Assembler directives — Statements beginning with a dot, such as
.long,.byte,.quad,.pushsection,.popsection, or.balign, instruct the assembler to organize data or format sections rather than generating CPU instructions. Comments are marked with the#character.
With these syntax rules defined, the structure of
arch_static_branch_jump() is highly transparent, passing the following
template to asm goto:
1:
jmp %l[l_yes]
<jump table entry for this site, via JUMP_TABLE_ENTRY>
1:establishes the local label pointing to the start of the branch instruction, satisfying thecodefield requirement described in §6.4.jmp %l[l_yes]represents an unconditional branch jump to the target address associated with the C labell_yes. This forms the pre-patched compiled branch.JUMP_TABLE_ENTRY(\"%c0 + %c1\",\"%l[l_yes]\")expands inline, embedding the metadata record for the call site. The directives.pushsection __jump_table, \"aw\"and.popsectiontemporarily redirect the assembler output to the custom__jump_tablesection, separating the metadata from the executable.textsegment.
The companion function arch_static_branch() is constructed
analogously, but delegates the generation of the local label and
instruction to the
ARCH_STATIC_BRANCH_ASM
macro:
#ifdef CONFIG_HAVE_JUMP_LABEL_HACK
#define ARCH_STATIC_BRANCH_ASM(key, label) \
"1: jmp " label " # `objtool` NOPs this \n\t" \
JUMP_TABLE_ENTRY(key " + 2", label)
#else /* !CONFIG_HAVE_JUMP_LABEL_HACK */
#define ARCH_STATIC_BRANCH_ASM(key, label) \
"1: .byte " __stringify(BYTES_NOP5) "\n\t" \
JUMP_TABLE_ENTRY(key, label)
#endif /* CONFIG_HAVE_JUMP_LABEL_HACK */
This preprocessor conditional is resolved at build time based on
HAVE_JUMP_LABEL_HACK.
Each compilation path outputs different assembly representations:
- With the hack enabled — The instruction layout matches
arch_static_branch_jump()but appends a descriptive comment for readability. The expression" + 2"performs assembler-level arithmetic on the key address, setting bit 1 as a metadata flag forobjtool. This flag indicates that the site should be processed during build-time optimization, as detailed in §6.5 and §6.6. - With the hack disabled — The instruction line is generated using
raw bytes:
1: .byte 0x0f,0x1f,0x44,0x00,0x00. The preprocessor converts the macroBYTES_NOP5into a comma-separated byte list. This sequence decodes as a valid 5-bytenopinstruction (nopl 0x0(%rax,%rax,1)), providing a safe, inactive default path.
The operand declaration : : \"i\" (key), \"i\" (branch) : : l_yes);
binds the C variables to the assembly template. Defining the inputs with
the "i" constraint forces the compiler to resolve these parameters as
compile-time constants. This constraint ensures that the metadata values
remain static and available during the assembly phase, even after the
compiler has inline-expanded the containing functions. The clobber list
is empty, and the goto-label list specifies the branch target l_yes.
6.4 The jump table entry (sidecar metadata)
Along with the inline instruction, each branch site emits a metadata
descriptor of type
struct jump_entry
into the
__jump_table
section. The
JUMP_TABLE_ENTRY
macro generates this metadata without outputting any executable CPU
instructions:
#define JUMP_TABLE_ENTRY(key, label) \
".pushsection __jump_table, \"aw\" \n\t" \
_ASM_ALIGN "\n\t" \
ANNOTATE_DATA_SPECIAL "\n" \
".long 1b - . \n\t" \
".long " label " - . \n\t" \
_ASM_PTR " " key " - . \n\t" \
".popsection \n\t"
| Field | Asm | Meaning |
|---|---|---|
code |
.long 1b - . |
relative offset to the patchable insn |
target |
.long label - . |
relative offset to the l_yes target |
key |
_ASM_PTR key - . |
relative offset to the static_key, low bits = flags |
Every directive in this macro instructs the assembler to format data rather than producing executable machine instructions. Each directive is evaluated by the assembler according to specific rules:
.pushsection __jump_table, "aw"— This switches the active assembly target section to the designated jump table. The flag string"aw"configures the properties of the section:adesignates the section as allocatable (loading it into memory at boot alongside code and data), whilewmarks it as writable. Setting the writable property is critical because the page tables of the CPU enforce write permissions at runtime. This configuration enables the patching engine to modify the section contents post-boot._ASM_ALIGN— This aligns the subsequent data, expanding to.balign 8on x86_64 and.balign 4on 32-bit platforms. This alignment ensures that the native-width pointer-sized fields are aligned to natural boundaries, preventing unaligned memory accesses.ANNOTATE_DATA_SPECIAL— This assembler macro emits metadata to a build-time section analyzed exclusively by theobjtoolverification engine. It signals that the succeeding bytes represent static data rather than executable instructions. This preventsobjtoolfrom attempting to parse the offsets as machine code..long 1b - .— This directive reserves 4 bytes of data initialized to the value of the relative expression1b - .. The operand1bresolves to the address of the nearest preceding local label1:(representing the patchable branch site), while the.symbol resolves to the address of the directive itself. The assembler computes this subtraction at build time, producing a signed relative distance in bytes..long " label " - ."— This repeats the relative distance calculation for the target label, which corresponds to the address of the target C labell_yes. This offset represents the relative distance from the metadata field to the branch target._ASM_PTR " " key " - ."— This reserves a pointer-sized field, expanding to.quad(8 bytes) on x86_64. The symbolkeyresolves to the expression"%c0 + %c1". At assembly time, the expression computes the relative offset from the metadata field to the address of the corresponding static key. The addition of the branch hint (%c1) effectively ORs the logical branch bit into bit 0 of the stored key address, which is retrieved at runtime byjump_entry_is_branch()..popsection— This directive terminates the active section redirect, restoring the assembly target back to the original section (typically.text).
The fields code and target utilize 32-bit displacements via .long,
whereas key requires the full pointer width of _ASM_PTR. Because the
patchable instruction and the target block reside within the same
function body, a 32-bit offset is guaranteed to reach the target.
Conversely, the target
struct static_key
may be located far from the call site under KASLR or within a separate
kernel module, requiring a full-width relocation.
Each field is encoded as a self-relative offset, meaning the distance is measured from the address of the metadata field itself rather than from the beginning of the structure or the array. This design allows each offset to be resolved using a uniform address calculation.
For the code field:
__jump_table[i].code lives at a memory address designated as F.
The patchable nop/jmp instruction in .text resides at address C.
The value written into the code field at assembly time represents the distance:
stored value = C - F
At runtime, the address of the instruction C is resolved by referencing
the address of the field F:
C = &entry->code + entry->code
^^^^^^^^^^^^ ^^^^^^^^^^^^
own address the distance
of this field stored in it
This addition represents the implementation of
jump_entry_code().
The fields for .target (resolving the branch destination) and .key
(resolving the target key) are computed using the same mechanism,
evaluating the offsets relative to their own field addresses. Thus,
jump_entry_target()
evaluates as &entry->target + entry->target, while
jump_entry_key()
masks out the flag bits and computes the address relative to
&entry->key. Measurement relative to individual fields eliminates the
need for field-specific offsets in the retrieval logic.
This relative layout also ensures that the metadata survives KASLR
relocations with no boot-time overhead. If the kernel image is shifted
by a constant offset at boot, both the field address F and the target
address C scale by the identical offset. The relative distance C - F
remains invariant, which allows
jump_label_init()
to read the jump table immediately without performing pointer
relocation.
6.5 HAVE_JUMP_LABEL_HACK: why sites are 2 or 5 bytes
The nop-default site of
arch_static_branch()
is defined using a hand-coded .byte BYTES_NOP5 (representing a fixed
5-byte NOP) under certain configurations. However, this is not the
instruction size that is guaranteed to land in a compiled kernel image.
Depending on the properties of the call site, the NOP that resides in
memory is either a compact 2-byte instruction or a full 5-byte
instruction. The resolution of this size is settled before boot: the
compilation process outputs a real jmp instruction, allowing the
compiler to select the shortest valid encoding, and a subsequent build
step converts this jmp into a NOP of identical width before the kernel
image is finalized.
This optimization mechanism is gated by the HAVE_JUMP_LABEL_HACK
configuration option. The architecture configuration file
arch/x86/Kconfig
enables this option for any build where objtool is available via
select HAVE_JUMP_LABEL_HACK if HAVE_OBJTOOL (which is satisfied on all
modern x86_64 builds). The preprocessor directive
ARCH_STATIC_BRANCH_ASM
resolves this compile-time branch.
This optimization solves a compiler limitation: generating a hard-coded
.byte BYTES_NOP5 always forces a 5-byte NOP, even when a 2-byte
instruction is sufficient, wasting instruction cache space at every
nop-default call site. The alternative is to let the compiler emit a
correctly sized jmp instruction, as the compiler can evaluate the
shortest encoding needed to reach the destination target. The build
process then converts this jmp into a same-sized NOP.
With the optimization hack enabled:
- The compiler emits a real
jmptol_yesas it would for an ordinary conditional branch. The assembler selects between two unconditional jump encodings based on the displacement distance.JMP rel8uses a 1-byte opcode (0xEB) and a signed 1-byte relative displacement (2 bytes total), which is valid if the branch target resides within -128 to +127 bytes of the next instruction.JMP rel32uses a 1-byte opcode (0xE9) and a 4-byte signed displacement (5 bytes total), which can reach any location within a 32-bit relative displacement.1 The terms “rel8” and “rel32” specify the bit width of the displacement field. The assembler calculates the real distance to the label and selects the 2-byterel8format when the target is close, falling back to the 5-byterel32format only when necessary. This selection ensures that the instruction size is optimal for each call site. - The key expression passed to the jump table macro is defined as
"%c0 + %c1 + 2", which sets bit 1 of the stored key address. This bit acts as a metadata marker that is processed during the subsequent build stage rather than being evaluated at runtime. - objtool
(
handle_jump_alt()) processes the compiled object files during the build phase. When it detects that bit 1 of the key address is set, it overwrites the correspondingjmpinstruction with a same-sized NOP and clears the associated relocation entry. This transformation is completed entirely at build time. When the kernel boots, the call sites already contain valid NOP instructions. Once this transformation is complete,jump_label_init()repurposes bit 1 of the key address to store the init-text flag viajump_entry_set_init(), as the build-time marker is no longer required.
By the time the kernel image is finalized, every nop-default call site
contains a valid NOP instruction optimized to the smallest possible
width: 2 bytes for sites where the original jmp utilized the rel8
format, or 5 bytes where it utilized rel32. The instruction size is
fixed during the build phase and inherited at boot.
However, the
struct jump_entry
metadata descriptor does not record the instruction size of each site.
Storing this size is impractical because the instruction width is
finalized during the objtool pass, which occurs after the compiler has
laid out the structure fields.
Consequently, when the patching engine toggles a site at runtime—for
example, during a call to
static_branch_enable()
or
static_branch_disable()—the
patching logic must examine the instruction bytes in memory to
dynamically resolve whether a 2-byte or 5-byte instruction is present.
It then generates a replacement instruction of identical width to ensure
that the surrounding instruction stream remains aligned. This decoding
and patching pipeline—implemented via
arch_jump_entry_size(),
insn_decode_kernel(),
and
__jump_label_patch()—is
detailed in §8. The
critical detail is that the instruction size established by objtool is
discovered and preserved during patching, preventing code relocation.
In configurations where the optimization hack is disabled (such as older
toolchains lacking objtool integration), the compiler bypasses this
pipeline. It compiles the #else branch of the assembly template,
emitting a fixed 5-byte NOP via .byte BYTES_NOP5 regardless of target
proximity. Later runtime patching always replaces this NOP with a 5-byte
jump instruction. While functional, this fallback increases the
instruction footprint by 3 bytes per nop-default call site compared to
optimized builds.
The build-time optimization only applies to nop-default branches
generated via the arch_static_branch() helper. The jmp-default helper,
arch_static_branch_jump(),
generates its asm goto directly, passing the key expression to
JUMP_TABLE_ENTRY
as a plain \"%c0 + %c1\" without the additional offset of 2.
Without this metadata flag, objtool does not modify the jump
instruction, which reaches boot as a real conditional jump that is
patched only when the key state is disabled. This instruction is still
compiled using either the 2-byte rel8 or 5-byte rel32 layout
depending on assembler-level distance calculation. The optimization hack
alters whether objtool transforms the instruction post-compilation,
rather than how the initial jump instruction is encoded by the
assembler.
Bit 1 of the key field—the metadata segment masked off by
jump_entry_key()
to retrieve the underlying
struct static_key
pointer—serves two independent purposes across the build and boot
boundary:
build phase --------------------------------> boot phase ---> execution
bit 1 = "objtool: NOP this jmp" jump_label_init()
(processed and discarded by overwrites the bit to represent:
objtool at build time) bit 1 = "jump_entry_is_init"
(marks the site as init-only text)
The boot-time flag is initialized using jump_entry_set_init() during
the execution of jump_label_init(), and is subsequently read using
jump_entry_is_init().
This dual usage is a common source of confusion when analyzing the
initialization sequence.
6.6 Assembly-level picture
The integration of the C helper functions, the sidecar jump-table
metadata, and the build-time instruction width optimization results in a
cohesive machine-level layout. For a nop-default call site on a build
configured with
HAVE_JUMP_LABEL_HACK
(the standard configuration on x86_64), the compiled object file
contains the following representation:
.text:
1: 0f 1f 44 00 00 ; 5-byte NOP if l_yes was far, OR 66 90 (2-byte) if
... ; close (objtool rewrote a real jmp into this
; at build time)
__jump_table: ; non-executable metadata
.long 1b - . ; code: self-relative offset to the NOP instruction
.long L - . ; target: self-relative offset to l_yes
.quad key+branch+2 - .; key: static_key address, branch bit 0,
; plus the objtool-only "+2" signal
The two potential byte sequences representing the inactive state of the
instruction at the 1: label represent distinct machine instructions
rather than arbitrary padding. The 5-byte sequence 0f 1f 44 00 00
represents the multi-byte NOP (nopl 0x0(%rax,%rax,1)), which
incorporates an unused addressing mode to achieve a 5-byte instruction
width. The 2-byte alternative 66 90 incorporates the operand-size
override prefix (0x66) stacked in front of the classic single-byte NOP
opcode (0x90, historically executing as xchg %ax,%ax). The prefix is
added to pad the instruction to exactly 2 bytes without modifying
execution behavior. Both instructions function as genuine no-ops: the
CPU decodes the instruction, consumes execution cycles, and falls
through to the next sequential instruction, leaving all architectural
registers and flags unmodified.
Two specific aspects of this layout merit close examination:
- The width of the NOP instruction (2 or 5 bytes) is not decided
dynamically by the compiler during initial code generation. Without
the optimization hack,
ARCH_STATIC_BRANCH_ASMalways outputs a hard-coded 5-byte NOP. With the hack enabled, the compiler outputs a standardjmpinstruction, which the assembler encodes as either a 2-byte or 5-byte branch based on the displacement distance tol_yes. The build-timeobjtoolverification pass subsequently converts this jump instruction into a NOP of identical width. This instruction width is finalized in the compiled object file and remains unchanged until the branch is toggled at runtime. - The
+2offset does not represent a runtime control state. It functions exclusively as the build-time instruction conversion signal, residing in bit 1 of the stored key address (while the branch hint occupies bit 0). This marker is processed and discarded during the build phase;jump_label_init()overwrites this bit position with the init-text flag during boot initialization.
The x86 architecture also defines
HAVE_JUMP_LABEL_BATCH,
which enables the batched and synchronized text-patching path. Without
this batched optimization, every individual jump-table entry would
require separate instruction patching and global CPU synchronization,
introducing significant performance overhead.
7 Core data structures
Runtime management of dynamic patching requires cooperative interaction between active state representation and static compiler metadata. The core subsystem models this relationship through unified control structures and metadata records that catalog every patching target across the system. These components link individual branch sites back to the central keys that govern them.
7.1 struct static_key
Every static key declared via
DEFINE_STATIC_KEY_TRUE
or
DEFINE_STATIC_KEY_FALSE
boils down to a single runtime representation defined by
struct static_key. Compile-time type wrappers maintain logical
distinction during compilation, but they resolve to this identical
underlying structure at runtime.
struct static_key {
atomic_t enabled;
#ifdef CONFIG_JUMP_LABEL
union {
unsigned long type;
struct jump_entry *entries;
struct static_key_mod *next;
};
#endif
};
The
enabled
field represents the active control state of the key. It is declared as
an
atomic_t
reference counter where a value of zero indicates the disabled state and
any positive value indicates the enabled state. This design allows
static_branch_inc
and
static_branch_dec
to coordinate multiple stacked owners on a single key. During
transitions, the field can temporarily hold the value -1, indicating
that an initial call to
static_key_slow_inc
is actively patching the instruction stream. To prevent confusing
readers during this transition window,
static_key_count
maps -1 back to 1 before returning.
The second member of the struct is a tagged union residing in a single machine word. The low two bits of the word act as metadata flags. This optimization is possible because pointers to struct jump_entry and struct static_key_mod are at least 4-byte aligned on supported architectures. The alignment guarantees that the two least significant bits of any valid pointer address are zero, leaving them available for bitwise tagging.
| Bit | Macro | Meaning |
|---|---|---|
| 0 | JUMP_TYPE_TRUE |
compile-time initial value of the key was true — this is the same type bit the type ^ branch formula from §6.1 uses |
| 1 | JUMP_TYPE_LINKED |
1: the rest of the word is a next pointer (a linked list); 0: it is an entries pointer (a flat array) |
This bit-packing strategy must not be confused with the tagging applied
to the key field of struct jump_entry. While both employ similar
bitwise operations on pointers, they serve different purposes. One tags
the initial state of the key and the structure of the associated
call-site list, while the other indicates the branch hint and
init-section status of a specific call site.
Bit 1 handles cases where call sites are distributed across separate
compilation boundaries. For keys utilized solely within the core kernel
image, the linker arranges all associated call sites into a single
contiguous array inside the __jump_table section of vmlinux. The
entries pointer then points directly to the start of this sequence.
However, kernel modules loaded at runtime carry private __jump_table
sections that cannot be merged post-link. Since modules are loaded and
unloaded dynamically, the active set of call sites must grow and shrink.
When a module introduces new call sites for an existing key, the
representation transitions. The word shifts from a direct pointer to the
head of a linked list composed of struct static_key_mod nodes, where
each node tracks the contribution of a specific module.
struct static_key_mod {
struct static_key_mod *next;
struct jump_entry *entries;
struct module *mod;
};
Subsystems outside the core jump label implementation in kernel/jump_label.c do not interact with these raw bits directly. Accessor helpers like static_key_entries, static_key_type, static_key_linked, and static_key_set_entries mask off the tag bits before returning pointers, insulating the rest of the kernel from the underlying representation.
The type wrappers struct
static_key_true
and struct
static_key_false
allow the compiler to distinguish key polarities at compile time. Each
contains a single struct static_key member:
struct static_key_true { struct static_key key; };
struct static_key_false { struct static_key key; };
These distinct C types allow the preprocessor and compiler to select the appropriate branch behaviors. The macros static_branch_likely and static_branch_unlikely use compiler builtins to inspect the type of the passed key and dispatch execution to the correct architecture-specific code generation path.
7.2 struct jump_entry (relative form)
This structure acts as the metadata record generated for each jump site.
The
JUMP_TABLE_ENTRY
assembler macro writes one instance of this record per call site into
the __jump_table section. The x86 architecture opts into a relative
variant of this metadata structure by selecting the configuration option
HAVE_ARCH_JUMP_LABEL_RELATIVE.
This configuration alters the layout of struct jump_entry:
struct jump_entry {
s32 code;
s32 target;
long key; /* full width: module↔vmlinux may be far under KASLR */
};
The code and target fields contain self-relative offset distances
rather than absolute pointer addresses. Reconstructing the absolute
address of the patchable instruction is achieved by adding the value of
code to the address of the code field itself. The target destination
address is reconstructed similarly. The key field, after masking off
the two least significant bits, resolves to the address of the governing
static_key using the same self-relative offset arithmetic.
Architectures that do not select HAVE_ARCH_JUMP_LABEL_RELATIVE employ
a standard fallback structure where the fields store absolute addresses.
This distinction is visible in the accessors of the fallback structure
(where
jump_entry_code
simply performs a direct pointer return with no arithmetic). On x86,
paying the minor computational cost of the addition on every lookup
yields significant advantages:
- Size Optimization: A running kernel contains tens of thousands
of individual branch sites. Storing offsets as 32-bit signed
integers (
s32) instead of full-width pointer values (8 bytes on x86_64) halves the size of two-thirds of each entry, substantially reducing the memory footprint of the total table. - KASLR Compatibility: Because both the patchable code site and
the target label reside within the boundaries of a single compiled
function, a 32-bit signed offset is always sufficient to span the
distance. The absolute address is resolved through relocation-free
arithmetic, avoiding boot-time relocation fixups under Kernel
Address Space Layout Randomization (KASLR). In contrast, the
keyfield is kept at full pointer width (longinstead ofs32) because it references astatic_keywhich can reside anywhere in the address space — including inside a different kernel module or the mainvmlinuxbinary. Because the distance between a dynamically loaded module and the core kernel image can exceed the range of a 32-bit signed integer, this field must maintain full pointer width.
The two least significant bits of the key pointer are reserved for
encoding additional metadata. This pointer tagging strategy conveys two
distinct pieces of information:
| Bit | Meaning |
|---|---|
| 0 | jump_entry_is_branch: Represents the branch direction hint. A value of 1 indicates that the call site uses likely(), while 0 indicates unlikely(). |
| 1 | jump_entry_is_init: Indicates whether the target instruction resides within an initialization section (__init text). Such sections are freed after boot, rendering the associated call sites unpatchable from that point forward. |
The inline accessors perform direct bitwise checks to retrieve these flags:
static inline bool jump_entry_is_branch(const struct jump_entry *entry)
{
return (unsigned long)entry->key & 1UL;
}
static inline bool jump_entry_is_init(const struct jump_entry *entry)
{
return (unsigned long)entry->key & 2UL;
}
Bit 1 is the same marker tracked through two distinct lifetimes. During
the initial build phase, it signals to objtool that a given jump
instruction must be converted to a NOP. After
jump_label_init
processes this instruction and completes boot-time setup, the bit
assumes the permanent meaning as the initialization-text flag for the
remainder of the kernel uptime.
7.3 Relationship
A standard struct
static_key
holds a direct entries pointer (the common, non-module-linked case)
referencing the associated call sites in memory:
struct static_key
+------------------------+
| enabled (atomic) |
| type / entries / next | tagged union
+-----------+------------+
|
| first jump_entry for this key
v
__jump_table[] (sorted by key, then by code address)
[ entries for key A ... ][ entries for key B ... ] ...
| |
+--> code / target / key <--+
The struct static_key structure does not store an explicit count of
referencing call sites; it retains only a pointer to the first
associated struct
jump_entry.
This single pointer is sufficient because of the ordering established in
the metadata array. The function
jump_label_sort_entries
sorts the whole __jump_table array. Sorting is performed not on the
raw bits stored in the key field, but on the decoded, absolute target
address returned by
jump_entry_key.
This decoding is necessary because, like the code and target fields,
the key field is self-relative. Rather than storing a direct absolute
address, it holds the distance to the target key, measured from the own
address of the field within the table entry. The inline helper function
jump_entry_key masks off the two low-order metadata flags to isolate
the offset, then adds the address of the key field itself to
reconstruct the absolute address of the governing key:
static inline struct static_key *jump_entry_key(const struct jump_entry *entry)
{
long offset = entry->key & ~3L;
return (struct static_key *)((unsigned long)&entry->key + offset);
}
Consequently, two distinct table entries referring to the same central
key but residing at different offsets within the table will measure
distances from different starting points. Because the base address of
each field differs per slot, the raw bit patterns stored in the key
fields will differ even though they resolve to the same underlying
control structure:
addr of raw key decode: addr + raw key
entry->key field (jump_entry_key())
---------- -------- -----------------------
slot 0 (A): 0x1000 +0x4000 0x1000 + 0x4000 = 0x5000 ─┐
... ├─ same
slot 5 (B): 0x2000 +0x3000 0x2000 + 0x3000 = 0x5000 ─┘ static_key!
static_key @ 0x5000
While the raw values 0x4000 and 0x3000 share no common bit pattern,
factoring in the absolute address of each entry yields the identical
result 0x5000. The sorting function uses this absolute address to
order the table.
Once sorted, all entries associated with a given key occupy a contiguous
sequence in memory. Locating every call site for a specific key requires
starting at the address specified by the entries pointer of the key
and scanning forward until the decoded key address changes. This layout
eliminates the need to maintain an explicit count or index of associated
call sites.
Within each contiguous key run, a secondary sort is performed using jump_entry_code to arrange the patchable instructions in ascending order of memory addresses. The instruction patching system requires this monotonic ordering to optimize batch patching operations and ensure predictable execution flows.
Sorting relative-offset metadata requires specialized swap logic. A standard sorting algorithm swaps array elements through a byte-for-byte copy. However, because the fields of a relative entry encode distances computed relative to the own address of the entry, moving the record to a different slot without adjusting the fields would corrupt the pointers.
The helper function
jump_label_swap
prevents this corruption. During a swap, the function calculates the
distance delta between the source and destination slots. It then
adjusts the offsets in each field by adding or subtracting delta so
that each relative field, now residing at a new address, continues to
resolve to the same absolute target:
static void jump_label_swap(void *a, void *b, int size)
{
long delta = (unsigned long)a - (unsigned long)b;
struct jump_entry *jea = a;
struct jump_entry *jeb = b;
struct jump_entry tmp = *jea;
jea->code = jeb->code - delta;
jea->target = jeb->target - delta;
jea->key = jeb->key - delta;
jeb->code = tmp.code + delta;
jeb->target = tmp.target + delta;
jeb->key = tmp.key + delta;
}
Architectures that do not use the relative form of struct jump_entry
are immune to this issue. Because absolute addresses remain valid
regardless of where the containing entry resides, those architectures
can rely on standard byte-for-byte swaps.
7.4 Linker section
Each translation unit expanding the
JUMP_TABLE_ENTRY
assembler macro emits a dedicated .pushsection __jump_table …
.popsection block. These metadata fragments are scattered across
numerous compiled object files during compilation. To form a cohesive,
contiguous table, the linker collects these fragments during the final
link phase of the kernel image. The core kernel linker script directs
this collation using the macro
BOUNDED_SECTION_BY
defined in
include/asm-generic/vmlinux.lds.h:
BOUNDED_SECTION_BY(__jump_table, ___jump_table)
This macro expands to three linker directives:
__start___jump_table = .;
KEEP(*(__jump_table))
__stop___jump_table = .;
The directive *(__jump_table) instructs the linker to extract the
__jump_table input section from every compiled object file and arrange
them sequentially in physical memory. This process concatenates the
independently generated metadata entries into a unified array,
leveraging the same linker-driven section aggregation mechanism utilized
for system initialization calls.
The KEEP modifier is crucial because the C code of the kernel does not
directly reference individual elements within this section or invoke
them like standard code symbols. Under aggressive dead-code elimination
optimizations, the linker might categorize the section as unused and
discard it. The KEEP instruction explicitly overrides this behavior,
forcing the linker to preserve the accumulated table.
The surrounding assignments __start___jump_table and
__stop___jump_table define the boundaries of the resulting array.
During boot-time initialization,
jump_label_init
references these linker-defined symbols to locate the table in memory.
Because these markers provide precise boundary addresses, the subsystem
does not require an explicit compile-time count of table entries.
This linker-driven aggregation is confined to the static vmlinux
binary. A kernel module loaded dynamically at runtime cannot participate
in the link phase of the core kernel. Instead, each module maintains a
private __jump_table section within the associated ELF object. When
the module loading subsystem maps a module into memory, the loader reads
this metadata and records the boundary addresses in two fields reserved
inside struct
module:
jump_entries
stores the base address of the array, and
num_jump_entries
stores the number of active entries. When a module shares a key with the
main kernel or another module, these dynamically mapped entries are
encapsulated inside struct
static_key_mod
nodes to integrate them into the central patching system.
8 Size of the patchable site on x86 (runtime)
At runtime, the instruction-patching subsystem on x86 must dynamically determine the size of each patchable site before performing any modification. Build-time optimizations discussed in §6.5 allow the assembler to emit either a 2-byte or a 5-byte instruction at a jump site, depending on the relative displacement to the target. However, struct jump_entry (detailed in §7.2) contains no field or metadata recording the length of the instruction that the assembler selected. When the kernel modifies a jump label on a running system, the patching engine must rediscover the instruction length by parsing the live machine bytes currently sitting in the executable memory.
The kernel delegates this discovery to arch_jump_entry_size():
/* arch/x86/kernel/jump_label.c */
int arch_jump_entry_size(struct jump_entry *entry)
{
struct insn insn = {};
insn_decode_kernel(&insn, (void *)jump_entry_code(entry));
BUG_ON(insn.length != 2 && insn.length != 5);
return insn.length;
}
The retrieval of the target address relies on
jump_entry_code()
(explained in §6.4). The x86
instruction decoder of the kernel,
insn_decode_kernel(),
parses the machine code at that location into struct
insn.
Rather than performing a simple byte count, the decoder fully parses the
opcode, prefixes, and displacement to determine the exact boundary of
the instruction, reporting the result in the length field of the
structure. Whether the memory location contains the compiler-generated
default nop, a branch instruction, or a previously patched instruction
from an earlier state transition, the decoder resolves the true
instruction boundaries.
A sanity check via the
BUG_ON()
macro, evaluating BUG_ON(insn.length != 2 && insn.length != 5), guards
against corruption. If the decoder encounters any length other than
these two supported sizes, the metadata table has desynchronized from
the executable stream, making further patching unsafe.
This dynamic decoding step is unique to the x86 architecture. Architectures with fixed instruction sizes, such as arm64, define a constant JUMP_LABEL_NOP_SIZE. On those platforms, every patchable site shares a uniform, known width, which eliminates the need for runtime discovery. The generic, architecture-independent jump_entry_size() helper encapsulates this architectural difference:
static inline int jump_entry_size(struct jump_entry *entry)
{
#ifdef JUMP_LABEL_NOP_SIZE
return JUMP_LABEL_NOP_SIZE;
#else
return arch_jump_entry_size(entry);
#endif
}
If the architecture defines a global constant size, the compiler
resolves jump_entry_size() to that constant. Otherwise, the helper
falls back to the dynamic decoder.
After resolving the instruction size, the helper __jump_label_patch() prepares both candidate byte sequences (the jump instruction and the corresponding no-op sequence) before selecting the sequence to install:
size = arch_jump_entry_size(entry);
switch (size) {
case JMP8_INSN_SIZE: /* 2 */
code = text_gen_insn(JMP8_INSN_OPCODE, addr, dest);
nop = x86_nops[size];
break;
case JMP32_INSN_SIZE: /* 5 */
code = text_gen_insn(JMP32_INSN_OPCODE, addr, dest);
nop = x86_nops[size];
break;
}
The constants
JMP8_INSN_SIZE,
JMP8_INSN_OPCODE,
JMP32_INSN_SIZE,
and
JMP32_INSN_OPCODE
represent the underlying instruction lengths and raw opcodes (0xEB and
0xE9) for short and near jumps on x86. The utility
text_gen_insn()
synthesizes the target jump sequence, calculating the self-relative
offset as dest - (addr + size). This calculation mirrors the standard
x86 instruction pointer offset convention, resolved here during
execution rather than at compile time.
To obtain the corresponding no-op sequence, the function indexes into
the lookup table
x86_nops
using the decoded size. This lookup supplies a pre-calculated nop
sequence matching the exact width of the site without requiring
instruction generation at runtime.
Generating both candidate sequences beforehand allows the kernel to perform a pre-patching safety validation. Before modifying the instruction stream, the patching engine determines the instruction sequence expected to reside in memory at the destination. This expected sequence is the logical inverse of the target sequence being installed. If the update request specifies JUMP_LABEL_JMP, the memory location must currently hold the no-op sequence. Conversely, if the update request specifies JUMP_LABEL_NOP, the memory location must hold the active jump instruction.
The patching engine invokes memcmp() to compare the live instructions in memory against this expected sequence. A mismatch indicates that the metadata table and the instruction stream have diverged somehow, suggesting memory corruption, a race condition, or an unhandled synchronization failure. In this scenario, proceeding is unsafe. The kernel reports the divergence using pr_crit() and immediately calls BUG(), halting the processor to prevent the execution of arbitrary or corrupted instructions.
9 Life of a static key: boot, enable, disable
Everything in this section is driven by one
atomic_t:
key->enabled.
Its value tells you both “is the feature on” and “is a patch pass
currently running”. The boolean on/off cycle looks like this:
set to -1 patch, set 1 cmpxchg(1,0), off
+----+ +----+ +----+ +----+
| 0 | ----------> | -1 | -------------> | 1 | ------------------> | 0 |
+----+ +----+ +----+ +----+
off enable on off
Read left to right, the three values in that diagram mean:
0— off. Every site for this key holds whichever instruction (noporjmp) means “disabled” for its polarity (§6.1).-1— transient: an enable is in progress,jump_label_update()is actively rewriting every site right now. This value exists so that no concurrent reader can ever observe0during that window and wrongly conclude the feature is still off while the patcher is mid-flight;static_key_count()deliberately reports-1as “1” (enabled) to close that gap.1— on, with the refcount sitting at exactly one net enable (the refcounting ladder below explains what “net” means here).
There is no symmetric -1-like transient for disable: going 1 -> 0
reads as briefly stale “on” instead (see §9.3), which is
harmless.
Above value 1, a second, independent ladder exists purely for
refcounting —
static_branch_inc()/static_branch_dec()
(§4.3) climb and descend it
without ever touching text, because the key is already known to be
enabled:
1 --inc--> 2 --inc--> 3 --inc--> ... (static_key_fast_inc_not_disabled:
1 <--dec-- 2 <--dec-- 3 <--dec-- ... pure atomic increment/decrement,
no jump_label_update() at all)
Only a dec that would land exactly on 1 -> 0 re-enters the boolean
cycle above and triggers a real patch-off. Put differently: of all the
edges across both diagrams, only the three that make up the boolean
cycle itself (0 -> -1, -1 -> 1, and 1 -> 0) ever touch instruction
bytes. Every step on the refcount ladder above 1 (1<->2, 2<->3, …)
is a bare atomic increment or decrement, with no jump_label_update()
call anywhere in it.
That is why the refcounted API in §4.3 exists: many callers can share a key without each one paying for a text-patch round-trip — the cost of patching is paid exactly once, by whichever caller happens to be the first to enable it or the last to disable it.
9.1 Boot: jump_label_init()
Before this function ever runs, the jump-label machinery is in a
half-built state. The linker has already concatenated the slice of
__jump_table
from every translation unit into one array (§7.4),
but that array is simply in link order — sites for the same key can be
scattered anywhere in it, and
key->entries
still holds whatever its static initializer left there, which for a
plain struct static_key key = STATIC_KEY_INIT_FALSE;
(§7.1) is nothing useful. In other words: the raw
table of call sites exists, but nothing yet knows which sites belong
to which key, so
static_branch_enable()
or
static_key_slow_inc()
(§4.3,
§9.2) would have
nothing to walk if called this early.
arch/x86/xen/multicalls.c
declares
static struct static_key mc_debug __ro_after_init;
and registers
xen_mc_debug
as an
early_param()
whose
xen_parse_mc_debug()
callback calls
static_key_slow_inc(&mc_debug)
directly, synchronously, while
parse_early_param()
is walking the command line. If that increment ran before
jump_label_init() had sorted the table and pointed mc_debug at its
entries, it would have nothing to patch and the key would silently stay
un-patched despite the user asking for it on the command line.
The kernel guards against exactly this ordering mistake with one
boolean,
static_key_initialized
— whose only job, per its own comment, is “to generate warnings if
static_key manipulation functions are used before jump_label_init is
called”;
jump_label_init_ro()
later even enforces it with a WARN_ON_ONCE(). That is why
jump_label_init() is called very early from
start_kernel()
in
init/main.c,
strictly before parse_early_param() gets a chance to run any handler
like xen_parse_mc_debug():
void __init jump_label_init(void)
{
...
jump_label_sort_entries(iter_start, iter_stop);
for (iter = iter_start; iter < iter_stop; iter++) {
if (jump_label_type(iter) == JUMP_LABEL_NOP)
arch_jump_label_transform_static(iter, JUMP_LABEL_NOP);
in_init = init_section_contains((void *)jump_entry_code(iter), 1);
jump_entry_set_init(iter, in_init);
iterk = jump_entry_key(iter);
if (iterk == key)
continue;
key = iterk;
static_key_set_entries(key, iter);
}
static_key_initialized = true;
}
Before the loop even starts,
jump_label_sort_entries()
sorts the whole of __jump_table by key, then by code address — the
precondition that lets the “walk all sites for this key as a linear
scan” from §7.3 and the “batch must stay
address-ordered” requirement from
§10.2 both work later.
The loop that follows does three things per entry, in this order, trusting that the table is already sorted:
- Optionally run
arch_jump_label_transform_staticfor NOP sites — a no-op on x86, sinceobjtooland the compiler have already left the right bytes in place (§6.5); other architectures without that build-time trick do real work here. - Mark
__initsites, so nothing later tries to patch a call site living in memory that will be freed once init finishes. - Point each key at the first entry of its contiguous run in the
now-sorted table, so
key->entriesis ready to use the moment something callsstatic_branch_enable().
None of this touches instruction bytes for sites that are already
correct: the compiled-in nop/jmp already matches the initial value
of each key (§6.1). What this
pass builds is bookkeeping — sort order, the __init flag, and the
entries pointer — so that a later toggle knows exactly which sites
to patch and in what order.
jump_label_init_ro() runs much later, from
mark_readonly().
It walks __jump_table a second time, but this pass skips almost
everything — it only acts on keys that
is_kernel_ro_after_init()
recognizes as living in
__ro_after_init
storage.2 For each matching key it calls:
static inline bool static_key_sealed(struct static_key *key)
{
return (key->type & JUMP_TYPE_LINKED) && !(key->type & ~JUMP_TYPE_MASK);
}
static inline void static_key_seal(struct static_key *key)
{
unsigned long type = key->type & JUMP_TYPE_TRUE;
key->type = JUMP_TYPE_LINKED | type;
}
static_key_seal()
keeps only the
JUMP_TYPE_TRUE
bit of the current
key->type
and throws the rest of the word away, replacing it with
JUMP_TYPE_LINKED
plus that one preserved bit. Recall from §7.1 that
this word is normally a tagged pointer — the low 2 bits are a type tag,
and everything above them is either the entries pointer the loop
earlier in this section just installed, or a module-chain next pointer
from §11. Sealing collapses it down to
just the tag, with nothing left standing above bit 1:
before sealing (a live `entries` or `next` pointer, tagged):
bit63 bit1 bit0
[ real pointer value ..................... ] [ L ] [ T ]
after static_key_seal():
bit63 bit1 bit0
[ 0000000000000000000000000000000000000000 ] [ 1 ] [ T ]
T is the preserved JUMP_TYPE_TRUE bit; L is JUMP_TYPE_LINKED.
static_key_sealed()
is exactly the test for the “after” picture — JUMP_TYPE_LINKED set and
nothing above bit 1 — so it can recognize a sealed key at a glance,
regardless of which of the two pointer kinds that word used to hold.
Several call sites can share the same key, so this loop visits the same
key more than once as it walks the table entry by entry.
static_key_sealed() is checked before static_key_seal() runs, so the
first visit seals the key and every later visit for that same key
becomes a no-op.
This is safe only because __ro_after_init is a promise that the key is
never toggled again after boot: once sealed,
key->entries
is gone, so nothing could walk it even if some later code mistakenly
tried.
The payoff shows up in §11. When a module
loaded afterward references a sealed key,
jump_label_add_module()
skips allocating and chaining a
struct static_key_mod
for it — the bookkeeping §11 otherwise
needs so a future toggle can find and patch call sites living in other
modules. Instead, it just patches the sites of that module once,
immediately, to match the already-final value of the key.
The ordering against
mark_rodata_ro()
follows from the same fact: sealing is the last write anything makes to
key->type,
and that field sits inside the very __ro_after_init section
mark_rodata_ro() is about to make actually read-only in the page
tables. Running first just means the field has already settled into its
final value before write access to it disappears.
9.2 Enabling: static_key_enable() / static_branch_enable()
jump_label_init()
(§9.1) only builds bookkeeping — it never flips
a key. Every site in the kernel image boots running whatever nop/jmp
the compiler and objtool baked in
(§6.1,
§6), on or off, and stays that way
until something calls static_branch_enable() or
static_key_slow_inc()
for the first time. This section is about that first flip, which can
happen years into uptime, on a system where other CPUs may already be
executing the very instructions about to be rewritten.
drivers/md/dm-stats.c
shows a concrete case of this
(§4.4): the first time a
user asks the device-mapper stats ioctl to start recording per-region
I/O counters, it flips
stats_enabled,
a key that has sat disabled since boot:
if (!static_key_enabled(&stats_enabled.key))
static_branch_enable(&stats_enabled);
The guard matters as much as the call: it is what keeps a second, third,
or hundredth request for the same device from re-triggering a full patch
round once the key is already on
(§4.4). But the first
time through, this
static_branch_enable(&stats_enabled)
does run, and from that call onward every stats_enabled check in the
I/O path — on every CPU, some of which may be running that exact code
right now — has to observe the new state, and none of them may ever see
a half-patched instruction.
static_key_enable_cpuslocked()
below is what makes that safe:
void static_key_enable_cpuslocked(struct static_key *key)
{
...
jump_label_lock();
if (atomic_read(&key->enabled) == 0) {
atomic_set(&key->enabled, -1); /* "enabling" */
jump_label_update(key); /* patch all sites */
atomic_set_release(&key->enabled, 1);
}
jump_label_unlock();
}
Walking through what that function does, in order:
- It first checks whether the key is already enabled, and bails out if so.
jump_label_mutexserializes all jump-label patching globally, so two callers enabling different keys at the same time still cannot have their text pokes race one another.- It sets
enabled = -1before touching any code — this is the0 → -1transient from the state diagram above, and it is what makesjump_label_update()below safe to run concurrently with readers:static_key_count()andstatic_key_enabled()both treat-1as enabled, so no concurrent reader ever sees a window of “disabled” while sites are only half-patched. jump_label_update()does the actual work: it patches every site for this key (§9.5).- Finally, it stores
1.
static_key_enable() wraps this in
cpus_read_lock()
so CPUs cannot come online mid-patch.
The refcounted path,
static_key_slow_inc_cpuslocked(),
is not just this same function reused for inc/dec
(§4.3). With multiple
independent owners, more than one CPU can call it at once, and only one
of them may actually be the one that flips 0 → 1:
bool static_key_slow_inc_cpuslocked(struct static_key *key)
{
lockdep_assert_cpus_held();
if (static_key_fast_inc_not_disabled(key))
return true;
guard(mutex)(&jump_label_mutex);
if (!atomic_cmpxchg(&key->enabled, 0, -1)) {
jump_label_update(key);
atomic_set_release(&key->enabled, 1);
} else {
if (WARN_ON_ONCE(!static_key_fast_inc_not_disabled(key)))
return false;
}
return true;
}
It tries a lock-free fast path first (below), and only takes
jump_label_mutex if that fails. Past that point it is the same 0 → 1
dance as static_key_enable_cpuslocked() above: exactly one caller
actually performs the transition and runs jump_label_update(); anyone
else who reaches the mutex finds the key already on and just falls back
to the fast path to add their own count.
static_key_fast_inc_not_disabled()
is what makes “further incs” in the table from
§4.3 cost nothing more than an
atomic — no lock, no text poke:
bool static_key_fast_inc_not_disabled(struct static_key *key)
{
int v;
STATIC_KEY_CHECK_USE(key);
/*
* Negative key->enabled has a special meaning: it sends
* static_key_slow_inc/dec() down the slow path, and it is non-zero
* so it counts as "enabled" in jump_label_update().
*
* The INT_MAX overflow condition is either used by the networking
* code to reset or detected in the slow path of
* static_key_slow_inc_cpuslocked().
*/
v = atomic_read(&key->enabled);
do {
if (v <= 0 || v == INT_MAX)
return false;
} while (!likely(atomic_try_cmpxchg(&key->enabled, &v, v + 1)));
return true;
}
The retry loop only succeeds once it can prove, atomically, that the
count is already an ordinary positive number. v <= 0 catches both a
genuinely disabled key (0) and an enable already in progress (-1),
sending both cases down to the slow path above instead of incrementing a
value that doesn’t mean what it looks like yet. v == INT_MAX guards
against overflows.
9.3 Disabling
§9.2 walked through
what happens when a key crosses from off to on. Disabling is the same
problem in reverse: patch every site back to its original instruction,
without letting any reader ever observe a torn one. It runs on the very
same rule as enabling — not every call to disable/dec needs to touch
an instruction at all.
jump_label_update()
only has to run on the one transition that actually flips what is
patched into the instruction stream: 1 → 0 for the boolean API,
N → 0 for the refcounted one. Every other call just moves a plain
integer and can be answered with a single atomic instruction — no lock,
no text poke, because nothing about the patched code needs to change.
That is exactly the same shape as the
static_key_fast_inc_not_disabled()
from §9.2 on the
enable side — a bare CAS loop, no lock, no patch, for every increment
that doesn’t cross 0 → 1. That rule (every transition that isn’t on a
patch-triggering boundary is a bare atomic op) is what both functions
below are built around, on either side of the key.
void static_key_disable_cpuslocked(struct static_key *key)
{
...
if (atomic_read(&key->enabled) != 1) {
WARN_ON_ONCE(atomic_read(&key->enabled) != 0);
return;
}
jump_label_lock();
if (atomic_cmpxchg(&key->enabled, 1, 0) == 1)
jump_label_update(key);
jump_label_unlock();
}
static_key_disable_cpuslocked()
is the mirror, in the boolean API, of the enable path from
§9.2, but it skips
the -1 choreography. It first checks that the key is currently exactly
1, bailing out (and warning if the value isn’t 0 either, which would
mean the boolean and refcounted APIs got mixed on this key —
§4.3) before doing anything
else. The actual disable is then a single
atomic_cmpxchg(&key->enabled, 1, 0): “if the value is currently 1,
replace it with 0, and tell me whether you succeeded.” If some other
CPU changed it first, the compare fails and this call does nothing
further. Only on success does it call jump_label_update() to patch
every site for this key back to its disabled instruction.
Compare that to enabling: there,
enabled
is deliberately set to -1 before any text is touched, specifically
so no reader can mistake in-progress patching for “off” (point 3 of
§9.2). Disabling has
no matching problem to solve. During the window between the cmpxchg
above and jump_label_update() finishing, enabled already reads 0
while some sites out there are still physically holding their “enabled”
instruction — the opposite kind of staleness from enabling (stale “on”
instead of stale “off”), but just as harmless. A reader who calls
static_key_enabled()
during that window is simply told “off” a few instructions before the
code itself has caught up;
§4.4 already covers why
these state reads only ever need to be eventually correct, never
instruction-exact.
The refcounted side runs through a different function,
__static_key_slow_dec_cpuslocked(),
which only reaches jump_label_update() on the one decrement that
actually drives the count to 0
(atomic_dec_and_test()
reports true exactly then, and only then). Every decrement that lands
above 1 — 3 -> 2, 2 -> 1, and so on — is intercepted earlier, by
static_key_dec_not_one(),
which performs a plain atomic decrement and returns without ever taking
the jump-label lock or looking at an instruction stream. That is the
same “every transition that isn’t on a patch-triggering boundary is a
bare atomic op” rule we discussed earlier, now seen from the decrement
side:
static bool static_key_dec_not_one(struct static_key *key)
{
int v;
v = atomic_read(&key->enabled);
do {
WARN_ON_ONCE(v < 0);
if (WARN_ON_ONCE(v == 0))
return true;
if (v <= 1)
return false;
} while (!likely(atomic_try_cmpxchg(&key->enabled, &v, v - 1)));
return true;
}
The CAS loop is what makes this safe against concurrent decrements: if
another CPU wins the race and changes
key->enabled
between the read of this CPU and its cmpxchg, v is refreshed and the
loop just re-checks the same two conditions against the new value rather
than clobbering it.
9.4 Deferred / rate-limited dec
This is the internal counterpart to the rate-limited disable in §4.7, seen here from the function-by-function angle of this section rather than the call-site angle of the cookbook. §4.7 already covers why a delayed decrement exists and how the coalescing works when toggles arrive faster than the timeout; this is only the piece that was left out there — where the state-machine logic actually lives.
__static_key_slow_dec_deferred()
opens with the exact same
static_key_dec_not_one()
check §9.3 just introduced: if this decrement would not
bring the count down to the 1 -> 0 boundary, it is a plain atomic
decrement and the function returns immediately, no different from the
non-deferred path.
The two paths only diverge on the one decrement that would disable the
key. Where the
__static_key_slow_dec_cpuslocked()
from §9.3 reaches straight for
jump_label_update()
at that point, this function instead calls
schedule_delayed_work()
— a standard kernel workqueue primitive that runs a callback once a
given delay has elapsed, rather than right away — and returns without
touching a single instruction. The count is deliberately left at 1
rather than dropped to 0; only the plan to disable has been
recorded, in the timer. In full:
void __static_key_slow_dec_deferred(struct static_key *key,
struct delayed_work *work,
unsigned long timeout)
{
if (static_key_dec_not_one(key))
return;
schedule_delayed_work(work, timeout);
}
When that timer eventually fires,
jump_label_update_timeout()
runs the ordinary, undeferred decrement path from §9.3. If
nothing else touched the key in the meantime, the count is still exactly
1, the decrement finally lands on the 1 -> 0 boundary, and
jump_label_update() runs for real. If instead another caller
incremented the key again while the timer was pending, the count is no
longer 1 by the time the timer fires — so static_key_dec_not_one()
intercepts that decrement too, as an ordinary atomic op, and
jump_label_update() is never reached. Nothing needed patching back,
because nothing was ever patched in the first place.
9.5 jump_label_update() → __jump_label_update()
Every path in §9.2 through §9.4 eventually funnels into this function — it is the one that actually walks the sites of a key and asks the architecture layer to patch each one. Reading the outer function first:
static void jump_label_update(struct static_key *key)
{
...
if (static_key_linked(key)) {
__jump_label_mod_update(key); /* walk module list */
return;
}
entry = static_key_entries(key);
if (entry)
__jump_label_update(key, entry, stop, init);
}
The first branch is the module case from §7.1: if
bit 1 of
key->type
is set
(LINKED),
the sites of this key are not one contiguous run inside
__jump_table,
but scattered across a linked list of per-module entry tables
(struct static_key_mod),
so a separate helper has to walk that list instead of a flat array —
§11 covers
__jump_label_mod_update()
in full. Otherwise,
static_key_entries(key)
recovers the pointer §9.1 stored during boot —
the first entry of the contiguous run for this key inside the sorted,
vmlinux-only table — and the real patching happens in
__jump_label_update().
On x86, which defines
HAVE_JUMP_LABEL_BATCH,
that function looks like this:
for (; entry < stop && jump_entry_key(entry) == key; entry++) {
if (!jump_label_can_update(entry, init))
continue;
if (!arch_jump_label_transform_queue(entry, jump_label_type(entry))) {
arch_jump_label_transform_apply();
BUG_ON(!arch_jump_label_transform_queue(...));
}
}
arch_jump_label_transform_apply();
The loop condition — advance while entry < stop and
jump_entry_key(entry)
== key — only works because of the sort from
§9.1: every site belonging to this key is
guaranteed to sit in one unbroken run starting at entry, so the loop
can walk forward blindly and stop the instant it reaches a site
belonging to some other key, with no need to search the rest of the
table.
For each site still in range,
jump_label_type(entry)
computes enabled ^ branch — the exact static/dynamic formula from the
table in §6.1, but now
evaluated against the current live state of the key rather than its
compile-time initial value — to decide whether this specific site should
end up holding a
JUMP_LABEL_NOP
or a
JUMP_LABEL_JMP.
Before acting on that answer,
jump_label_can_update()
filters out two kinds of site that must not be touched at all: one still
living in __init text after boot has finished (that memory may already
have been freed, which is exactly what the
jump_entry_is_init
flag from §9.1 was recorded to detect), and one
that
kernel_text_address()
does not even recognize as live kernel text — built-in code that is
__exit-only and therefore can never run, so patching it would be
pointless even though it is technically still present:
static bool jump_label_can_update(struct jump_entry *entry, bool init)
{
if (!init && jump_entry_is_init(entry))
return false;
if (!kernel_text_address(jump_entry_code(entry))) {
WARN_ONCE(!jump_entry_is_init(entry),
"can't patch jump_label at %pS",
(void *)jump_entry_code(entry));
return false;
}
return true;
}
Rather than patching each surviving site immediately, one at a time, the
loop splits the work into two separate jobs:
arch_jump_label_transform_queue()
builds up a batch, and
arch_jump_label_transform_apply()
executes it. Splitting them apart is what lets every site belonging to
one key ride through a single INT3-synchronized patch round
(§10.2) instead of paying for one
round per site.
Queuing a site. Each call to arch_jump_label_transform_queue()
computes the replacement bytes for that one site and hands
(address, new bytes, length) to
smp_text_poke_batch_add().
That function appends the request to a pending array; nothing is written
to memory yet. The one exception is early boot: only one CPU is running,
so there is no concurrent fetcher to synchronize against and nothing
worth batching — the function calls the non-batching
arch_jump_label_transform()
directly instead.
Applying the batch. arch_jump_label_transform_apply() executes
everything queued so far. It calls
smp_text_poke_batch_finish()
which runs the three-step INT3 dance once for the whole batch instead of
once per site. __jump_label_update() calls it once, after its loop
ends, to flush whatever is still pending.
10 x86 text patching: the gory details
This is where the torn-write argument from §5.3 turns into working code.
10.1 Early boot vs live SMP
Very early in boot, only the boot CPU is running and .text is still
writable.
__jump_label_transform(),
the function every x86 patch eventually funnels through, checks for
exactly that window and takes the cheap route whenever it still holds:
/* arch/x86/kernel/jump_label.c: __jump_label_transform() */
if (init || system_state == SYSTEM_BOOTING) {
text_poke_early(...); /* IRQ-disabled + `memcpy()` + sync_core() */
return;
}
smp_text_poke_single(...); /* or batch_add during queueing (§9.2) */
system_state == SYSTEM_BOOTING is a proxy for one fact: only the boot
CPU exists so far. That alone is reason enough to skip the whole
IPI-synchronized protocol of
§10.3 — with
nobody else around to race the write,
text_poke_early()
can just do a plain IRQ-disabled + memcpy() +
sync_core()
and be done. .text also happens to still be writable at this point, so
the alias trick from §5.4
isn’t needed either — but that is a bonus the check gets for free, not
something it verifies directly:
smp_init()
wakes every other CPU well before
mark_rodata_ro()
ever runs, so there is a real stretch of boot where other CPUs are
already up while .text is still writable. Jump labels take the full
protocol for that entire stretch anyway, because the only thing this
check ever verifies is whether this CPU is still provably alone.
The leading init in that condition is not the per-site __init-text
flag from §9.1
(jump_entry_is_init()).
It is a separate, whole-system flag threaded down from
init = system_state < SYSTEM_RUNNING inside
jump_label_update()
itself (§9.5), true a little
longer than SYSTEM_BOOTING alone. On the batching path of x86, though,
that value never actually reaches here:
arch_jump_label_transform_queue()
only calls this function through its own
system_state == SYSTEM_BOOTING fallback, passing a hardcoded 0 for
init every time it does. So on x86 this condition is, in practice,
exactly system_state == SYSTEM_BOOTING — the init || half of it only
ever matters on architectures that call this function directly, without
going through batching at all.
10.2 Batching API used by jump labels
§10.1 settled how a single site gets patched once the decision to patch it is made; this section is about when jump labels actually pull that trigger. A busy tracepoint can have thousands of call sites sharing one key, and paying the full IPI-synchronized protocol (§10.3) separately for each one would be needless — the sites can be collected first and the expensive part paid once for the whole group. Two arch-level hooks make that possible: one that queues the newly computed bytes for a site without touching hardware yet, and one that flushes everything queued so far in a single synchronized pass:
bool arch_jump_label_transform_queue(...)
{
if (system_state == SYSTEM_BOOTING) {
arch_jump_label_transform(entry, type);
return true;
}
mutex_lock(&text_mutex);
jlp = __jump_label_patch(entry, type);
smp_text_poke_batch_add(addr, jlp.code, jlp.size, NULL);
mutex_unlock(&text_mutex);
return true;
}
void arch_jump_label_transform_apply(void)
{
mutex_lock(&text_mutex);
smp_text_poke_batch_finish();
mutex_unlock(&text_mutex);
}
The queue collects many sites (a busy tracepoint may have thousands);
one
batch_finish()
amortizes the IPI syncs. The queue must stay address-sorted; if a
new address would break order, or the page-sized array is full,
smp_text_poke_batch_add()
flushes early
(text_poke_addr_ordered()
in
alternative.c).
That is why
jump_label_cmp
sorts by code address within each key.
How many fit in one batch? The pending patches live in a single
statically-allocated page,
struct smp_text_poke_loc:
struct smp_text_poke_loc {
s32 rel_addr; /* addr := _stext + rel_addr */
s32 disp; /* branch displacement, for emulation */
u8 len; /* 1, 2, 5, or 6 bytes */
u8 opcode; /* first opcode byte, for emulation */
u8 text[5]; /* the new instruction bytes */
u8 old; /* byte that used to be there, for perf/PT tracing */
}; /* 16 bytes, naturally aligned */
#define TEXT_POKE_ARRAY_MAX (PAGE_SIZE / sizeof(struct smp_text_poke_loc))
/* 4096 / 16 = 256 entries per flush on a 4K-page x86_64 build */
rel_addr is relative to
_stext3,
not to the entry itself like
jump_entry
— cheap, because every patch site is, by definition, in kernel text. A
tracepoint or jump-label key with more than 256 call sites needs more
than one batch_finish() round (i.e. more than 3 IPI rounds,
§10.3) to fully
enable/disable.
__jump_label_update()
(§9.5) does contain a generic
“queue full → apply → retry” branch, and it looks like the natural place
to expect this 256-site overflow to be handled.
On x86 the “queue full → apply → retry” branch never runs, though: it
only fires when
arch_jump_label_transform_queue()
itself returns false, and the implementation of that function on this
architecture never returns false. So the branch is dead code here.
for (; (entry < stop) && (jump_entry_key(entry) == key); entry++) {
if (!jump_label_can_update(entry, init))
continue;
if (!arch_jump_label_transform_queue(entry, jump_label_type(entry))) {
/*
* Queue is full: Apply the current queue and try again.
*/
arch_jump_label_transform_apply();
BUG_ON(!arch_jump_label_transform_queue(entry, jump_label_type(entry)));
}
}
arch_jump_label_transform_apply();
The overflow is actually caught one layer further down, entirely inside
smp_text_poke_batch_add() — the entire guard is these three lines:
void smp_text_poke_batch_add(void *addr, const void *opcode, size_t len, const void *emulate)
{
if (text_poke_array.nr_entries == TEXT_POKE_ARRAY_MAX || !text_poke_addr_ordered(addr))
smp_text_poke_batch_finish();
__smp_text_poke_batch_add(addr, opcode, len, emulate);
}
Every iteration of that for loop does one thing: it calls
arch_jump_label_transform_queue(), which computes the new bytes for
one site and appends one
smp_text_poke_loc
to the array — the cheap, per-site step this section has been
describing. The if (!arch_jump_label_transform_queue(...)) branch is
the “queue full → apply → retry” path discussed above; on x86 it never
runs, since that function never returns false. Nothing expensive
happens inside the loop.
arch_jump_label_transform_apply()
sits outside the loop, called exactly once after it exits, for every
site the loop just queued. That single call is what finally triggers
batch_finish(), the three-phase IPI-synchronized protocol from
§10.3.
So the entire set of call sites for a key — whether it has one or close
to 256 — rides through on that one shared batch_finish(), and only a
key with more than 256 sites forces a second round.
10.3 The INT3 SMP algorithm (smp_text_poke_batch_finish)
This is the payoff of everything
§5.3 through
§10.2 built toward. The key fact
from §5.3 was that
only a single-byte store is atomic with respect to instruction fetch —
nothing wider is. The protocol below never trusts a multi-byte write to
be safe on its own; instead it uses one atomic single-byte store to
plant a trap on top of the site, uses that trap to absorb any CPU
unlucky enough to fetch through mid-update, and only then fills in the
rest. Three writes, three synchronizations, one site at a time across
the whole batch. Documented at the top of the function in
alternative.c:
For each site in the vector:
(1) Write INT3 (0xCC) over the first byte
→ IPI sync all CPUs (serialize pipelines / I-caches)
(2) Write bytes 1..N-1 of the new instruction
→ IPI sync again unnecessary, according to Intel,
but better safe than sorry
(3) Write byte 0 of the new instruction (replaces INT3)
→ IPI sync again
Here are the actual bytes of one site across the three phases, patching
a 5-byte NOP (0f 1f 44 00 00) into a 5-byte JMP rel32 (e9 + 4-byte
displacement, shown as ?? ?? ?? ??):
start (before) 0f 1f 44 00 00 any fetch: executes the NOP
phase 1 (INT3 in) cc 1f 44 00 00 any fetch: #BP -> handler emulates
^^ the NEW instruction (jumps to l_yes)
trap byte
-------- IPI sync --------
phase 2 (tail in) cc ?? ?? ?? ?? same as phase 1: byte 0 is still
^^ INT3, so any fetch still traps and
still traps gets emulated — the real tail bytes
underneath are now correct, but
nothing reads them yet
-------- IPI sync --------
phase 3 (done) e9 ?? ?? ?? ?? any fetch: executes the real JMP
^^ real opcode directly, no trap needed anymore
-------- IPI sync --------
The subtle point this diagram is here to make: the observable behavior
of the site flips the instant the sync in phase 1 completes, not at
phase 3. From phase 1 onward, any CPU landing on this address — whether
by falling into it in a hot loop or by literally executing byte 0 — gets
the effect of the new instruction, because the #BP handler always
emulates the pending new instruction (it was computed and stashed in
the queue back at
smp_text_poke_batch_add()
time, long before phase 1 starts). A CPU that hits the address
mid-update and one that hits it after phase 3 land on the same outcome —
the only difference is whether it got there by trapping into the handler
or by executing the finished bytes directly. Phases 2 and 3 exist to
make that direct path available, so steady-state execution stops paying
the #BP tax.
The writing side. All of the above is driven by
smp_text_poke_batch_finish() (it early-returns immediately if
text_poke_array.nr_entries is 0 — nothing queued, nothing to do).
Trimmed of the
cond_resched()
softlockup guard, the perf/Intel-PT tracing hook, and a 6-byte-opcode
edge case, it opens by arming the refcount of every CPU — the release
side of the release/acquire pairing the “Why INT3?” sidebar below
explains — then runs the three phases in order:
for_each_possible_cpu(i)
atomic_set_release(per_cpu_ptr(&text_poke_array_refs, i), 1);
smp_wmb();
Phase 1 writes INT3 over the first byte of every site, saving the
byte it replaces (for the perf/PT tracing hook trimmed out above), then
syncs once for the whole batch:
for (i = 0; i < text_poke_array.nr_entries; i++) {
text_poke_array.vec[i].old = *(u8 *)text_poke_addr(&text_poke_array.vec[i]);
text_poke(text_poke_addr(&text_poke_array.vec[i]), &int3, INT3_INSN_SIZE);
}
smp_text_poke_sync_each_cpu();
Phase 2 writes bytes 1..len-1 of the new instruction for every
site — safe, since byte 0 is still INT3 — and syncs again, but only if
some site actually has tail bytes to write (a single-byte patch has
none):
for (do_sync = 0, i = 0; i < text_poke_array.nr_entries; i++) {
int len = text_poke_array.vec[i].len;
if (len - INT3_INSN_SIZE > 0) {
text_poke(text_poke_addr(&text_poke_array.vec[i]) + INT3_INSN_SIZE,
text_poke_array.vec[i].text + INT3_INSN_SIZE,
len - INT3_INSN_SIZE);
do_sync++;
}
}
if (do_sync)
smp_text_poke_sync_each_cpu();
Phase 3 writes byte 0, replacing the INT3 — skipped for the corner
case (discussed next) where the new opcode is itself 0xCC — and
syncs again if anything actually changed:
for (do_sync = 0, i = 0; i < text_poke_array.nr_entries; i++) {
u8 byte = text_poke_array.vec[i].text[0];
if (byte == INT3_INSN_OPCODE)
continue;
text_poke(text_poke_addr(&text_poke_array.vec[i]), &byte, INT3_INSN_SIZE);
do_sync++;
}
if (do_sync)
smp_text_poke_sync_each_cpu();
Draining after phase 3. After the last sync, the writer cannot yet
assume every CPU has left the handler — a CPU could be inside
smp_text_poke_int3_handler()
right up until that sync completes (it entered before the sync, is still
emulating). So smp_text_poke_batch_finish() ends with this drain and
the final reset:
for_each_possible_cpu(i) {
atomic_t *refs = per_cpu_ptr(&text_poke_array_refs, i);
if (unlikely(!atomic_dec_and_test(refs)))
atomic_cond_read_acquire(refs, !VAL);
}
text_poke_array.nr_entries = 0;
atomic_dec_and_test() decrements the refcount of every CPU; for any
that don’t immediately hit zero, atomic_cond_read_acquire(refs, !VAL)
spin-waits — i.e. for stragglers still mid-handler to finish and drop
their own reference. Only then does the function reset
text_poke_array.nr_entries = 0, making the buffer safe to reuse for
the next batch.
In the common case — jump labels, static calls, ftrace, none of which
ever replace a site with a literal 0xCC — phase 3 already wrote every
final byte, and
smp_text_poke_sync_each_cpu()
already fenced out any in-flight handler. So this drain loop observes
zero immediately, and the comment in the source calls it out explicitly:
“unless the replacement instruction is INT3, this case goes unused.”
It exists for the corner case (other clients, not jump labels) where the
new opcode of a site is 0xCC itself. Byte 0 is therefore left alone
in phase 3 (writing 0xCC over an existing 0xCC would be a no-op
anyway, so that per-entry sync is skipped), and the only thing standing
between “batch done” and “safe to reuse the array” is the
atomic_cond_read_acquire() spin-wait itself.
Why INT3? The whole point of writing
INT3first (phase 1 of the three-phase protocol above) is that it gives the kernel a way to intercept any CPU that would otherwise have executed half-written bytes, and make it run the finished instruction instead. While the site is mid-update, any CPU that hits it takes#BP(a fault, vector 3), and that fault always lands insmp_text_poke_int3_handler(), which is wired up as the#BPhandler intraps.c. Its locals are justtpl(the matchedstruct smp_text_poke_loc*),ret, andip; here is what it actually does, one check at a time.1. Bail on user mode. This handler only ever concerns itself with traps hit while executing kernel text:
if (user_mode(regs)) return 0;2. Confirm a batch is actually in flight, via a per-CPU refcount,
text_poke_array_refs:smp_rmb(); if (!try_get_text_poke_array()) return 0;
try_get_text_poke_array()is just an atomic increment-if-nonzero:static __always_inline bool try_get_text_poke_array(void) { atomic_t *refs = this_cpu_ptr(&text_poke_array_refs); return raw_atomic_inc_not_zero(refs); /* 0 → fails: no batch active */ }Before touching any site,
smp_text_poke_batch_finish()arms the refcount of every CPU to 1 withatomic_set_release(), thensmp_wmb()s, and only then writes the INT3 bytes. Thesmp_rmb()above is the mirror-image acquire. That release/acquire pairing is what guarantees: if the#BPof a CPU fires (meaning it must have fetched the INT3 the writer stored), that same CPU is also guaranteed to see a fully-populated, non-zero-refcounttext_poke_array— never a half-written vector.3. Find which site trapped, binary-searching
text_poke_array.vecfor thisregs->ip - 1(skipping straight to a direct compare when there is exactly one entry):ip = (void *) regs->ip - INT3_INSN_SIZE; if (unlikely(text_poke_array.nr_entries > 1)) { tpl = __inline_bsearch(ip, text_poke_array.vec, text_poke_array.nr_entries, sizeof(struct smp_text_poke_loc), patch_cmp); if (!tpl) goto out_put; } else { tpl = text_poke_array.vec; if (text_poke_addr(tpl) != ip) goto out_put; } ip += tpl->len;4. Emulate the new instruction by editing
regsand returning — the return address of the#BPhandler becomes the effect of the emulated instruction, soiretresumes execution as if the new bytes had actually executed:switch (tpl->opcode) { case INT3_INSN_OPCODE: goto out_put; /* explicit INT3, not ours; do not consume */ case RET_INSN_OPCODE: int3_emulate_ret(regs); break; case CALL_INSN_OPCODE: int3_emulate_call(regs, (long)ip + tpl->disp); break; case JMP32_INSN_OPCODE: case JMP8_INSN_OPCODE: int3_emulate_jmp(regs, (long)ip + tpl->disp); break; case 0x70 ... 0x7f: /* Jcc */ int3_emulate_jcc(regs, tpl->opcode & 0xf, (long)ip, tpl->disp); break; default: BUG(); } ret = 1;
JMP rel8/rel32(also a pending NOP, encoded asJMPwithdisp == 0, i.e. “jump to the next instruction”):int3_emulate_jmp()just overwritesregs->ip.CALL/RET/Jcc: same idea (push a fake return address / pop one / conditionally add the displacement) — jump labels never generate these, but static calls and ftrace share this exact engine and do.- If the replacement opcode is itself
0xCC(some other text-poke client is intentionally installing a breakpoint, not passing through this emulator): the handler does not consume the trap — it falls through so that debugging infrastructure (kgdb, kprobes) gets its#BP.Jump labels only ever exercise the
JMP32/JMP8case — theRET/CALL/Jccarms exist because static calls and ftrace share this exact handler.5. Release the refcount and report the trap as handled:
out_put: put_text_poke_array(); return ret;So no CPU ever executes a torn multi-byte instruction: control either sees the old bytes (before the sync in phase 1 completes everywhere), takes the INT3+emulation path (the entire window from phase 1 to phase 3), or sees the finished new bytes (after the sync in phase 3).
10.4 What “sync” means
The three-phase protocol from §10.3 names “IPI sync” as a step three times over, without ever saying what happens when that IPI lands. This section answers that — not what the patching CPU does (§10.3 already covered that), but what every other CPU is forced to do in response.
smp_text_poke_sync_each_cpu()
is the function invoked at each of those three points, and its job is
deceptively narrow: send an IPI to every CPU other than the one doing
the patching, have each of them run
sync_core(),
and block until every last one has reported back. That blocking is not
incidental — the whole structure of
§10.3 depends on
each phase being globally visible before the next one begins. If the
patching CPU raced ahead and wrote the tail bytes of phase 2 while some
other core was still mid-fetch on the phase-1 view of the site, the
entire point of planting the INT3 trap first would be defeated.
What actually happens on the receiving end is where the CPU model from §5.1 finally pays for itself. A modern core does not execute an instruction the moment it sees its bytes: it fetches ahead of where it is currently retiring, decodes into microcodes, and may be holding several instructions of that pipeline in flight at once. A CPU that fetched the old bytes moments before the patch landed can still be sitting on a stale decode of them — and no ordinary memory write, however carefully sequenced, undoes that. Something has to reach into the pipeline itself and discard the stale work. That is exactly what a serializing instruction is architecturally defined to do: retire everything already in flight, drop any speculative or partially-decoded work, and guarantee that the very next fetch goes out fresh.
Two ways to get that guarantee are available, and which one runs depends
on the CPU generation.
X86_FEATURE_SERIALIZE,
present on most CPUs manufactured since around 2020, provides a
dedicated instruction whose only job is this flush — cheap and direct:
static __always_inline void serialize(void)
{
/* Instruction opcode for SERIALIZE; supported in binutils >= 2.35. */
asm volatile(".byte 0xf, 0x1, 0xe8" ::: "memory");
}
That comment is the real reason this is written as raw .byte values
rather than a mnemonic the assembler recognizes: SERIALIZE was only
added to binutils in 2.35, so a kernel built with an older assembler
still needs to be able to emit the three raw opcode bytes of the
instruction (0F 01 E8) by hand. The "memory" clobber tells the
compiler this call is a full optimization barrier — it must not reorder
ordinary memory accesses across it.
On older hardware, the kernel falls back to
iret_to_self().
An interrupt frame is just the handful of words — return SS,
RSP, RFLAGS, CS, and RIP — that the CPU itself pushes onto the
stack whenever a real interrupt or exception fires, and which iret
later pops to hand control back to whatever was running. Normally
software never builds one by hand; the CPU builds it automatically at
the moment of a real trap, the same way it did for the #BP frame the
INT3 handler above edited before its own iret. iret_to_self() does
the half of that job the CPU would do: it pushes those same five words
by hand, with no real interrupt behind them, sets the saved RIP field
to the very next instruction after the iret, and then executes iret
against that fabricated frame. As far as the CPU can tell, this is a
completely ordinary return from an interrupt, so it does everything
architecture requires of one — including the serializing flush this
function exists to get. But because the fabricated return address is
simply “keep going from here,” nothing about the actual control flow
changes: execution lands right back where it would have anyway, one
instruction later.
Here is the actual body, from
arch/x86/include/asm/sync_core.h
(the 64-bit variant — a 32-bit build takes a shorter path that skips the
SS/RSP pushes, since a same-privilege 32-bit iret doesn’t pop
them):
static __always_inline void iret_to_self(void)
{
unsigned int tmp;
asm volatile (
"mov %%ss, %0\n\t" /* SS is a segment register; it can't be */
/* pushed directly in this form, so copy */
/* it into a GPR first */
"pushq %q0\n\t" /* frame field 1 (bottom): return SS */
"pushq %%rsp\n\t" /* frame field 2: return RSP — but this */
/* captures RSP *after* the SS push above */
/* already moved it down by 8 */
"addq $8, (%%rsp)\n\t" /* ...so correct the just-pushed copy back */
/* up by 8, to the RSP value from before */
/* this function started pushing anything */
"pushfq\n\t" /* frame field 3: return RFLAGS */
"mov %%cs, %0\n\t" /* same GPR trick as SS, for CS this time */
"pushq %q0\n\t" /* frame field 4: return CS */
"pushq $1f\n\t" /* frame field 5 (top): return RIP — the */
/* address of local label "1:" below, i.e. */
/* the instruction right after this one */
"iretq\n\t" /* pop all five fields and "return" — the */
/* CPU treats this exactly like returning */
/* from a genuine interrupt */
"1:" /* execution resumes here, indistinguishable */
/* from simply falling through to this point */
: "=&r" (tmp), ASM_CALL_CONSTRAINT : : "cc", "memory");
}
The iret is architecturally required to be a serializing event on
every CPU, which is precisely the guarantee this fallback needs — it
works identically at any privilege level (so it survives under
paravirtualization) and never exits to a hypervisor, both properties
this code cannot give up. The price is that it measures a bit more than
twice as slow as the dedicated instruction, and it unconditionally
unmasks NMIs, which SERIALIZE does not.
One more candidate is conspicuously missing. CPUID also serializes,
and on paper looks like the most portable option of all. The kernel
avoids it here for a practical reason, not a correctness one: under
virtualization, CPUID commonly traps out to the hypervisor, and this
is exactly the kind of hot, latency-sensitive path — run on every online
CPU, on every single key toggle — that cannot tolerate an unpredictable
VM exit in the middle of it.
The choice is a single feature check, decided once per call:
static __always_inline void sync_core(void)
{
if (static_cpu_has(X86_FEATURE_SERIALIZE)) {
serialize();
return;
}
iret_to_self();
}
10.5 Writing through RO mappings
Every store §10.3
walked through — the INT3, the tail bytes, the final first byte — is not
a raw write to .text. Each one is a full call to
text_poke(),
meaning each one pays the entire
§5.4 dance in full: build
the temporary alias in
text_poke_mm,
switch %cr3 onto it, copy through STAC/CLAC, switch back, tear the
alias down. Nothing about these writes being unusually small — as small
as the single INT3 byte in phase 1 — or unusually frequent lets any of
them skip a step.
All of that happens under
text_mutex,
and this is the piece §5.4
leaned on without yet saying where it comes from: text_mutex is what
stops two unrelated patchers — jump labels, static calls, ftrace,
kprobes, the alternatives machinery — from ever building two competing
temporary mappings onto the same text_poke_mm address space at once.
Jump labels layer a second lock,
jump_label_mutex,
on top of that, but the two are not protecting the same thing:
text_mutex serializes individual pokes at the hardware level, while
jump_label_mutex serializes the higher-level operation of enabling or
disabling one whole key — the
enabled
counter update and the walk over the
jump_entry
run for that key, not just the bytes it eventually writes.
The two are also held for deliberately different spans. On the queueing
path (§10.2), text_mutex is
acquired and released once per call to
arch_jump_label_transform_queue()
— bracketing only the
__jump_label_patch()
computation for that one site and its single
smp_text_poke_batch_add()
append. It is free again in between sites, so some unrelated
text_poke() caller elsewhere in the kernel is free to interleave its
own single-site poke while jump labels are still accumulating theirs for
this key. Only once the whole batch is ready does
arch_jump_label_transform_apply()
take text_mutex back and hold it continuously across the entire
three-phase
smp_text_poke_batch_finish()
from §10.3 — that
phase genuinely cannot tolerate a second patcher walking in mid-batch,
since the correctness argument of the INT3 protocol assumes the batch
array it iterates is exactly the one it built. jump_label_mutex, by
contrast, stays held across all of that from the first line of
jump_label_update()
onward — there is no benefit to releasing it early, since doing so would
only let a second, unrelated
static_branch_enable()
call start interleaving its own bookkeeping with that of this one, not
let this one finish any faster.
10.6 End-to-end timeline for one enable
§10.1 through
§10.5 examined the machinery one piece
at a time — which patch path early boot takes, how sites get batched,
the three INT3 phases, what “synchronize” actually does on the wire, and
how each of those writes reaches memory that is nominally read-only.
Laid end to end, a single
static_branch_enable()
call looks like this:
static_branch_enable(&key)
cpus_read_lock()
jump_label_mutex
enabled = -1
jump_label_update(key)
for each jump_entry of key:
__jump_label_patch() # compute nop↔jmp bytes, sanity memcmp
smp_text_poke_batch_add() # append to vector (may flush if full)
smp_text_poke_batch_finish()
arm text_poke_array_refs = 1 on every CPU (smp_wmb)
text_poke INT3 × N
IPI sync # step 1
text_poke tails (bytes 1..N-1) × N
IPI sync # step 2 ("paranoid")
text_poke first bytes × N
IPI sync # step 3
drain: wait for text_poke_array_refs == 0 on every CPU
text_poke_array.nr_entries = 0 # buffer free for next batch
enabled = 1 (release)
unlock…
After this, every previously-nop site for that key is a jmp (or vice
versa), and the hot path behavior has flipped — without any flag load.
Notice what does, and does not, scale with the number of call sites. A
key can have one
jump_entry
or a few hundred, but the number of IPI rounds is fixed at three,4
because
smp_text_poke_batch_finish()
runs its three phases once over the entire batched vector, not once
per site (§10.3).
The only thing that grows that fixed cost is the
§10.2 overflow case: past 256
queued sites, a second full
batch_finish()
round is required, so a tracepoint with, say, 300 call sites pays six
IPI rounds total, not 300 x 3.
11 Modules: the trickiest part
Static keys often live in vmlinux (or module A) while call sites
live in module B — a tracepoint defined in the core kernel, say, with
trace_*() call sites scattered across several drivers loaded as
modules. The picture from §7.1 of
key->entries
as one pointer into one contiguous run of
jump_entry
records only holds while every call site for a key lives in a single
object.
Modules break that assumption in the least convenient way possible: they
load and unload independently of vmlinux and of each other, in an
order nothing can predict ahead of time, so the representation has to be
able to grow and shrink at runtime instead of being settled once at boot
the way §9.1 settles it for the vmlinux-only
case. That is what makes this section “the trickiest part” — every step
has to stay correct no matter how many objects currently contribute to a
key, or in what order they arrived. Over its lifetime, that one
word/pointer union from §7.1 can end up in exactly
three states:
- One home only: every call site for the key lives in the same
single object (
vmlinux, or one module), sokey->entriespoints directly at the contiguous run of that object. This is the fast path, and the common case. - Linked: a second object has registered call sites for the same
key, so one pointer is no longer enough.
key->nextinstead heads a list ofstruct static_key_modnodes, one per contributing object, each pointing at the run that object owns:
/*
next next next
key (LINKED) -----> static_key_mod -----> static_key_mod -----> NULL
(mod A / vmlinux) (mod B)
| |
entries entries
| |
v v
[jump_entry…] [jump_entry…]
*/
struct static_key_mod {
struct static_key_mod *next;
struct jump_entry *entries;
struct module *mod;
};
- Sealed: a key living in
__ro_after_initstorage has its entries pointer permanently forgotten oncejump_label_init_ro()runs (§9.1 covers the mechanics and the bit layout in full). A sealed key never needs linking into anything again, by any module, ever.
Concretely: say the
static_key
of a tracepoint is declared in vmlinux, but nothing in the core kernel
itself calls it — only the driver module A does, loaded first. Module
A becomes the sole contributor for the key, so
key->entries
points straight at the own run of module A, no list involved (the
one-home-only case above, even though the struct of the key lives in
vmlinux while its only call site lives in a module — those are two
independent facts, and only the second one matters here). Module B now
loads and also calls the same tracepoint. A second object just started
contributing, so the key flips into linked mode: one
static_key_mod
node is built to wrap what
key->entries
already pointed at (the run of module A), a second node is built for
module B, and
key->next
now heads that two-node list — exactly the diagram above.
Exactly two functions drive every transition between those three states:
jump_label_add_module()
when a module loads, and
jump_label_del_module()
when one unloads. Neither is called directly — both run from a module
notifier5
(jump_label_module_notify(),
registered with
.priority = 1)
that hooks
MODULE_STATE_COMING/
MODULE_STATE_GOING
— the two transitions the module loader fires while a module is being
mapped in and while it is being torn down.
The priority value is not an arbitrary tie-breaker — it enforces a
strict order. Notifier chains run higher-priority callbacks first, and
while jump labels register at
.priority = 1,
tracepoints register their own, separate module notifier at
.priority = 0
(lower) — so the jump-label notifier always runs first. That ordering
matters because tracepoints are themselves built on static keys. When a
new module loads (MODULE_STATE_COMING), its notifier wants to start
touching the static keys behind its own tracepoints — but those keys are
only ready to be touched once jump_label_add_module() has registered
(and patched where needed), the jump entries of that module. Running the
jump-label notifier first guarantees that the setup is already done by
the time the tracepoint notifier runs.
11.1 Loading: jump_label_add_module()
This is the function that actually produces every transition described
above, once per key contributed by the newly-loaded module. Trimmed of
its early-return-if-empty check, the __init-flag bookkeeping
§9.1 already covered, and its -ENOMEM error
paths, here it is, one piece at a time.
The signature, its local state, and the first thing it does:
static int jump_label_add_module(struct module *mod)
{
struct jump_entry *iter_start = mod->jump_entries;
struct jump_entry *iter_stop = iter_start + mod->num_jump_entries;
struct jump_entry *iter;
struct static_key *key = NULL;
struct static_key_mod *jlm, *jlm2;
jump_label_sort_entries(iter_start, iter_stop);
jump_label_sort_entries()
sorts the jump table of this module — the same routine
§9.1 used on the vmlinux-wide table, for the
same reason. After sorting, all entries that reference the same
static_key
sit next to each other in the array.
That adjacency matters because the work this function does — allocating
static_key_mod
nodes, wiring them into the list, deciding whether to patch — is
per-key, not per-call-site. A module that calls trace_sched_switch()
at ten different places still needs just one list node for that
tracepoint key, not ten. With entries sorted by key, the loop can handle
each distinct key exactly once: run the per-key body when the first
entry for a key appears, then skip all subsequent entries for the same
key with a single pointer comparison.
That skip pattern is what the top of the loop implements:
for (iter = iter_start; iter < iter_stop; iter++) {
struct static_key *iterk = jump_entry_key(iter);
if (iterk == key)
continue;
key = iterk;
The for loop advances iter through every entry in the table, one at
a time. On each iteration,
jump_entry_key()
reads which static_key that entry belongs to. If it is the same key
the loop just finished processing, continue skips the entry — all
per-key work was already done when the first entry of that group was
reached. Only when iterk differs from key does execution fall
through into the per-key body below.
Two consequences follow from that design. First, the per-key body runs
once per distinct key in this module, not once per call site, and each
of the three states from the introduction of
§11 corresponds to exactly one path
through it. Second, iter at the point where the body runs always
points at the first entry of a key group in the sorted table — the start
of a contiguous run. That matters at the end of the function, where
iter is passed as a range start to
__jump_label_update(),
which walks forward from there to patch every call site for that key in
this module.
Once inside the per-key body, the first check asks whether this module even needs to consider the linked-list machinery at all — the one-home-only case:
if (within_module((unsigned long)key, mod)) {
static_key_set_entries(key, iter); /* one home only */
continue;
}
within_module()
asks whether the static_key struct itself — not a call site, not a
jump entry, but the struct that holds
enabled
and
entries
— lives inside memory owned by mod.
If it does, then mod is, by construction, the very first and only
object that has ever contributed call sites for this key. The reason is
physical: before this module loaded, the memory backing that
static_key was not even mapped. No other module or vmlinux could
have built a jump entry referencing an address that did not yet exist.
The one-home-only direct-pointer form is therefore not just adequate
here, it is the only form this key has ever needed — which is why this
branch continues past all the linked-list machinery that follows.
A key that fails that check has call sites outside this module, so before building any list node the sealed case has to be checked:
if (static_key_sealed(key))
goto do_poke; /* sealed: patch once, keep no link */
A sealed key has already forgotten its entries/next union for good
(§9.1), so there is nothing left to link this
module into. Instead the goto skips straight past all the
list-building below, to a single comparison shared with the ordinary
case — reached down in the text.
An ordinary, unsealed key that reaches this point is the linked case. This is where the list from the introduction of §11 actually gets built or grown, in two steps.
The first step handles a one-time transition. Up to this point the key
might still be in direct-pointer form — one pointer, one contributing
object, no list. Before anything can be prepended to it, that existing
pointer has to be wrapped in a list node so there is a next field to
link through:
jlm = kzalloc(sizeof(struct static_key_mod), GFP_KERNEL);
if (!static_key_linked(key)) {
jlm2 = kzalloc(sizeof(struct static_key_mod), GFP_KERNEL);
scoped_guard(rcu)
jlm2->mod = __module_address((unsigned long)key);
jlm2->entries = static_key_entries(key);
static_key_set_mod(key, jlm2);
static_key_set_linked(key);
}
static_key_linked()
returns false only when the key is still in direct-pointer form, meaning
exactly one object has contributed to it so far. The code allocates
jlm2 to retroactively wrap whatever
key->entries
already pointed at — the run belonging to that first object — discovers
which module owns that object via
__module_address(),
and installs jlm2 as the head of a new one-node list. Once
static_key_set_linked()
flips the linked bit, this wrapping step never runs again for this key —
every later module finds static_key_linked() already true and skips
straight into the second step.
The second step prepends a node for the newly-arriving module onto the now-guaranteed-to-exist list:
jlm->mod = mod;
jlm->entries = iter;
jlm->next = static_key_mod(key);
static_key_set_mod(key, jlm);
static_key_set_linked(key);
jlm->entries is set to iter — the first entry of this key group in
the sorted table, the same pointer the one-home-only case would have
stored directly into
key->entries.
static_key_set_mod()
makes jlm the new head of the list, with the previous head linked
behind it via jlm->next.
Both the sealed-key
goto
and the ordinary path above fall into the same final check, once per
key:
do_poke:
if (jump_label_type(iter) != jump_label_init_type(iter))
__jump_label_update(key, iter, iter_stop, true);
}
return 0;
}
Every call site a module ships with is compiled around one fixed
assumption: the default value the key had at compile time — the same
type ^ branch formula from
§6.1, exposed here as
jump_label_init_type()
— which decided whether the assembler macro emitted this site as a nop
or a jmp in the first place (§6.2),
frozen into the module from then on. But the live jump label type
value can have moved away from that compiled-in default already, if
something else called
static_branch_enable()/disable()
on this key before the module ever loaded.
do_poke
catches exactly that mismatch and patches the new sites immediately,
before the code of the module has a chance to run and observe them in
the wrong state — whether it arrived there via the sealed-key goto
above, or by falling through normally after linking the key into the
list.
11.2 Toggling an already-linked key
When someone calls
static_branch_enable()/static_branch_disable()/static_branch_inc()/static_branch_dec()
on a key whose call sites span more than one object, the flat array scan
from §9.5 is not enough — the
sites are scattered across separate per-object tables, each with its own
bounds, and the patcher has to visit every one of them.
jump_label_update()
detects that case: if
static_key_linked()
returns true, it hands off to
__jump_label_mod_update()
instead of doing the walk itself. That function walks the linked list,
calling
__jump_label_update()
once per node:
static void __jump_label_mod_update(struct static_key *key)
{
struct static_key_mod *mod;
for (mod = static_key_mod(key); mod; mod = mod->next) {
struct jump_entry *stop;
struct module *m;
if (!mod->entries)
continue;
m = mod->mod;
if (!m)
stop = __stop___jump_table;
else
stop = m->jump_entries + m->num_jump_entries;
__jump_label_update(key, mod->entries, stop,
m && m->state == MODULE_STATE_COMING);
}
}
The loop visits each
static_key_mod
node in the list and passes that entries pointer and matching stop
bound to __jump_label_update(). Three details in that loop deserve
explanation.
First, stop cannot be a single kernel-wide constant the way it is in
the flat walk from §9.5. Each
module has its own private
__jump_table
section (§7.4), bounded by the
jump_entries/num_jump_entries
fields of that module. Without the correct per-node bound,
__jump_label_update() would walk forward past the end of a module
table into unrelated memory. The loop has to look up that bound for each
node individually.
Second, m — the
mod
field of the list node, not the function parameter — can be NULL. That
is not an error: it is the node that represents vmlinux itself. The
NULL originates in §11.1: the
first-time linking step calls
__module_address()
to discover which module owns the key, and __module_address() returns
NULL for addresses inside the core kernel. For that node the correct
bound is the global
__stop___jump_table,
which marks the end of the vmlinux-wide table.
Third,
mod->entries
can genuinely be NULL. That happens when an object defines a key but
contains no call sites for it — the linking code from
§11.1 still creates a list node to
represent that object (it is a contributor), but there is nothing
there to patch. Concretely: a module exports a
DEFINE_STATIC_KEY_FALSE,
and call sites in other modules are the only consumers. Skipping the
node is correct.
The final argument, m && m->state == MODULE_STATE_COMING, tells
__jump_label_update() whether to also patch entries living in the
__init section of the module. A module still in
MODULE_STATE_COMING
has not finished running its init function yet, so its init section is
still mapped and its call sites there are reachable. A module already in
MODULE_STATE_LIVE
has discarded that section — patching into freed memory would be a
use-after-free, not a harmless no-op.
11.3 Unloading: jump_label_del_module()
jump_label_del_module() is the mirror image of
jump_label_add_module(),
run on
MODULE_STATE_GOING:
for each distinct key this module contributes to, it finds and removes
the
static_key_mod
node that §11.1 created. The
text-patching machinery does not need to undo anything here — the
.text section of the module is about to be unmapped entirely, so the
patchable sites in it simply cease to exist. What does need cleaning
up is the linked-list bookkeeping that still references them.
The function uses the same sorted-table, skip-by-key loop structure as loading. Three of the four skip cases mirror §11.1 directly:
for (iter = iter_start; iter < iter_stop; iter++) {
if (jump_entry_key(iter) == key)
continue;
key = jump_entry_key(iter);
if (within_module((unsigned long)key, mod))
continue;
/* No @jlm allocated because key was sealed at init. */
if (static_key_sealed(key))
continue;
/* No memory during module load */
if (WARN_ON(!static_key_linked(key)))
continue;
The first three are familiar from
§11.1: skip duplicate entries for the
same key (the sorted-table grouping from the sort step), skip keys whose
struct lives inside this module (the struct vanishes with the module, so
there is no list to update), and skip sealed keys (no node was ever
created for them). The fourth is a defensive check: if the key is not in
the linked state at this point, something went wrong during loading —
likely an allocation failure that jump_label_add_module() could not
recover from. The
WARN_ON
flags the inconsistency without crashing6, and the continue skips
the key rather than dereferencing a pointer that was never set up.
For every key that passes all four checks, the function walks the linked list to find and splice out the node belonging to this module:
prev = &key->next;
jlm = static_key_mod(key);
while (jlm && jlm->mod != mod) {
prev = &jlm->next;
jlm = jlm->next;
}
/* No memory during module load */
if (WARN_ON(!jlm))
continue;
if (prev == &key->next)
static_key_set_mod(key, jlm->next);
else
*prev = jlm->next;
kfree(jlm);
The while loop advances through the list until it finds the node whose
mod
field matches the departing module, keeping prev pointed at the next
pointer of the preceding node so the splice has something to patch. If
no matching node is found — again a sign that something went wrong
during loading — a second WARN_ON fires and the key is skipped.
Otherwise, the standard singly-linked-list splice removes the node: if
it was the head of the list (prev == &key->next), the next pointer
of the key itself is updated via
static_key_set_mod();
if it was somewhere in the middle, the next pointer of the predecessor
is patched directly.
After the splice, one more step checks whether the list can be eliminated entirely:
jlm = static_key_mod(key);
/* if only one etry is left, fold it back into the static_key */
if (jlm->next == NULL) {
static_key_set_entries(key, jlm->entries);
static_key_clear_linked(key);
kfree(jlm);
}
If exactly one node remains, the key no longer needs the list form —
only one object still contributes to it. The code folds the entries
pointer of that last node back into
key->entries
directly, clears the
LINKED
bit, and frees the node. This is the linked-state transition from
§11.1 running in reverse: a key that
needed the list form only because two objects happened to overlap
returns to the one-home-only direct-pointer form the moment that overlap
ends.
12 Fallback: CONFIG_JUMP_LABEL=n
JUMP_LABEL
is optional. Most distro kernels end up with it on — arm64 selects it
outright, and on x86
PREEMPT_DYNAMIC
pulls it in — but a minimal config can legitimately leave it off.
Everything from §5 onward
assumed the option was enabled. This section covers what happens when it
is not.
Without CONFIG_JUMP_LABEL, the struct shrinks to a bare counter. The
entries/next/type union from §7 is
compiled out entirely:
struct static_key {
atomic_t enabled;
};
jump_label_init() sets
static_key_initialized
to true and returns. There is no jump table to sort, no entries to
pre-patch.
The call-site macros turn into ordinary branch-hinted conditionals.
static_branch_likely() and static_branch_unlikely() reduce to:
#define static_branch_likely(x) likely_notrace(static_key_enabled(&(x)->key))
#define static_branch_unlikely(x) unlikely_notrace(static_key_enabled(&(x)->key))
No asm goto, no jump table, no patching — just a
likely_notrace()/unlikely_notrace()
hint around a read of enabled.
static_key_count() is a plain
raw_atomic_read():
static __always_inline int static_key_count(struct static_key *key)
{
return raw_atomic_read(&key->enabled);
}
Compare this with the CONFIG_JUMP_LABEL=y version in
kernel/jump_label.c,
which clamps negative values (n >= 0 ? n : 1). That clamp exists
because
static_key_enable()
temporarily sets enabled to -1 while the patching pass runs (see
§9.2). Without
patching, enabled never goes negative, so the clamp is unnecessary.
static_key_enable() and static_key_disable() are simpler for the
same reason. Each one checks whether enabled already holds the target
value and returns early if so. If it holds something unexpected (neither
0 nor 1), a
WARN_ON_ONCE
fires. Otherwise, a plain atomic_set() writes the new value. No
cmpxchg, no intermediate -1, no
jump_label_update()
call.
jump_label_lock()/jump_label_unlock()
are empty stubs — there is no patch pass to serialize.
The net effect on every hot path is exactly the cost
§2 opened with: a memory load of
enabled,
a compare, and a conditional branch. The
likely()/unlikely()
hint steers the branch predictor the same way a compiled-in NOP or JMP
would, but it cannot eliminate the branch itself — that is the
optimization that CONFIG_JUMP_LABEL=y adds.
Jump labels are an optimization, not a correctness feature: behavior matches; only the mechanism changes.
13 Worked micro-example (bytes on the wire)
Every mechanism described so far — the asm helpers, the jump-table
entry, the objtool hack, the size-discovery check, the
enabled
state machine, and the INT3 protocol — touches one concrete call site at
some point in its life. This section traces a single, minimal site
through that entire life, from compile time through one enable and one
disable, close enough to see the actual instruction bytes change rather
than just the names of the steps that change them. Suppose:
DEFINE_STATIC_KEY_FALSE(k);
void f(void)
{
if (static_branch_unlikely(&k))
printk("on\n");
something();
}
Two small assumptions turn this from a symbolic description into an actual trace, and neither changes anything about how the mechanism works, only which specific numbers show up:
f()is called after boot, on a live, multi-CPU system — the interesting §10 INT3 path, not the single-CPU §10.1 boot shortcut (a boot-time toggle of this same key would usetext_poke_early()instead, with no INT3 involved at all).- The compiler places
l_yes— the out-of-line block containing theprintk()call — 80 bytes past the end of the patch site. 80 fits in a signed byte, so the build-time trick from §6.5 picks the 2-byteJMP rel8encoding rather than the 5-byterel32form. A fartherl_yeswould just mean 5 bytes instead of 2 everywhere below (§5.2); nothing else about the trace would change.
13.1 Compile / link / objtool (HAVE_JUMP_LABEL_HACK)
From the source code above, the compiler, linker, and objtool produce
the bytes that sit in vmlinux at the patch site. For orientation, here
is where the three pieces end up — the patch site in .text, the
out-of-line target, and the sidecar entry in __jump_table:
.text (function f) __jump_table (one entry)
────────────────────────── ──────────────────────────
... code: delta to 1:
1: [ 2 bytes ] patch site target: delta to l_yes
... key: delta to &k.key + 2
call something bit 0 = 0 (branch)
ret bit 1 = 1 (objtool)
...
l_yes: 80 bytes past 1:+2
call printk
jmp back ------> (after 1:+2)
The six steps below trace how those bytes arrive at their final state:
-
kis aFALSEkey read withstatic_branch_unlikely(). Per the table in §6.2, that combination calls the nop-default helperarch_static_branch(&k.key, false), not the jmp-defaultarch_static_branch_jump(). That is thetype ^ branchformula from §6.1 at work:type = 0(FALSE),branch = 0(unlikely), sotype ^ branch = 0— the hint agrees with the default, and the nop-default path gets chosen. -
The
asm gotoinsidearch_static_branch()emits1: jmp l_yesat the patch site, plus one raw row in__jump_tableviaJUMP_TABLE_ENTRY()(§6.4):codeis the self-relative distance to1:,targetis the self-relative distance tol_yes, andkeyis the self-relative distance to&k.key + 0 + 2.That
+ 2puts a1in bit 1 of the storedkeyaddress. Bit 0 isbranch— here0, matching theunlikely()hint. Bit 1 is the build-time signal toobjtool: “NOP this jmp” (§6.5); after boot,jump_entry_set_init()repurposes this same bit as the__init-text flag. At runtime,jump_entry_key()masks both bits off to recover the real address ofk. -
The assembler picks the encoding by the real distance to
l_yes— here, 80 bytes forward (assumption 2 above), well inside the -128..+127 reach of a signed byte. It emits the 2-byteJMP rel8form: opcodeEB, followed bydisp = dest - (addr + insn_size) = 80 = 0x50(the disp formula from §5.2). The two bytes actually sitting at1:right after assembly, beforeobjtoolever runs, areEB 50. -
During the build,
objtoolcallshandle_jump_alt()(§6.5 step 3), which sees bit 1 set in the storedkeyoperand and rewrites those exact two bytes, in place, from thejmp(EB 50) to the 2-byte NOP (66 90, the NOP encoding from §5.2) — same size, so nothing around the site shifts. By the timevmlinuxis linked, the live bytes at1:are already66 90, andobjtoolitself is long gone. -
At boot,
jump_label_init()sorts__jump_tableby key, then loops over every entry (§9.1). For the one entry belonging tok, the loop body does the following:jump_label_init() ├── jump_label_sort_entries() sort __jump_table by key └── for each entry: (k has exactly one) ├── jump_label_type() = NOP enabled(0) ^ branch(0) │ └── arch_jump_label_transform_static() no-op on x86 ├── jump_entry_set_init(entry, false) code not in __init └── static_key_set_entries(&k, entry) wires k.entries → entryjump_label_type()computesenabled(0) ^ branch(0) = NOP— the expected state matches the live bytes, which are already66 90. Soarch_jump_label_transform_static()is a genuine no-op: x86 never overrides the generic fallback, whose entire body is one comment,/* nothing to do on most architectures */. No instruction bytes get rewritten at this site. -
The same loop iteration does two pieces of bookkeeping that wire
kinto the runtime data structures. First,jump_entry_set_init()checks whether the code at1:lives in__inittext — it does not, so bit 1 of the storedkeyaddress (the same bitobjtoolused in step 2) gets cleared to0. Second,static_key_set_entries()pointsk.entriesat this entry — the pointer thatjump_label_update()will follow whenstatic_branch_enable(&k)runs later. With that in place, the hot path inf()from the very first time it runs is: decode66 90(falls through, no load ofk, §5.1), thencall something.
In summary, the same two bytes at 1: passed through three stages
before the kernel ever ran a line of f():
Stage Bytes at 1: Why
───────────────── ─────────── ──────────────────────────────────
After assembly EB 50 (JMP) assembler picks rel8 for +80 distance
After objtool 66 90 (NOP) handle_jump_alt() sees bit 1 in key
At boot 66 90 (NOP) jump_label_init(): NOP expected, NOP found
13.2 Runtime static_branch_enable(&k)
§10.6 already lays out the full
call chain a batch of sites goes through on enable; k has exactly one
entry, so this is that same chain with the actual bytes for every step
filled in. Both directions of the public API are one-line macros in
include/linux/jump_label.h:
#define static_branch_enable(x) static_key_enable(&(x)->key)
#define static_branch_disable(x) static_key_disable(&(x)->key)
The full call chain for this one enable, with the concrete byte values
for k filled in at each level:
static_branch_enable(&k)
└─ static_key_enable(&k.key)
├── cpus_read_lock / jump_label_lock
├── enabled: 0 --> -1 callers already see "on"
├── jump_label_update(key)
│ └─ __jump_label_update()
│ ├── type = enabled(true) ^ branch(0) = JMP
│ ├── arch_jump_label_transform_queue()
│ │ └─ __jump_label_patch(entry, JMP)
│ │ ├── size = 2 (live-decode of 66 90)
│ │ ├── code = EB 50 (text_gen_insn)
│ │ ├── nop = 66 90 (x86_nops[2])
│ │ ├── memcmp(addr, nop) -- pre-flight OK
│ │ └── smp_text_poke_batch_add(addr, EB 50, 2)
│ └── arch_jump_label_transform_apply()
│ └─ smp_text_poke_batch_finish()
│ └── INT3 three-phase: 66 90 --> EB 50
├── enabled: -1 --> 1 atomic_set_release
└── jump_label_unlock / cpus_read_unlock
The numbered steps below walk through this chain in detail:
static_branch_enable(&k)expands tostatic_key_enable(&k.key), which acquiresjump_label_lock()and walks the0 → -1 → (patch) → 1state machine from §9.2.enabledstarts at0, gets set to-1(“enabling in progress”), andstatic_key_count()already reports-1as “on” — so no caller sees a false “off” window while patching runs. Thenjump_label_update(&k.key)does the actual patching, and only after it returns doesenabledget published as1with release ordering.jump_label_update()finds the one entry ofkand computesjump_label_type(entry)=enabled ^ branch.enabledis the transient-1from step 1, whichstatic_key_enabled()reports astrue;branchis the stored hint bit,0.true ^ false = JMP— the live nop becomes a jmp.arch_jump_label_transform_queue()calls__jump_label_patch(), which re-derives the size by decoding the live bytes at1:—arch_jump_entry_size()returns 2, matching whatobjtoolleft behind (§8). The function then builds both sequences:nop = 66 90(fromx86_nops[2]) andcode = EB 50(fromtext_gen_insn()— the same bytes computed in §13.1 step 3, sinceaddranddesthave not moved). Because this is a nop→jmp transition, it callsmemcmp()to verify the live bytes are currently66 90; a mismatch would be aBUG(). They match, so the patch66 90→EB 50gets queued viasmp_text_poke_batch_add()(§10.2).- The loop of
jump_label_update()over the entries ofkends here — there is only the one — soarch_jump_label_transform_apply()immediately callssmp_text_poke_batch_finish(), which runs the three-phase protocol from §10.3 on this one queued site, now with the real two bytes instead of a placeholder:
start (before) 66 90 any fetch: executes the 2-byte NOP
phase 1 (INT3 in) cc 90 any fetch: #BP -> handler emulates
^^ the pending JMP (jumps to l_yes)
trap byte
-------- IPI sync --------
phase 2 (tail in) cc 50 same as phase 1: byte 0 still
^^ traps and gets emulated — byte 1
still traps is now its final value underneath
-------- IPI sync --------
phase 3 (done) eb 50 any fetch: executes the real
^^ real opcode JMP rel8 directly, no trap needed
-------- IPI sync --------
- From the instant the sync in phase 1 completes, any CPU landing on
this address already gets the effect of the jump via emulation
(§10.3);
phases 2-3 only retire the trap-and-emulate path in favor of the
real bytes. Once phase 3 lands, every subsequent call to
f()decodesEB 50, jumps 80 bytes forward intol_yes, runsprintk("on\n"), then hits the compiler-emittedjmp back(§3) and falls intosomething().
13.3 static_branch_disable(&k): same protocol, asymmetric math
Disabling runs the mirror call,
static_branch_disable(&k)
→
static_key_disable(&k.key)
— a single
atomic_cmpxchg(&key->enabled, 1, 0)
instead of the enable state machine, because disable has no in-progress
state to protect (§9.3: a reader who sees stale “on” for a
few more instructions is exactly the harmless case, unlike stale “off”).
If that cmpxchg succeeds,
jump_label_update()
runs again, this time computing
jump_label_type()
= false ^ false = NOP
(enabled
now 0, branch still 0).
__jump_label_patch()
rebuilds the same two sequences as before — nop = 66 90,
code = EB 50, both unchanged, since addr, dest, and size haven’t
moved — but this time expects the live bytes to be the jmp and installs
the nop. That is the one genuine asymmetry in this whole worked example:
enabling always needs a fresh, target-specific displacement computed by
text_gen_insn();
disabling never does, because the encoding of a NOP does not depend on
where the branch would have gone — the exact same fixed bytes go back
every time a site of this size is disabled, no matter what it was
jumping to.
The same three-phase protocol (§10.3) runs again, in the opposite byte direction:
start (before) eb 50 executes the JMP rel8
phase 1 (INT3 in) cc 50 #BP -> handler now emulates the
^^ pending NOP (a JMP with disp == 0
trap byte — not a special case)
-------- IPI sync --------
phase 2 (tail in) cc 90 byte 1 now its final value; byte 0
^^ still traps
still traps
-------- IPI sync --------
phase 3 (done) 66 90 executes the real NOP directly
^^ real opcode
-------- IPI sync --------
Put together, the build-time 66 90 of this one site, the EB 50 of
the first enable, and the 66 90 of this disable again are the entire
lifecycle that
§5-§10
spent this whole tutorial describing in the abstract — the same two
bytes, chosen and re-derived by a different mechanism at each stage, but
never touched by anything other than the three sanctioned writers: the
assembler once at compile time, objtool once at build time, and
__jump_label_patch() through the protocol from
§10.3 every time
after that.
| Stage | Live bytes at 1: |
Who wrote them |
|---|---|---|
After assembly, before objtool |
EB 50 |
Compiler/assembler (§6.3, §5.2) |
After objtool, at boot |
66 90 |
handle_jump_alt() (§6.5) |
After static_branch_enable(&k) |
EB 50 |
__jump_label_patch() via INT3 protocol (§10.3) |
After static_branch_disable(&k) |
66 90 |
Same, reverse direction |
14 Further reading in-tree
- Jump label (LWN, 2010) — the original write-up of the idea.
Documentation/staging/static-keys.rst— historical overview and oldgetppidinstruction dump (still useful; ignore fixed “always 5 bytes” claims for modern x86).arch/x86/kernel/alternative.c—smp_text_poke_batch_finishcomment block (canonical INT3 algorithm write-up).tools/objtool/check.c—handle_jump_alt()for the jmp→nop hack.include/linux/jump_label.h— the top comment in the header: deprecated vs. current API, behavioral model, the “absolute slow paths” warning, deferred-decrement rationale.include/linux/tracepoint-defs.h—struct tracepointembeds both astatic_key_falseand astatic_call_key— the largest real-world consumer of static keys, and the intersection of both sister mechanisms.- Sister mechanisms with the same poke engine: static
calls,
ftrace,
kprobes
— same
text_poke/ INT3 infrastructure, different metadata sections.
-
A 4-byte signed displacement covers a range of ±2 GiB. Because x86_64 kernels are compiled using the
-mcmodel=kerneloption, the entire kernel image is restricted to the upper 2 GiB of the address space. This model guarantees that any relative offset within the kernel binary falls within the reach of arel32displacement. ↩ -
__ro_after_initis a section attribute (include/linux/cache.h,__section(".data..ro_after_init")) for data that is written during boot but never again afterward — unlikeconst, which the compiler must be able to enforce at compile time, this is a promise the author makes about runtime behavior. The kernel makes that promise real atmark_rodata_ro()time, when the whole.data..ro_after_initsection is remapped read-only in the page tables, so any later write attempt — a bug, or an author breaking their own promise — faults instead of silently corrupting state. ↩ -
_stextis a linker-defined symbol, not a C variable — it marks the address where the.textsection of the kernel begins, set directly inarch/x86/kernel/vmlinux.lds.S. Every function in core kernel text lives at some fixed offset from it, which is what letssmp_text_poke_loc.rel_addr(a plains32) address any patch site with 4 bytes instead of the full 8-byte pointerjump_entry.keyneeds (§6.4) for its potentially-far-awaystatic_key. ↩ -
Strictly, phases 2 and 3 each sync only if at least one site in the batch actually needed a write in that phase, tracked by a
do_synccounter insidesmp_text_poke_batch_finish(). The phase 3 sync, for instance, is skipped if the final first byte of every site already happens to equalINT3. An ordinary nop-to-jmp toggle always writes something in both phases, so three rounds is what actually happens in practice; the fixed count just isn’t unconditional at the code level. ↩ -
A module notifier is a callback registered against the module-load notifier chain (
module_notify_list) viaregister_module_notifier(). The module loader walks that chain withblocking_notifier_call_chain()at each state transition (MODULE_STATE_COMING,MODULE_STATE_GOING, etc.), invoking every registeredstruct notifier_blockin priority order — this is the generic mechanism subsystems use to react to modules loading/unloading, not something jump labels invented. ↩ -
Unless
panic_on_warn = 1↩