Skip to content

Persist z_seq across znode eviction - #18573

Closed
ixhamza wants to merge 3 commits into
openzfs:masterfrom
truenas:persist_znode_across_eviction
Closed

ixhamza wants to merge 3 commits into
openzfs:masterfrom
truenas:persist_znode_across_eviction

Conversation

@ixhamza

@ixhamza ixhamza commented May 21, 2026

Copy link
Copy Markdown
Member

Motivation and Context

Commit 312bdab advertises STATX_ATTR_CHANGE_MONOTONIC to knfsd and builds the NFSv4 change_cookie from (ctime.tv_sec << 32) | zp->z_seq. zp->z_seq is reset to a magic constant in zfs_znode_alloc(), so any event that drops the znode from cache (memory pressure, remount, reboot) brings the file back with the same ctime.tv_sec upper bits but a smaller z_seq in the lower bits, regressing the cookie within the same second.

NFSv4 clients that trust the monotonicity contract treat this as metadata they cannot rely on. VMware ESXi over NFSv4.1 reliably reproduces it with The file specified is not a virtual disk, causing a VM stored on the affected ZFS dataset to fail to power on.

Description

Persist zp->z_seq via a new SA attribute SA_ZPL_SEQ so it survives znode eviction. A new pflag bit ZFS_HAS_SEQ marks the file as carrying SA_ZPL_SEQ in its layout, mirroring the existing ZPL_PROJID/ZFS_PROJID pattern. The bit gates may_grow at SA tx-hold sites, choosing B_TRUE on the first add per file and B_FALSE thereafter, so steady-state operations pay no extra reservation.

A ZFS_PERSIST_SEQ() macro captures z_seq and sets the bit into the caller's bulk in one step, persisting both atomically alongside the file's other SA attributes. Every site that bumps z_seq uses it. zfs_znode_alloc() restores z_seq from SA_ZPL_SEQ when the bit is set.

No on-disk format change requiring a feature flag is needed. Older binaries preserve the new attribute and bit opaquely. The first modify by a patched binary lazily migrates each file.

How Has This Been Tested?

  • Before: ESXi VM fails to power on over NFSv4.1.
  • After: VM powers on successfully.
  • CI Testing

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Performance enhancement (non-breaking change which improves efficiency)
  • Code cleanup (non-breaking change which makes code smaller or more readable)
  • Quality assurance (non-breaking change which makes the code more robust against bugs)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Library ABI change (libzfs, libzfs_core, libnvpair, libuutil and libzfsbootenv)
  • Documentation (a change to man pages or other documentation)

Checklist:

Comment thread include/sys/zfs_znode.h Outdated
@ixhamza
ixhamza force-pushed the persist_znode_across_eviction branch from eeb4661 to 56245e3 Compare June 1, 2026 21:14
@behlendorf
behlendorf self-requested a review June 1, 2026 22:03
@ixhamza
ixhamza force-pushed the persist_znode_across_eviction branch from 56245e3 to 3b204af Compare June 2, 2026 21:28
Comment thread include/sys/zfs_znode.h
Comment thread module/zfs/zfs_vnops.c
Comment thread module/zfs/zfs_vnops.c Outdated
@ixhamza
ixhamza force-pushed the persist_znode_across_eviction branch 2 times, most recently from 95dc629 to c7dbebd Compare June 5, 2026 11:47
@ixhamza

ixhamza commented Jun 5, 2026

Copy link
Copy Markdown
Member Author

Updated per @amotin's private feedback. Rebased onto master.

Comment thread module/os/freebsd/zfs/zfs_znode_os.c Outdated
Comment thread module/os/linux/zfs/zfs_vnops_os.c
@ixhamza
ixhamza force-pushed the persist_znode_across_eviction branch 3 times, most recently from a695cd9 to 28c1e17 Compare June 8, 2026 19:07
@ixhamza
ixhamza requested a review from amotin June 10, 2026 13:58

@robn robn left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sort of a low-form of drive-by review from me (ride-by review?).

@ixhamza ran this by me early and I said "looks plausible, and SAs have the nice property of being forward and backward compatible if assembled sensibly". I haven't looked since, but have reread through it now, and I still think that, and I trust from the review comments and the diffs that its been thought about in enough detail that I don't have to think it all through from first principles. (but I will if someone tells me to).

One thought though, have we actually tested that things still work right (fsvo) when importing a pool that has run on this change back into an older ZFS that hasn't? And/or, Linux<->FreeBSD? I mean, that's weird, and coming from cold I don't expect it'd make much different for NFS clients so long as the number doesn't go backwards. Just so long as we're not inadvertently causing a pool not to import or anything. I think its fine, but doesn't hurt to ask.

(Aside: SA API always feel so difficult...)

@ixhamza

ixhamza commented Jun 11, 2026 •

Copy link
Copy Markdown
Member Author

Thanks @robn for taking a look.

One thought though, have we actually tested that things still work right (fsvo) when importing a pool that has run on this change back into an older ZFS that hasn't? And/or, Linux<->FreeBSD? I mean, that's weird, and coming from cold I don't expect it'd make much different for NFS clients so long as the number doesn't go backwards. Just so long as we're not inadvertently causing a pool not to import or anything. I think its fine, but doesn't hurt to ask.

Linux<->FreeBSD imports cleanly both ways. Older <-> Newer ZFS imports fine too since SA_ZPL_SEQ is just an opaque SA attribute.

@ixhamza
ixhamza force-pushed the persist_znode_across_eviction branch from 28c1e17 to 8f193da Compare June 15, 2026 12:16
@ixhamza

ixhamza commented Jun 15, 2026 •

Copy link
Copy Markdown
Member Author

Pushed a follow up commit covering a few additional modification sites where z_seq wasn't being bumped, surfaced while testing more thoroughly. Since the NFSv4 change_cookie now relies entirely on z_seq, wanted to make sure every modification path advances it. Rebased onto current master.

@ixhamza
ixhamza requested review from amotin and robn June 15, 2026 12:19

@robn robn left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bunch of comments, but all of the kind I would point at in a pair review and say "hmm, what's that about?" and you'll likely tell me and I'll learn something new (man that'd be SO MUCH FUN).

Assuming you know what you're doing and approving on that basis. This is a proper tour de force, nice work!

Comment thread module/os/linux/zfs/zpl_file.c
Comment thread module/os/linux/zfs/zpl_file.c Outdated
Comment thread module/os/linux/zfs/zpl_file.c
Comment thread module/os/linux/zfs/zpl_file.c
Comment thread module/os/linux/zfs/zpl_inode.c Outdated
Comment thread module/os/linux/zfs/zpl_xattr.c Outdated
Comment thread module/os/linux/zfs/zfs_znode_os.c
@ixhamza
ixhamza force-pushed the persist_znode_across_eviction branch from 8f193da to 7d01e1b Compare June 24, 2026 19:09
Comment thread module/os/freebsd/zfs/zfs_vnops_os.c Outdated
@ixhamza
ixhamza force-pushed the persist_znode_across_eviction branch from 7d01e1b to bc5865e Compare June 25, 2026 11:58
@behlendorf behlendorf added Status: Accepted Ready to integrate (reviewed, tested) and removed Status: Code Review Needed Ready for review and testing labels Jun 25, 2026
behlendorf pushed a commit that referenced this pull request Jun 25, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes #18573
behlendorf pushed a commit that referenced this pull request Jun 25, 2026
The SA bulk arrays are filled by series of conditional SA_ADD_BULK_ATTR
calls whose worst-case count is easy to miscount, so a future attribute
add could silently overrun the fixed-size array.
Add ASSERT3S(count, <=, ARRAY_SIZE(bulk)) before each sa_bulk_update so
an overrun trips in debug builds, and use ARRAY_SIZE so the bound stays
tied to the declaration. Linux zfs_setattr allocates its arrays on the
heap sized by 'bulks', so it asserts against that instead.
Also tighten FreeBSD zfs_setattr's xattr_bulk from [7] to [6], its
actual worst case.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes #18573
bugclerk pushed a commit to truenas/zfs that referenced this pull request Jun 30, 2026
Commit 312bdab advertises STATX_ATTR_CHANGE_MONOTONIC and builds
the NFSv4 change_cookie from (ctime.tv_sec << 32) | zp->z_seq.
zp->z_seq is reset to a magic constant in zfs_znode_alloc(), so any
event that drops the znode from cache (memory pressure, remount,
reboot) regresses the lower bits of the cookie, a backward step
within the same second.

NFSv4 clients that trust this contract treat a regressed cookie as
evidence that the file's metadata cannot be relied on. VMware ESXi
over NFSv4.1 surfaces this as "The file specified is not a virtual
disk", and a VM stored on the affected NFS-exported ZFS dataset
fails to power on.

Widen z_seq to 64 bit and present it directly as the change_cookie,
dropping the ctime packing, so the cookie is a single monotonic
counter that no longer depends on the clock. FreeBSD's va_filerev
consumer also takes the wider value.

Persist z_seq via a new SA attribute SA_ZPL_SEQ. An in-core marker
zp->z_has_seq records whether the file already carries SA_ZPL_SEQ in
its layout; it is derived at load time and never stored on disk, so
no global pflag bit is consumed. ZFS_SEQ_MAY_GROW() keys off the
marker to grow the SA layout only on the first add per file;
ZFS_PERSIST_SEQ() then sets the marker and adds SEQ to the caller's
bulk alongside the file's other SA attributes. zfs_znode_alloc()
restores z_seq from SA_ZPL_SEQ when present and sets the marker;
zfs_rezget() recomputes the marker in place on rollback/recv without
disturbing the in-core z_seq, keeping the cookie monotonic.

A file written before this change carries no SA_ZPL_SEQ; on Linux it
is seeded with (ctime.tv_sec + 1) << 32 so the counter starts above
any pre-change cookie and stays monotonic across the upgrade. A
missing attribute is simply treated as not-yet-migrated, not an
error. FreeBSD never folded ctime into va_filerev, so it needs no
seed.

No feature flag or on-disk format change is needed: the new SA
attribute is keyed by name, so an implementation that does not know
it preserves it opaquely, and the first modify lazily migrates each
file. Covers both the Linux and FreeBSD ZPL.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
(cherry picked from commit 8c5ac8e)
bugclerk pushed a commit to truenas/zfs that referenced this pull request Jun 30, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
(cherry picked from commit 32c1d94)
bugclerk pushed a commit to truenas/zfs that referenced this pull request Jun 30, 2026
The SA bulk arrays are filled by series of conditional SA_ADD_BULK_ATTR
calls whose worst-case count is easy to miscount, so a future attribute
add could silently overrun the fixed-size array.
Add ASSERT3S(count, <=, ARRAY_SIZE(bulk)) before each sa_bulk_update so
an overrun trips in debug builds, and use ARRAY_SIZE so the bound stays
tied to the declaration. Linux zfs_setattr allocates its arrays on the
heap sized by 'bulks', so it asserts against that instead.
Also tighten FreeBSD zfs_setattr's xattr_bulk from [7] to [6], its
actual worst case.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
(cherry picked from commit 2d6b51c)
lundman pushed a commit to openzfsonosx/openzfs-fork that referenced this pull request Jul 30, 2026
Commit 312bdab advertises STATX_ATTR_CHANGE_MONOTONIC and builds
the NFSv4 change_cookie from (ctime.tv_sec << 32) | zp->z_seq.
zp->z_seq is reset to a magic constant in zfs_znode_alloc(), so any
event that drops the znode from cache (memory pressure, remount,
reboot) regresses the lower bits of the cookie, a backward step
within the same second.

NFSv4 clients that trust this contract treat a regressed cookie as
evidence that the file's metadata cannot be relied on. VMware ESXi
over NFSv4.1 surfaces this as "The file specified is not a virtual
disk", and a VM stored on the affected NFS-exported ZFS dataset
fails to power on.

Widen z_seq to 64 bit and present it directly as the change_cookie,
dropping the ctime packing, so the cookie is a single monotonic
counter that no longer depends on the clock. FreeBSD's va_filerev
consumer also takes the wider value.

Persist z_seq via a new SA attribute SA_ZPL_SEQ. An in-core marker
zp->z_has_seq records whether the file already carries SA_ZPL_SEQ in
its layout; it is derived at load time and never stored on disk, so
no global pflag bit is consumed. ZFS_SEQ_MAY_GROW() keys off the
marker to grow the SA layout only on the first add per file;
ZFS_PERSIST_SEQ() then sets the marker and adds SEQ to the caller's
bulk alongside the file's other SA attributes. zfs_znode_alloc()
restores z_seq from SA_ZPL_SEQ when present and sets the marker;
zfs_rezget() recomputes the marker in place on rollback/recv without
disturbing the in-core z_seq, keeping the cookie monotonic.

A file written before this change carries no SA_ZPL_SEQ; on Linux it
is seeded with (ctime.tv_sec + 1) << 32 so the counter starts above
any pre-change cookie and stays monotonic across the upgrade. A
missing attribute is simply treated as not-yet-migrated, not an
error. FreeBSD never folded ctime into va_filerev, so it needs no
seed.

No feature flag or on-disk format change is needed: the new SA
attribute is keyed by name, so an implementation that does not know
it preserves it opaquely, and the first modify lazily migrates each
file. Covers both the Linux and FreeBSD ZPL.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
lundman pushed a commit to openzfsonosx/openzfs-fork that referenced this pull request Jul 30, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
lundman pushed a commit to openzfsonosx/openzfs-fork that referenced this pull request Jul 30, 2026
The SA bulk arrays are filled by series of conditional SA_ADD_BULK_ATTR
calls whose worst-case count is easy to miscount, so a future attribute
add could silently overrun the fixed-size array.
Add ASSERT3S(count, <=, ARRAY_SIZE(bulk)) before each sa_bulk_update so
an overrun trips in debug builds, and use ARRAY_SIZE so the bound stays
tied to the declaration. Linux zfs_setattr allocates its arrays on the
heap sized by 'bulks', so it asserts against that instead.
Also tighten FreeBSD zfs_setattr's xattr_bulk from [7] to [6], its
actual worst case.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
tonyhutter pushed a commit to tonyhutter/zfs that referenced this pull request Aug 12, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
tonyhutter pushed a commit to tonyhutter/zfs that referenced this pull request Aug 20, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
robn pushed a commit to robn/zfs that referenced this pull request Aug 20, 2026
Commit 312bdab advertises STATX_ATTR_CHANGE_MONOTONIC and builds
the NFSv4 change_cookie from (ctime.tv_sec << 32) | zp->z_seq.
zp->z_seq is reset to a magic constant in zfs_znode_alloc(), so any
event that drops the znode from cache (memory pressure, remount,
reboot) regresses the lower bits of the cookie, a backward step
within the same second.

NFSv4 clients that trust this contract treat a regressed cookie as
evidence that the file's metadata cannot be relied on. VMware ESXi
over NFSv4.1 surfaces this as "The file specified is not a virtual
disk", and a VM stored on the affected NFS-exported ZFS dataset
fails to power on.

Widen z_seq to 64 bit and present it directly as the change_cookie,
dropping the ctime packing, so the cookie is a single monotonic
counter that no longer depends on the clock. FreeBSD's va_filerev
consumer also takes the wider value.

Persist z_seq via a new SA attribute SA_ZPL_SEQ. An in-core marker
zp->z_has_seq records whether the file already carries SA_ZPL_SEQ in
its layout; it is derived at load time and never stored on disk, so
no global pflag bit is consumed. ZFS_SEQ_MAY_GROW() keys off the
marker to grow the SA layout only on the first add per file;
ZFS_PERSIST_SEQ() then sets the marker and adds SEQ to the caller's
bulk alongside the file's other SA attributes. zfs_znode_alloc()
restores z_seq from SA_ZPL_SEQ when present and sets the marker;
zfs_rezget() recomputes the marker in place on rollback/recv without
disturbing the in-core z_seq, keeping the cookie monotonic.

A file written before this change carries no SA_ZPL_SEQ; on Linux it
is seeded with (ctime.tv_sec + 1) << 32 so the counter starts above
any pre-change cookie and stays monotonic across the upgrade. A
missing attribute is simply treated as not-yet-migrated, not an
error. FreeBSD never folded ctime into va_filerev, so it needs no
seed.

No feature flag or on-disk format change is needed: the new SA
attribute is keyed by name, so an implementation that does not know
it preserves it opaquely, and the first modify lazily migrates each
file. Covers both the Linux and FreeBSD ZPL.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
tonyhutter pushed a commit to tonyhutter/zfs that referenced this pull request Aug 20, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
tonyhutter pushed a commit to tonyhutter/zfs that referenced this pull request Aug 20, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
bugclerk pushed a commit to truenas/zfs that referenced this pull request Aug 26, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
(cherry picked from commit 83eabaf)
bugclerk pushed a commit to truenas/zfs that referenced this pull request Aug 26, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
(cherry picked from commit 0da2e6b)
creatorcary pushed a commit to truenas/zfs that referenced this pull request Aug 26, 2026
…431)

* CI: FreeBSD 15.1 STABLE

Update the freebsd15-1s builder to the released STABLE image.

Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18524

* CI: Fix 99.99 META version

We have an option in zfs-qemu-packages to test against a specific kernel
version.  However, qemu-3-deps.sh was incorrectly hard coded to look
at $2 for a kernel version argument (which could come in $2 or $3
depending on if --poweroff was also passed).  This caused the CI
to incorrectly edit META with a max supported kernel version of 99.99
when we didn't want that.

Fix this by looking at all the arguments for something that looks
like a kernel version and set that as the kernel max in META.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18526
Closes #18531

* CI: Remove deprecated Fedora 42

Fedora 42 was deprecated on May 13 2026.  Remove it from CI tests.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18545

* CI: Allow testing with a newer GCC on ARM builder

Add a text box to specify a custom GCC version (like '16') when
running the zfs-arm builder.  This allows you to test with a newer
GCC than the Ubuntu default.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18540

* CI: remove FreeBSD 13.5 (EOL April 30, 2026)

FreeBSD 13.5 and stable/13 reached End-of-Life on April 30, 2026 and no
longer receive security support, so they fall outside README.md's stated
support policy.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Christos Longros 
Closes #18553

* CI: Add Ubuntu 26.04 builder

The Ubuntu 26.04 LTS, named "Resolute Raccoon, was released on
April 23, 2026.  Add to the supported releases in README.md and
add a CI builder for it.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* CI: Fix qemu-guest-agent systemd enable

The qemu-guest-agent.service for Debian and Ubuntu does
not contain an install section which prevents it from
being enabled.  Add a drop-in override file so it can
be enabled and the service started on boot.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* ZTS: zfs_unshare_006_pos.ksh enable usershares

Ensure samba usershares are enabled in the CI test environment for
the zfs_unshare_006_pos test case.  By default they are disabled
in the Ubuntu 26.04 LTS and must be enabled.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* CI: Build custom branch from zfs-qemu-packages

The zfs-qemu-packages workflow allows us to easily build RPMs for the
current branch.  However, there can be cases where we want to use the
current CI environment to build older releases.  This can happen when
the VM or runner environment changes, and the older CI doesn't have
the updates needed to run with it anymore.

This commit adds in a text box to specify a specific branch/tag to build
using the current CI environment.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18569

* CI: enable FreeBSD 15.0-RELEASE in matrix

Add freebsd15-0r to the FreeBSD presets

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Christos Longros 
Closes #18561

* CI: skip full CI runs on push events

Full CI runs for proposed changes always occur in the PR where the
review is done and patch approved.  Once merged the full CI is run
again using the merged commit.  This is somewhat overkill.  In the
interest of reducing the CI load only run the zloop and checkstyle
workflows which are enough to verify the build on the master branch.
Push events to forks will continue to trigger a full CI run.

Reviewed-by: George Melikov 
Signed-off-by: Brian Behlendorf 
Closes #18571

* CI: run full CI when a workflow YAML changes

FULL_RUN_REGEX in generate-ci-type.py covered .github/workflows/scripts/
but not the workflow YAML files, so a PR that only edited zfs-qemu.yml
got "quick" CI and never tested its own matrix change. Add the YAML
files to the list.

Reviewed-by: George Melikov 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18577

* .github: update workflows README

Describe the current zfs-qemu pipeline, ci_type selection, supported
guests, and the code-checking and other auxiliary workflows.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18590

* CI: Update checkstyle checkout action to v6

The checkstyle workflow was the only one still pinned to
actions/checkout@v4; the other workflows already use v6.
Bump it to match.

Reviewed-by: Tony Hutter 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18600

* CI: Lustre 6.16 kernel compatibility fix (#18602)

Almalinux 9,10 kernels now include a backport of Linux commit
v6.15-13744-g41cb08555c41 which renames the from_timer() function
to timer_container_of().  Apply the upstream Lustre compatibility
patch to our builds.  This patch should be included in the next
Lustre release and can be dropped then.

ZFS-CI-Type: quick

Signed-off-by: Brian Behlendorf 
Reviewed-by: Tony Hutter 

* CI: skip smatch, zloop, and zfs-arm for documentation-only changes

Follow-up to #18518, which skipped the qemu matrix on doc-only PRs.
zloop, zfs-arm, and smatch are irrelevant to doc-only changes.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Tony Hutter 
Signed-off-by: Christos Longros 
Closes #18601

* build: add ZFS_DEBUG Kconfig for copy-builtin

... so we can toggle ZFS debug assertions from the
Linux kernel build without having to regenerate the
ZFS patch.

Update the qemu test script to also set this kernel
config.

Reviewed-by: Tony Hutter 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Timothy Day 
Co-authored-by: Timothy Day 
Closes #18595

* CI: apt-get update before purging host packages

The package removal ran against a stale package index and failed to
fetch a package that had been removed from the repository. Refresh
the index first.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18607
Closes #18609

* CI: add concurrency support to zfs-arm

The zfs-arm workflow was the only build/test workflow without a
concurrency block, so superseded runs were not cancelled.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18608

* nvpair: Check for un-terminated strings in packed nvlist

Add additional checks to verify a packed string or string array nvpair
is terminated.  Or more specifically, verify doing a strlen() on the
prospective string does not overrun the packed nvlist buffer.

Also add additional checks in the libzfs_input_checks test case to
verify un-terminated strings, and add in a nvlist ioctl payload
fuzz test for good measure.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18604

* Fix the integer type in zfs_ioc_userspace_many()

Fix the mismatched type in zfs_ioc_userspace_many() and limit the
number of entries returned to 1000.  When a size larger than this
is requested the response is truncated, zfs_userspace() already
correctly handles short responses.  Historically, zfs_userspace()
has requested 100 entries at a time, this cap allows for 10x larger
batch sizes if needed in the future.

Reported-by: Yuxiang Yang, Yizhou Zhao, Ao Wang, Xuewei Feng, Qi Li,
Reported-by: and Ke Xu from Tsinghua University using GLM-5.1 from Z.ai
Reviewed-by: Alexander Motin 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18615

* sharenfs: Check for invalid characters

Check for invalid characters in sharenfs/sharesmb dataset props.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18613

* Fix uninitialized variable warning in vdev_prop_get()

Update vdev_prop_get_objid() to set objid on error as the comment
in vdev_prop_get() describes.

    "objid is set to 0 when absent and the few cases that call
    zap_lookup directly guard against this below."

This resolves the following possible uninitialized variable warning.

    module/zfs/vdev.c: In function ‘vdev_prop_get’:
    module/zfs/vdev.c:6913:12: error: ‘objid’ may be used uninitialized
    in this function [-Werror=maybe-uninitialized]

Reviewed-by: Alexander Motin 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18616

* Extend dataset zfs_ioc_set_prop() secpolicy

When zc->zc_cookie is set this indicates to zfs_ioc_set_prop() that
these are received properties and ZPROP_HAS_RECVD will be set on the
dataset.  This is only done as part of a `zfs receive` so additionally
apply the zfs_secpolicy_recv() policy.  Individual property checks
continue to be handled by zfs_check_settable().

Reviewed-by: Tony Hutter 
Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18617

* pam: use open fd instead of path

Instead of performing multiple operations on the path name in
zfs_key_config_modify_session_counter() open the file once and
perform the fchown, fchmod, and openat on the open file handle.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18618

* Remove /etc/sudoers.d/zfs

The smartctl exception in /etc/sudoers.d/zfs doesn't cover devices
like NVMe or symlinked devices.  Just get rid of it rather than
keep maintaining it.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Tony Hutter 
Closes #18626

* CI: Re-enable CodeQL workflows on push

This workflow was disabled 'on push' recently in commit 1916c2c5
to reduce redundant CI runs.  However, this check is fairly quick
and we want it run regularly against the branches.  Enable it.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18627

* CI: Update CodeQL actions to v4

CodeQL Action v3 has been deprecated and will be retired
December 2026.  Update codeql.yml to use CodeQL Action v4
and update the runner to ubuntu-24.04.

Signed-off-by: Brian Behlendorf 
Closes #18629

* CI: Increase default RCU stall timeout on Linux

When CONFIG_RCU_CPU_STALL_TIMEOUT is configured an RCU stall which
exceeds the default timeout will trigger an NMI and panic the VM.
Given the heavily virtualized nature of the CI environment we want
to make sure to only trigger this due to a real deadlock and not
due to over-subscription of the systems resources.  This timeout
normally defaults to 20-30 seconds and this change increases it
to 120 seconds.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18624

* CI: Add alternative URLs for CentOS stream

Fallback to trying the "CentOS Strean Composes" repo for the qcow2
images if the regular URLs fail.  The Composes repo contains the daily
autobuilt Stream images.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18628

* Add additional verification of size fields and strings (#18623)

- Check for size fields that convert to smaller integers.
- Explicitly terminate bootenv string.
- Initialize variables that could be returned in an error case.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Reviewed-by: Alexander Motin 
Signed-off-by: Tony Hutter 
Closes #18623

* Fix uninitialized variable warning in zil_parse()

This resolves the following possible uninitialized variable warning
when building with --enable-code-coverage and gcc 8.5.0.

    module/zfs/zil.c: In function ‘zil_parse’:
    module/zfs/zil.c:549:47: warning: ‘end’ may be used uninitialized
    in this function [-Wmaybe-uninitialized]

Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18633

* ZTS: relax zpool_import_parallel_pos.ksh timing

Occasionally in the CI this test will fail because the parallel import
took longer than half of the serial time (but still less than the full
serial time).  Increase the cutoff to 3/4 of the serial time to preserve
the intent yet try and avoid these false positive failures.

Reviewed-by: Chris Longros 
Signed-off-by: Brian Behlendorf 
Closes #18634

* linux: verify stale znodes in legacy fallocate

The mode=0 and FALLOC_FL_KEEP_SIZE preallocation path can reach
zfs_freesp() directly and call zfs_statvfs() before going through the
normal zpl_enter_verify_zp() boundary.

When zfs_rezget() tears down a failed SA reload, a stale inode may
remain alive in the VFS with z_sa_hdl cleared. The unchecked
fallocate path can then reach sa_lookup(zp->z_sa_hdl, ...) through
zfs_statvfs() or zfs_freesp() and crash on a NULL SA handle.

Use zfs_enter_verify_zp() in zfs_statvfs() so stale znodes are
rejected under the teardown lock for both fallocate and statfs.
Also wrap the direct zfs_freesp() call in
zpl_enter_verify_zp()/zfs_exit() so this path follows the same
validation rules as the other Linux ZPL file operations.

Fixes: f734301d2267
("linux: add basic fallocate(mode=0/2) compatibility")

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Co-authored-by: gality369 
Closes #18458

* Linux: annotate nested xattr setattr znode locks

zfs_setattr() updates both the target znode and its hidden xattr
directory when ownership, mode, or project ID changes. The xattr
directory uses the same z_acl_lock and z_lock classes as the
parent znode, so lockdep reports recursive locking when the
second znode's mutexes are acquired.

This is a lockdep false positive rather than a real deadlock.
attrzp is the target file's hidden xattr directory, and the code
does not acquire these znode mutexes in the reverse order.
Acquire the attrzp mutexes with mutex_enter_nested() so lockdep
treats them as nested.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Co-authored-by: gality369 
Closes #18506

* Linux: avoid znode list lock inversion during resume

Lockdep reports a circular locking dependency during mounted filesystem
rollback.  zfs_resume_fs() walks z_all_znodes under z_znodes_lock and
calls zfs_rezget(), which takes the per-object znode hold lock via
zfs_znode_hold_enter().

The normal zget path takes these locks in the opposite order.
zfs_zget() takes the per-object hold lock before zfs_znode_alloc()
inserts the znode on z_all_znodes under z_znodes_lock.  Resume can
therefore establish z_znodes_lock -> zh_lock while normal lookup
creates zh_lock -> z_znodes_lock.

Pin the current and next znodes with igrab() while holding the list
lock, then drop the list lock before reloading the znode.  Existing
stale inode handling is preserved, and both the suspended reference
and temporary walk reference are released asynchronously.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes #18517

* linux/zpl_super: handle 'source' option directly

vfs_parse_fs_param_source() didn't appear until 5.14, and was not
backported to kernel.org LTS kernels. It's simple enough that it's
easier to just handle it ourselves rather than use a configure check.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18529

* linux: suppress reclaim lockdep in zfs_inactive via rwlock wrappers

kswapd can enter zfs_inactive() from inode reclaim while holding
fs_reclaim. The z_teardown_inactive_lock still serializes teardown,
but the reclaim-thread acquire/release pair can produce a lockdep
cycle through zfs_zinactive() and zfs_rmnode().

Add Linux rwlock nolockdep wrappers alongside the existing rwlock
macros and use them only for the reclaim-thread
z_teardown_inactive_lock acquire/release in zfs_inactive(). Keep
the real rwsem semantics unchanged and leave CONFIG_LOCKDEP
handling in the platform rwlock layer.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes #18505

* config: show progress output for kernel API checks

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18554

* linux/super: properly apply ro/rw mount option to superblock

f5a9e3a622 changed how SB_RDONLY was applied to the new mount in a way
that was too simplistic - it only sets readonly on the filesystem if the
mount was 'ro', but it never clears it if the mount was 'rw'. This
causes the 'rw' option to effectively be ignored, and so the readonly=
property wins out.

This fixes it by doing it the right way: checking the flags mask to see
if it was actually provided as an option at all, and then setting or
clearing it as appropriate.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18557
Closes #18563

* Linux 5.6 compat: fix fs_parse API mismatch

Added m4 macro to check fs_parse API signature and wrappers.  Before 
5.6, fs_parse() took a struct fs_parameter_description which wraps
the parameter specs with name and enum pointers. From 5.6, the 
description struct was removed and fs_parse() accepts the 
fs_parameter_spec directly.

Reviewed-by: Rob Norris 
Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18585

* Fix aarch64 build failure by removing earlyclobber (#18532)

The UVR macros used "+&w" (read-write + earlyclobber) as the
constraint for NEON register operands that are declared as explicit
hard-register variables via:

register unsigned char wN asm("vN") __attribute__((vector_size(16)));

The + modifier implicitly makes the operand also an input (reading the
register before the asm runs). The & (earlyclobber) modifier says "this
output may be written before all inputs are consumed." Having an
earlyclobber output on the same hard-register that is simultaneously
an input is a contradiction — GCC 16 now strictly diagnoses this.

The fix removes the & from "+&w", yielding "+w". The earlyclobber
was both incorrect (contradicts the implicit input) and unnecessary
(the physical registers are already hard-bound, so the compiler has no
freedom to assign conflicting registers anyway).



Issue #18525

Signed-off-by: Brian Behlendorf 
Co-authored-by: Claude Sonnet 4.6 
Reviewed-by: Tony Hutter 

* Fix "panic: cache_vop_rename: lingering negative entry"

A FreeBSD ZFS filesystem with properties "utf8only=on" and
"normalization=formD" consistently produces this panic when
building the lang/perl-5.42.0 port.

A ZFS file system with "utf8only=off" and "normalization=none"
works fine.

The cause of the panic seems to be incorrectly using the FreeBSD
namecache when normalisation is present. This commit adds a
predicate to prevent that.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Jan Martin Mikkelsen 
Closes #18430

* key lookup failure should always return EACCES

spa_do_crypt_abd() already maps a missing key to EACCES. However
spa_do_crypt_mac_abd(), spa_do_crypt_objset_mac_abd(), and
spa_crypt_get_salt() still return the raw
spa_keystore_lookup_key() error (ENOENT). This is inconsistent
As we want to treat all “no key” failures as a permission
failure. Standardize on EACCES for the unloaded-key case.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alek Pinchuk 
Closes #18448

* Fix off-by-one in PREVIOUSLY_REDACTED handler that drops last block

In send_reader_thread(), the PREVIOUSLY_REDACTED handler computed
file_max as MIN(dn->dn_maxblkid, range->end_blkid).  dn_maxblkid is
an inclusive maximum block ID while range->end_blkid is exclusive (one
past the last block).  The resulting file_max was then used as an
exclusive loop bound, causing the last block of any file (at index
dn_maxblkid) to be silently skipped when a PREVIOUSLY_REDACTED range
covered the end of the file.

The block was never written to the send stream so the receiver kept
zeros there.  ZFS reported no error because the stream itself was
valid; the data was simply absent.

Fix: use dn_maxblkid + 1 so file_max is consistently exclusive.

Add a regression test (redacted_max_blkid.ksh) that modifies only the
last block of a file in one clone, creates a redaction bookmark from
it, then sends an unmodified clone incrementally from that bookmark.
The PREVIOUSLY_REDACTED path must fill in the last block; the test
verifies it is not zeros and matches the original.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Paul Dagnelie 
Reviewed-by: Reviewed-by: Tony Hutter 
Signed-off-by: Manoj Joseph 
Closes #18477

* Avoid flushing unrelated NFS exports on snapshot unmount

zfsctl_snapshot_unmount() called exportfs_flush() before every umount
attempt to drop NFS export cache references that pin the snapshot
mountpoint.  The flush has global effect on the host's NFS exports and
clients, so paying it on every snapshot unmount (including auto-expire
rounds for snapshots that were never NFS-accessed) impacts unrelated
snapshots and clients.

ZFS cannot invalidate individual export cache entries because the
relevant sunrpc cache APIs are exported GPL-only.  Defer the global
flush so it runs only when the umount has actually failed, then retry
once.  Snapshots that are not NFS-pinned succeed on the first attempt
and never trigger the flush.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Youzhong Yang 
Signed-off-by: Ameer Hamza 
Closes #18476

* zfs: annotate nested dd_lock in reservation sync accounting

When reservation sync updates a child's reserved space, it rolls the
delta into ancestor space accounting while still holding the child's
dd_lock.  That locking order is intentional, but Linux lockdep sees
the ancestor acquisition as recursive because it lacks a nested lock
subclass annotation.

Teach the reservation-sync space-accounting path to acquire ancestor
dd_lock instances with a nested subclass.  Keep the existing public
interfaces and accounting behavior unchanged by routing only the
ancestor rollup through local helpers.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Signed-off-by: gality369 
Closes #18497

* sa: fix sa_add_projid lock ordering

sa_add_projid() currently acquires hdl->sa_lock before zp->z_lock.
Several same-znode update paths take zp->z_lock and then call
sa_update() or sa_bulk_update() on the same SA handle.

On Linux, FS_IOC_FSSETXATTR reaches zfs_setattr() through
zpl_ioctl_setxattr() without outer inode serialization. This makes
the reversed lock order a real ABBA deadlock rather than a lockdep
false positive when projid is added to an old-format inode while
another thread updates the same znode.

Acquire zp->z_lock before hdl->sa_lock in sa_add_projid() to match
the existing znode update ordering.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Co-authored-by: gality369 
Closes #18503

* zdb: detect BRT and DDT leaks during block traversal

During -b traversal, track BRT and DDT reference counts and report
blocks claimed more times than their reference tables account for
if it causes claim errors, instead of just asserting it.  Also
report entries with references not fully consumed by the traversal.

Add zdb leaks checks to cloning and dedup tests. This should make
sure the pools are in a sane state after completing the functional
tests.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #18494

* zarcstat: detect attached L2ARC device with no data

zarcstat and zarcsummary detected L2ARC presence using the l2_size
kstat, which is data held in L2ARC, not whether a cache device is
attached. When a cache device was attached but empty (freshly added,
or fully evicted):

  - zarcstat rejected "-f l2*" with "Incompatible field specified!"
  - zarcsummary printed "L2ARC not detected, skipping section",
    hiding cumulative I/O history and health counters

Expose the existing l2arc_ndev counter as a new kstat l2_dev_count.
It is maintained by l2arc_add_vdev() and l2arc_remove_vdev(), so it
tracks attachment in real time. Use it in both tools, falling back to
l2_size for compatibility with older kernel modules.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Ameer Hamza 
Closes #18499

* Fix double free for blocks cloned after DDT prune

Before this change, for blocks marked with D flag but absent in DDT
(pruned from it), zio_ddt_free() fell back to ZIO_STAGE_DVA_FREE
without trying ZIO_STAGE_BRT_FREE first.  Same time such blocks
might be present in BRT, and not handling that would result in
double/multiple free.

This change makes ZIO_DDT_FREE_PIPELINE include ZIO_FREE_PIPELINE,
just adding required ZIO_STAGE_ISSUE_ASYNC and ZIO_STAGE_DDT_FREE,
and moves DDT stages before BRT.  This way, if the block is found
in DDT by zio_ddt_free(), the pipeline is short-circuited to
ZIO_INTERLOCK_PIPELINE, similar to what zio_brt_free() does.  If
not, then BRT is checked, and if also no match, the block is freed.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Rob Norris 
Signed-off-by: Alexander Motin 
Closes #18520

* arc: export additional required symbols

External consumers of arc_read() need to be able to destroy the
returned arc_buf_t.  Add the arc_buf_destroy() interface as an
exported symbol.

Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18533

* zap_impl: use flex array field for mzap_phys_t.mz_chunks

mz_phys_t is always a full-block allocation, with mz_chunks[] as an
array over the rest of the block past the header.

Recent Linux compiled with CONFIG_UBSAN will complain about this:

    UBSAN: array-index-out-of-bounds in module/zfs/zap.c:1236:28
    index 2 is out of range for type 'mzap_ent_phys_t [1]'

The fix is straightforward; simply convert this field to a flex member.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18550

* spl_kvmalloc: remove __GFP_COMP before calling vmalloc()

In cb1833023 we stopped using it for KM_VMEM allocations, since its not
a valid flag for vmalloc(). However, there's a fallback path for
non-KM_VMEM allocations to use vmalloc(), and we need to remove
__GFP_COMP there too to avoid a warning.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18558

* FreeBSD: Make it possible to build openzfs.ko with sanitizers

Add make options which let one respectively compile the kernel modules
with the address sanitizer, memory sanitizer, and undefined behaviour
sanitizer enabled.  This makes it much easier to run the ZTS with those
sanitizers enabled.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Signed-off-by: Mark Johnston 
Closes #18596

* enforce exact decompressed length for lz4, gzip, and zstd

Decompressors must expand a ZFS block to exactly the expected number
of bytes. Treat decompression to an unexpected length as failure, so
truncated or short output is not accepted as valid decompression. This
makes our handling of decompress return values consistent with the
decompression functions' APIs.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Alek Pinchuk 
Closes #18599

* dsl_scan: close errorscrub cursor on pause

If the cursor were ever to actively hold resources, not finalising it
would mean leaking those resources whenever the scrub is paused.

The cursor is already reinitialized from the stored serialized form
if/when it is resumed, so there's nothing we need from the old one, just
to release it.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18603

* When reading a vdev label skip libzfs_core_init()

There's no need to call libzfs_core_init() when `zdb -l` is used to
read a vdev label.

Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18606

* Simplify dnode_level_is_l2cacheable()

We should not dereference through dn_handle->dnh_dnode once we
already have a dnode pointer.  The result will be the same.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #18212

* Remove parent ZIO from dbuf_prefetch()

I am not sure why it was added there 10 years ago, but it seems not
needed now.  According to my tests removing it improves sequential
read performance with recordsize=4K by 5-10% by reducing the CPU
overhead in prefetcher.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Rob Norris 
Reviewed-by: Ameer Hamza 
Reviewed-by: Akash B 
Signed-off-by: Alexander Motin 
Closes #18214

* Fix log vdev removal issues

When we clear the log, we should clear all the fields, not only
zh_log.  Otherwise remaining ZIL_REPLAY_NEEDED will prevent the
vdev removal.  Handle it also from the other side, when zh_log
is already cleared, while zh_flags is not.

spa_vdev_remove_log() asserts that allocated space on removed log
device is zero.  While it should be so in perfect world, it might
be not if space leaked at any point.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #18277

* ZVOL: Add encryption key check for block cloning

Somehow during block cloning porting from file systems was missed
the check for identical encryption keys.  As result, blocks cloned
between unrelated ZVOLs produced authentication errors on later
reads.  Having same or different encryption root does not matter.

This patch copies dmu_objset_crypto_key_equal() call from FS side.

Reviewed-by: Ameer Hamza 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #18315

* abd: Fix stats asymmetry in case of Direct I/O

abd_alloc_from_pages() does not call abd_update_scatter_stats(),
since memory is not really allocated there.  But abd_free_scatter()
called by abd_free() does.  It causes negative overflow of some
ABD and possibly ARC counters.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Rob Norris 
Signed-off-by: Alexander Motin 
Closes #18390

* Tag zfs-2.4.3

META file and changelog updated.

Signed-off-by: Tony Hutter 

* Constify some rrd_*() functions

These don't modify the db, so just constify them while we're in the
area.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Signed-off-by: Kyle Evans 

* Add dbrrd_latest_time() to grab the latest timestamp in the db

Returns 0 if the database is empty, otherwise it returns the highest
value of the minutely db.  dbrrd_add() will already enforce the property
that these are monotonically increasing, so we won't try to second-guess
it.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Signed-off-by: Kyle Evans 

* freebsd: set mnt_time on the rootfs at mountroot time

FreeBSD's vfs_mountroot() will collect `mnt_time` from every filesystem
that we mounted and use the highest timestamp as a source for the system
time if we didn't get anything from an attached RTC.

Use the rrd mechanism added to gather up a notion of the latest time
and set it on mnt_time.  If the timestamp db is empty, we just fallback
to the uberblock timestamp and hope that that is in the right ballpark.

Relevant: FreeBSD PR254058[0] reporting the problem downstream

[0] https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=254058

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Signed-off-by: Kyle Evans 

* zio_ddt_write: compute have_dvas after taking dde_io_lock

In zio_ddt_write(), have_dvas and is_ganged were computed before
dde_io_lock was taken. A concurrent zio_ddt_child_write_done() error
path calls ddt_phys_unextend() under dde_io_lock, which can zero
DVA[0] while another thread is between computing have_dvas and taking
dde_io_lock. That thread then uses the stale have_dvas=1 to call
ddt_bp_fill(), copying the zeroed DVA into the BP. A zero DVA resolves
as a hole, producing blocks that read back as zeros with no checksum
error (silent data corruption).

Fix by moving have_dvas and is_ganged computation to after dde_io_lock
is taken, so they always reflect the current state of dde->dde_phys.

Regression introduced by a41ef36858 ("DDT: Reduce global DDT lock
scope during writes").

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Co-authored-by: Saju Palayur 
Signed-off-by: Saju Palayur 
Closes #18366
Closes #18544
(cherry picked from commit 6fb72fda0f60d9efb591e320f83f78b19ec451cc)
(cherry picked from commit 0a07efe3283df42db41a2060dc3ae0c455a7c056)

* ZTS: Pass dec instead of hex to mknod

On Ubuntu 26.04 the default mknod command returns an error when
provided the major and minor numbers in hex.  Switch to passing
decimal values.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* ZTS: statx_dioalign.ksh update to stride_dd

The uutils 0.8.0 version of dd appears to diverge from GNU behavior
and does not fail when an unaligned write O_DIRECT write is issued.
Update the test case to use stride_dd which is provided by the ZTS
so the expected syscall behavior can be verified.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* ZTS: delegate: add test for send sub-permissions

Regular send and raw send are actually separate operations with separate
permissions. This adds a test to test the combinations properly using
the existing permission test infrastructure.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18672
(cherry picked from commit 4d1d00f9fe0c87b2334613a0efaa758275f938b8)

* ZTS: remove send_delegation tests

These tests are doing the same tests as delegate/zfs_allow_send, and are
hard to follow and maintain. There's no need for them now, so drop them.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18672
(cherry picked from commit 562b96ced9f28b8517913e5e9f7eaf1686db7cd0)

* ZTS: delegate: add encryption option for test fixture datasets

The delegate test framework doesn't care about the encryption status of
the dataset under test, so by adding an option to create with encryption
the framework can be used to check encryption-related permissions
without any further fanfare.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit 8303a36488da79c13d0fcca4365d71d5180407c3)

* ZTS: delegate: check send permissions on encrypted datasets

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit bce9a8ef7d0b4c1b8e323ef2b045379177c20410)

* ZTS: delegate: test send:encrypted

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit 166a6672502c6398b3ef549d4a17f113f5cb2e8d)

* zfs_secpolicy_send: lift checks to common function for both

The permissions checks for send are a little involved because different
permissions grant different abilities, and there's two ways to initiate
a send.

This lifts the common permissions checks into a single function, and
ensures that we maintain a single dataset hold across all checks. This
will become important in the next commit when we need to check a
specific dataset property as part of the permission check.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit f0d69e6b16d93ab49f1914a29e64d5df6e4f7bc2)

* delegate: add 'send:encrypted' permission

send:encrypted is like send:raw, but only permits encrypted datasets to
be sent - raw send is not permitted for unencrypted datasets.

This commit creates the permission, wires it up, and adds the check for
it in zfs_secpolicy_send_impl(), if it is the last send permission
standing, the dataset is checked for its encryption state.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit 97b9ba7a982e36d39a3cf7db271da2fab110769d)

* zbookmark_compare: handle "marker" bookmarks with negative levels

"Marker" bookmarks (those with zb_level == ZB_ROOT_LEVEL, ZB_ZIL_LEVEL
or ZB_DNODE_LEVEL) represent valid blocks, but are associated with a
dataset directly rather than with a specific object within it. They end
up on bookmark lists during scan prefetch, and so need to be sorted
ahead of any "true" object blocks.

The problem is that for negative levels, BP_SPANB produces a negative
shift, which is not legal C. Fortunately the results are used only for
comparison, so the worst possible behaviour in a forgiving compilation
environment is a mis-sort, which for the scan/traverse cases, means that
we haven't prefetched certain metadata before we actually need it. But
there _is_ UB in there, and UBSAN does rightly complain.

Here we fix all this by handling these bookmarks directly - sorting them
ahead of "true" object blocks, which is usually what scan/traverse will
prefer. And we don't do any interesting math on these bookmarks, so we
sidestep the whole UB thing.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #14777
Closes #18652

* arc: add a few invariant checks in release builds

Convert a couple ASSERTs invariants to VERIFYs to enforce them
in release builds to be able to root-case #18782 kernel panic,
whenever it happens again.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Andriy Tkachuk 
Closes #18840
(cherry picked from commit 023d44b9ef68f1eff6b97d98c6a16cdb4b93cf76)

* CI: Have zfs-build-packages workflow build tarballs on Alma (#18662)

Previously, zfs-build-packages would only build source tarballs
on Fedora due to problems with building them on RHEL 7.  That's
a relic of the past now, as we haven't supported RHEL 7 since
it went EOL in 2024.  With this change, we now build the tarballs
on both Alma and Fedora.

Signed-off-by: Tony Hutter 
Reviewed-by: Olaf Faaland 
Reviewed-by: Chris Longros 

* Update our CI runners to the newest FreeBSD 15.1 RELEASE (#18667)

Signed-off-by: Christos Longros 
Reviewed-by: Alexander Motin 
Reviewed-by: Tony Hutter 

* Linux 7.1 compat: META (#18682)

Update the META file to reflect compatibility with the 7.1
kernel.

Signed-off-by: Tony Hutter 
Signed-off-by: Rob Norris 
Reviewed-by: Chris Longros 

* Fix handling of _PC_HAS_HIDDENSYSTEM for FreeBSD

The hidden and system flags are only supported for
ZFS pools if the z_use_fuids is true.  Fix
zfs_freebsd_pathconf() to check this.

Reviewed-by: Alexander Motin 
Signed-off-by: Rick Macklem 
Closes #18688

* zfs_ioctl: fix EBUSY race between quota queries and mount

zfsvfs_hold() fell back to zfsvfs_create() -> dmu_objset_own()
(exclusive) for unmounted datasets.  A concurrent zfs_domount()
also calls dmu_objset_own(), causing EBUSY on the same dataset.

Introduce zfsvfs_create_hold() using dmu_objset_hold() (shared
hold) instead.  Shared holds do not conflict with exclusive owns,
eliminating the race.  The release path (zfsvfs_rele,
zfsvfs_create_impl error) uses dmu_objset_ds()->ds_owner to
determine whether to disown or rele, avoiding the need for an
extra flag in zfsvfs_t.

Added tests userspace_005, groupspace_005, projectspace_006
(50 iter race test).

Reviewed-by: Brian Behlendorf 
Reviewed-by: Tony Hutter 
Signed-off-by: HeonJe Lee 
Closes #18611

* CI: Re-allow workflow_dispatch on zfs-qemu

Allow zfs-qemu to be invoked from a workflow_dispatch event (a.k.a,
manually running a workflow).  This may have been accidentally disabled
in 1916c2c55.

Reviewed-by: Chris Longros 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18680

* initramfs-zfs should not try to copy directories

We had find only return files from the beginning for libgcc.so, but not
libfetch/libcurl. This oversight affected a user when vmware installed
its own libcurl.so.4 in a directory called libcurl.so.4, since our code
then tried to copy a directory, which fails.

Reviewed-by: Chris Longros 
Reviewed-by: Brian Behlendorf 
Suggested-by: Carsten Härle 
Signed-off-by: Richard Yao 
Closes #18582
Closes #18686

* Clean up embedded slog metaslab across txgs

On a read-write import, metaslab_set_fragmentation() can dirty a
metaslab via vdev_dirty() while still in the txg==0 load path when its
space map has an unexpected bonus size (e.g. a makefs-created pool
whose space-map dnodes use the boot loader's 24-byte space_map_phys_t
with nblkptr=3, giving db_size=64). If that metaslab is then selected
as the embedded slog, vdev_metaslab_init() only removed it from
vdev_ms_list when txg != 0, so the txg==0 case left it queued and
metaslab_fini() tripped VERIFY(!txg_list_member(&vd->vdev_ms_list,
msp, t)).

Remove slog_ms from the dirty list for every TXG_SIZE slot before
metaslab_fini() so the cleanup is correct regardless of txg.

Reported on FreeBSD as PR 281520:
External-issue: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=281520

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Nick Price 
Closes #18693

* README: update supported FreeBSD release to 15.1

Our CI runners moved to FreeBSD 15.1 in 0a4b59765 (#18667), but the
README still lists 15.0. Update it to match the CI version.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Christos Longros 
Closes #18696

* honor file argument in file_wait_event

grep the log path passed by the caller instead of always using
ZED_DEBUG_LOG.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Alek Pinchuk 
Closes #18700

* Update mtime/ctime when fallocate grows a file

Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes #18573

* ddt_log: Fix refcount tagging for begin/commit

Sponsored-by: Klara, Inc.
Sponsored-by: Wasabi Technology, Inc.
Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Igor Ostapenko 
Closes #18706

* CI: Increase default watchdog NMI timeout on Linux

When the watchdog driver is configured and enabled an NMI will be
generated when the watchdogd process fails to regularly reset the
watchdog timer.  Given the heavily virtualized and potentially
over-subscribed nature of the CI environment increase the default
timeout to 120 seconds (normally defaults to 30 seconds).

Reviewed-by: Christos Longros 
Signed-off-by: Brian Behlendorf 
Closes #18704

* Fix race between device removal completion and pool export

vdev_remove_complete() finalizes a device removal in two phases under
the spa lock framework. Between the two phases it called
spa_vdev_exit(), which drops both the config locks (SCL_ALL) and
spa_namespace_lock and blocks on a txg sync. By that point
vdev_remove_replace_with_indirect() has already set svr->svr_thread =
NULL, and that is the only thing the export path (spa_export_common()
-> spa_async_suspend() -> spa_vdev_remove_suspend()) waits on. Once the
namespace lock is dropped, a concurrent export or destroy can acquire
it and set spa->spa_export_thread. When the removal thread re-enters
for its second phase via spa_vdev_enter(), it trips the
ASSERT0P(spa->spa_export_thread) assertion.

Hold spa_namespace_lock across both phases instead of dropping and
re-taking it: the intermediate spa_vdev_exit() becomes
spa_vdev_config_exit(), which drops only SCL_ALL and syncs the txg
while keeping the namespace lock held, and the second spa_vdev_enter()
becomes spa_vdev_config_enter(). Because the namespace lock is never
dropped between the phases, a concurrent export blocks in
spa_namespace_enter() and cannot set spa_export_thread until removal
finalization is done. This mirrors the multi-phase locking pattern
already used by the attach/detach and split paths in spa.c.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Prakash Surya 
Closes #18657

* linux: handle mmap read beyond file size

When performing a mmap read past the end of a file there is no data to
read, so simply zero-fill the page and return success.  zfs_getpage()
limits the range lock appropriately to cover the offset being read.

Reported-by: Iliya Polihronov (@vnsavage) (Automattic)
Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18715

* Using net/cloud-init to unpin specific python

These days freebsd 15/16 fail when fetching py311-cloud-init.
Switch to net/cloud-init to avoid python version pinning.

Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18717

* linux/abd: remove BIO support functions

Not used since the "classic" vdev_disk implementation was removed in
5764e218ba.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18719

* Linux 7.2: zpl_super: convert to sget_fc()

The old sget() superblock matcher has been removed in favour of the
fscontext-based sget_fc(). This converts to it. Its largely a signature
change, no functional change.

sget_fc() has existed since fscontext was introduced, so there's no need
for separate feature tests.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18677

* zpl_ctldir: remove comments describing ancient kernel behaviour

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18722

* Fix insufficient locking in dedup verify

Introduction of dde_io_lock removed global DDT lock acquisition
from write completion.  As result, white ZIO ABD could be freed
while zio_ddt_collision() is comparing against it.  Taking there
dde_io_lock should fix the issue.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #17960
Closes #18712
Closes #18720

* libzfs: fix MS_CRYPT/MS_OVERLAY collision with umount2(2) flags

MS_CRYPT and MS_OVERLAY are libzfs-internal mount flags, but their
values (0x8 and 0x4) aliased the umount2(2) flags UMOUNT_NOFOLLOW and
MNT_EXPIRE. A consumer that legitimately set UMOUNT_NOFOLLOW therefore
had that bit read as MS_CRYPT, so libzfs unloaded the dataset's
encryption key as a side effect.

Move both flags to high bits unused by umount2(2) and strip them before
the unmount syscall in do_unmount() (umount2(2) on Linux, unmount(2) on
FreeBSD) and in cmd/zfs. MS_CRYPT is a compile-time macro, so consumers
that set it (for example truenas_pylibzfs) must be rebuilt against the
new header.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes #18713

* ZTS: ctime_001_pos increase tolerance

The ctime_001_pos test checks that timestamp updates occur for a
file after performing certain operations (read, write, chown, etc).
The test case allowed for a +4 second tolerance in the timestamp
value which is generous but up to +7 second discrepencies have been
seen in the CI.  Bump the tolerance to +10 seconds to prevent these
false positives.  As long as the value increases and is reasonably
close to the expected value consider that to be sufficient.

Reviewed-by: Christos Longros 
Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18733

* build: Fix for building dist target outside of project root

Do not change the working directory when doing copying because $distdir
is a path relative to the working directory, and thus not valid after
changing the working directory. This previously worked, when configure
was run from the project root, because @srcdir@ was the same as the
build working directory.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Glenn Washburn 
Closes #18744

* Linux: fix zfs_write() infinite loop on unfaultable buffer

On Linux, zfs_write() copies from the source buffer with page faults
disabled while the transaction is open, relying on
zfs_uio_prefaultpages() to make the pages resident beforehand.  When
dmu_write_uio_dbuf() returns EFAULT, the retry path only subtracts
the bytes consumed so far from the prefault accounting.  On zero
progress pfbytes does not change, the "pfbytes < nbytes" check never
triggers another prefault, and the loop retries the same failing
copy.  If the buffer can never be faulted in again, e.g. the owning
process was torn down while a thread was in pwrite(2), that thread
spins unkillably at 100% CPU holding the file range lock, blocking
every other accessor of the file.

This is a regression from commit b0cbc1aa9a ("Use big transactions
for small recordsize writes."), which dropped the unconditional
re-prefault the EFAULT path had carried since commit 779a6c0bf6
("deadlock between mm_sem and tx assign in zfs_write() and page
fault").  Restore those semantics by resetting pfbytes on EFAULT, as
suggested in the issue analysis: the next iteration then faults the
pages back in before retrying, and for a permanently inaccessible
buffer zfs_uio_prefaultpages() fails, breaking the loop with EFAULT,
which is already propagated to userspace.  Transient faults, such as
mmap'ed source pages evicted under memory pressure, retry as before.

Built and tested on Linux aarch64: the ZTS mmap and write-path groups
pass and a munmap-versus-write stress run leaves no stuck writers.  The
teardown race itself is not reproducible on demand, so the retry paths
were also checked by inspection against the reproducers in the issue.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #17129
Closes #18740

* linux: batch DMU reads of non-resident pages in mappedread()

When a range being read has at least one page in the page cache,
zfs_read() routes the whole chunk through mappedread(), which falls
back to a separate dmu_read_uio_dbuf() call for every non-resident
PAGE_SIZE piece.  Since cached pages outlive munmap(), a file which
was mapped at some point may sit mostly outside the page cache and
still pay this cost: one DMU call per 4K page instead of one per
chunk, measured in #16031 as a 4-10x sequential read slowdown.
Commit 39be46f43 ("Linux 5.18+ compat: Detect filemap_range_has_page")
fixed the detection side so fully uncached chunks bypass mappedread()
again, but a chunk holding even one resident page still degrades to
page-sized DMU reads for everything else.

Instead of issuing one DMU read per absent page, probe the page cache
with find_get_page() and extend the read over the whole run of
non-resident pages which follows, restoring chunk-sized DMU reads for
the uncached parts of a mapped file.

A page can be faulted in after it was observed absent and before the
DMU read covering it completes, but this is safe for the same reason
the existing single-page window is.  zfs_read() holds the znode
rangelock as reader across mappedread(), so the DMU contents of the
range are stable: zfs_write(), zfs_putpage() and truncation all
require the writer lock, and zfs_getpage() fills concurrently faulted
pages from those same contents.  A faulted page can only diverge from
the DMU once dirtied through a writable mapping, making that store
concurrent with this read, for which returning the pre-store data is
a valid outcome.  Stores which completed before the read began cannot
be missed: a dirty page cannot be cleaned and reclaimed while the
reader lock is held (writeback takes the writer lock), so it is still
found resident, or zfs_putpage() already copied its data into the
DMU.

FreeBSD's mappedread() has the same per-page fallback and could be
batched the same way in a follow-up.

Measured in a VM with a 1 GiB file held in the ARC and read
sequentially with dd: one resident page per 1 MiB read request
degrades throughput from ~11.4 GB/s (no resident pages) to ~5.7 GB/s
on the baseline, and this change restores ~11.4 GB/s; one resident
page per 32 MiB chunk, read in 32 MiB requests, improves from
~5.4 GB/s to ~7.3 GB/s.  Reads of a fully resident file are
unaffected (~20 GB/s before and after).  All tests in the ZTS mmap
group pass, including the mmap_read and mmap_seek cases.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #16031 
Closes #18741

* Add SECURITY.md policy file

Add a basic SECURITY.md file to establish the repository's security
reporting policy.  Includes guidance for reporting security issues
and what to expect.

Reviewed-by: Allan Jude 
Reviewed-by: George Melikov 
Signed-off-by: Brian Behlendorf 
Closes #18766

* Fix receive -x according to comment

The old condition skipped the -x for ANY property whose source wasn't 
explicitly ZPROP_SOURCE_VAL_RECVD - which caught inherited/default 
properties too, not just locally-set ones. The new condition correctly 
skips only when the property is locally-set on the destination 
(source == fsname), which is the documented intent.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Richard Kojedzinszky 
Closes #18737
Closes #18738

* Fix receive of split large blocks with a short trailing chunk

A dataset with a large recordsize can store a single-block file whose
block size is not a power of two.  When such a block is sent without
large blocks (no -L), the sender splits it into SPA_OLD_MAXBLOCKSIZE
(128K) chunks, and the final chunk is smaller than the block size.

flush_write_batch_impl() already handles any WRITE record whose size
differs from the object's block size by doing a normal dmu_write(), but
it first asserted the record was always larger than the block size.  The
shorter trailing chunk violates that assertion, so receiving such a
stream panicked the receive_writer thread on debug builds; production
builds took the correct dmu_write() path and were unaffected.

Drop the assertion and describe both size-mismatch cases in the comment;
the dmu_write() path already handles a record smaller than the block
size.  Add an rsend test that sends such a block without -L (initial and
incremental, plain and compressed) and verifies the received file
matches.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Issue #17829
Closes #18749

* ZTS: migration/setup: clear stale zfs_member label before new_fs

During a full ZTS run functional/migration/setup fails intermittently
when it mounts the non-ZFS device.  That device is often one an earlier
test used as a pool vdev.  'zpool destroy' leaves the vdev labels in
place and new_fs only overwrites the front of the device, so the
trailing labels can survive.  libblkid then probes the device as
ambiguous (both the new filesystem and zfs_member) and the
auto-detecting mount refuses, which setup reports as a spurious failure.

Wipe any residual signatures with wipefs before laying down the new
filesystem so the device carries a single, unambiguous type, and let
udev settle before the mount.  Skip the wipe in the single-disk case,
where the non-ZFS device is the same one the test pool was just created
on, so the live pool is left untouched.

Verified on Linux: after a pool create and destroy the scratch device
still carries a zfs_member label (blkid -p reports zfs_member); a
wipefs -a removes it so the following new_fs is the only signature and
the mount succeeds.  The functional/migration group passes with the
change.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #18492
Closes #18753

* ZTS: stop zpool_initialize tests racing initialize to completion

zpool_initialize_import_export and zpool_initialize_suspend_resume start
initializing a one-disk pool, wait a fixed couple of seconds, and then
expect initializing to still be running so it can be observed across an
export/import and suspended.  On a small or fast vdev the default 1 MiB
initialize chunk lets the whole disk finish within that window, after
which "zpool initialize -s" fails with "there is no active
initialization" and the test fails.

Throttle initializing with zfs_initialize_chunk_size, exactly as the
zpool_wait_initialize_* tests already do, so it stays active long enough
to observe regardless of vdev size or speed.  The tunable is saved and
restored per test.  Cleanup destroys the pool before restoring the chunk
size: the initialize thread rereads zfs_initialize_chunk_size on every
write but allocates its fill buffer once at the smaller size, so raising
it back while the thread is still running would issue a write larger
than that buffer.

With this both tests pass reliably, so drop their entries from the
zts-report.py.in maybe list (#11948, #17311).

Tests: ran both tests against the reproduced race in a VM -- with the
default chunk size the suspend step fails with "there is no active
initialization"; with the smaller chunk size initializing is still
running across the export/import and suspend, and the tunable is
restored on exit.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #11948
Closes #17311
Closes #18771

* zpl_inode: remove zpl_rename no-flags variants

Removed in 4.9.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18769

* CI: Fix race caused by shared ctr file updates

CTR is shared between the VMs and used as a global counter.  This 
uncoordinated shared access can result in a CI failure due to the
racing updates.  From the log:

`qemu-6-tests.sh: line 27: 1`
`6: syntax error in expression
(error token is "6")`

Resolve the issue by using separate ctr files by appending the ID.
The output now prints the total test cases count along with each VMs
individual count.

Reviewed-by: Tino Reichardt 
Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18778

* ZTS: fix zpool_initialize_multiple_pools suspend race

zpool_initialize_multiple_pools verifies that "zpool initialize -a -s"
suspends initializing on all four pools. It starts initializing and,
without slowing it down, expects every pool to still be initializing
when the suspend runs.

On fast storage a pool can finish initializing before the suspend, so
"zpool initialize -a -s" reports "there is no active initialization"
and the pool's status is "completed" rather than "suspended". The
suspend command then fails the test intermittently.

Throttle zfs_initialize_chunk_size before the suspend phase, the same
way the zpool_wait_initialize_* tests do, so initializing stays active
through the suspend and cancel checks. The earlier phases that wait for
initializing to finish ("zpool wait" and "-w -a") are left at the
default chunk size so they still complete promptly.

The chunk size is saved and restored, and the pools are destroyed
before it is restored so a running initialize thread never sees a
larger value than the buffer it allocated at the smaller size.
…
creatorcary pushed a commit to truenas/zfs that referenced this pull request Aug 26, 2026
#430)

* ZTS: Pass dec instead of hex to mknod

On Ubuntu 26.04 the default mknod command returns an error when
provided the major and minor numbers in hex.  Switch to passing
decimal values.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* zbookmark_compare: handle "marker" bookmarks with negative levels

"Marker" bookmarks (those with zb_level == ZB_ROOT_LEVEL, ZB_ZIL_LEVEL
or ZB_DNODE_LEVEL) represent valid blocks, but are associated with a
dataset directly rather than with a specific object within it. They end
up on bookmark lists during scan prefetch, and so need to be sorted
ahead of any "true" object blocks.

The problem is that for negative levels, BP_SPANB produces a negative
shift, which is not legal C. Fortunately the results are used only for
comparison, so the worst possible behaviour in a forgiving compilation
environment is a mis-sort, which for the scan/traverse cases, means that
we haven't prefetched certain metadata before we actually need it. But
there _is_ UB in there, and UBSAN does rightly complain.

Here we fix all this by handling these bookmarks directly - sorting them
ahead of "true" object blocks, which is usually what scan/traverse will
prefer. And we don't do any interesting math on these bookmarks, so we
sidestep the whole UB thing.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #14777
Closes #18652

* CI: Have zfs-build-packages workflow build tarballs on Alma (#18662)

Previously, zfs-build-packages would only build source tarballs
on Fedora due to problems with building them on RHEL 7.  That's
a relic of the past now, as we haven't supported RHEL 7 since
it went EOL in 2024.  With this change, we now build the tarballs
on both Alma and Fedora.

Signed-off-by: Tony Hutter 
Reviewed-by: Olaf Faaland 
Reviewed-by: Chris Longros 

* Update our CI runners to the newest FreeBSD 15.1 RELEASE (#18667)

Signed-off-by: Christos Longros 
Reviewed-by: Alexander Motin 
Reviewed-by: Tony Hutter 

* Linux 7.1 compat: META (#18682)

Update the META file to reflect compatibility with the 7.1
kernel.

Signed-off-by: Tony Hutter 
Signed-off-by: Rob Norris 
Reviewed-by: Chris Longros 

* CI: Re-allow workflow_dispatch on zfs-qemu

Allow zfs-qemu to be invoked from a workflow_dispatch event (a.k.a,
manually running a workflow).  This may have been accidentally disabled
in 1916c2c55.

Reviewed-by: Chris Longros 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18680

* initramfs-zfs should not try to copy directories

We had find only return files from the beginning for libgcc.so, but not
libfetch/libcurl. This oversight affected a user when vmware installed
its own libcurl.so.4 in a directory called libcurl.so.4, since our code
then tried to copy a directory, which fails.

Reviewed-by: Chris Longros 
Reviewed-by: Brian Behlendorf 
Suggested-by: Carsten Härle 
Signed-off-by: Richard Yao 
Closes #18582
Closes #18686

* Clean up embedded slog metaslab across txgs

On a read-write import, metaslab_set_fragmentation() can dirty a
metaslab via vdev_dirty() while still in the txg==0 load path when its
space map has an unexpected bonus size (e.g. a makefs-created pool
whose space-map dnodes use the boot loader's 24-byte space_map_phys_t
with nblkptr=3, giving db_size=64). If that metaslab is then selected
as the embedded slog, vdev_metaslab_init() only removed it from
vdev_ms_list when txg != 0, so the txg==0 case left it queued and
metaslab_fini() tripped VERIFY(!txg_list_member(&vd->vdev_ms_list,
msp, t)).

Remove slog_ms from the dirty list for every TXG_SIZE slot before
metaslab_fini() so the cleanup is correct regardless of txg.

Reported on FreeBSD as PR 281520:
External-issue: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=281520

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Nick Price 
Closes #18693

* README: update supported FreeBSD release to 15.1

Our CI runners moved to FreeBSD 15.1 in 0a4b59765 (#18667), but the
README still lists 15.0. Update it to match the CI version.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Christos Longros 
Closes #18696

* honor file argument in file_wait_event

grep the log path passed by the caller instead of always using
ZED_DEBUG_LOG.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Alek Pinchuk 
Closes #18700

* Update mtime/ctime when fallocate grows a file

Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes #18573

* ddt_log: Fix refcount tagging for begin/commit

Sponsored-by: Klara, Inc.
Sponsored-by: Wasabi Technology, Inc.
Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Igor Ostapenko 
Closes #18706

* CI: Increase default watchdog NMI timeout on Linux

When the watchdog driver is configured and enabled an NMI will be
generated when the watchdogd process fails to regularly reset the
watchdog timer.  Given the heavily virtualized and potentially
over-subscribed nature of the CI environment increase the default
timeout to 120 seconds (normally defaults to 30 seconds).

Reviewed-by: Christos Longros 
Signed-off-by: Brian Behlendorf 
Closes #18704

* Fix race between device removal completion and pool export

vdev_remove_complete() finalizes a device removal in two phases under
the spa lock framework. Between the two phases it called
spa_vdev_exit(), which drops both the config locks (SCL_ALL) and
spa_namespace_lock and blocks on a txg sync. By that point
vdev_remove_replace_with_indirect() has already set svr->svr_thread =
NULL, and that is the only thing the export path (spa_export_common()
-> spa_async_suspend() -> spa_vdev_remove_suspend()) waits on. Once the
namespace lock is dropped, a concurrent export or destroy can acquire
it and set spa->spa_export_thread. When the removal thread re-enters
for its second phase via spa_vdev_enter(), it trips the
ASSERT0P(spa->spa_export_thread) assertion.

Hold spa_namespace_lock across both phases instead of dropping and
re-taking it: the intermediate spa_vdev_exit() becomes
spa_vdev_config_exit(), which drops only SCL_ALL and syncs the txg
while keeping the namespace lock held, and the second spa_vdev_enter()
becomes spa_vdev_config_enter(). Because the namespace lock is never
dropped between the phases, a concurrent export blocks in
spa_namespace_enter() and cannot set spa_export_thread until removal
finalization is done. This mirrors the multi-phase locking pattern
already used by the attach/detach and split paths in spa.c.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Prakash Surya 
Closes #18657

* linux: handle mmap read beyond file size

When performing a mmap read past the end of a file there is no data to
read, so simply zero-fill the page and return success.  zfs_getpage()
limits the range lock appropriately to cover the offset being read.

Reported-by: Iliya Polihronov (@vnsavage) (Automattic)
Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18715

* Using net/cloud-init to unpin specific python

These days freebsd 15/16 fail when fetching py311-cloud-init.
Switch to net/cloud-init to avoid python version pinning.

Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18717

* Linux 7.2: zpl_super: convert to sget_fc()

The old sget() superblock matcher has been removed in favour of the
fscontext-based sget_fc(). This converts to it. Its largely a signature
change, no functional change.

sget_fc() has existed since fscontext was introduced, so there's no need
for separate feature tests.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18677

* zpl_ctldir: remove comments describing ancient kernel behaviour

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18722

* libzfs: fix MS_CRYPT/MS_OVERLAY collision with umount2(2) flags

MS_CRYPT and MS_OVERLAY are libzfs-internal mount flags, but their
values (0x8 and 0x4) aliased the umount2(2) flags UMOUNT_NOFOLLOW and
MNT_EXPIRE. A consumer that legitimately set UMOUNT_NOFOLLOW therefore
had that bit read as MS_CRYPT, so libzfs unloaded the dataset's
encryption key as a side effect.

Move both flags to high bits unused by umount2(2) and strip them before
the unmount syscall in do_unmount() (umount2(2) on Linux, unmount(2) on
FreeBSD) and in cmd/zfs. MS_CRYPT is a compile-time macro, so consumers
that set it (for example truenas_pylibzfs) must be rebuilt against the
new header.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes #18713

* ZTS: ctime_001_pos increase tolerance

The ctime_001_pos test checks that timestamp updates occur for a
file after performing certain operations (read, write, chown, etc).
The test case allowed for a +4 second tolerance in the timestamp
value which is generous but up to +7 second discrepencies have been
seen in the CI.  Bump the tolerance to +10 seconds to prevent these
false positives.  As long as the value increases and is reasonably
close to the expected value consider that to be sufficient.

Reviewed-by: Christos Longros 
Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18733

* build: Fix for building dist target outside of project root

Do not change the working directory when doing copying because $distdir
is a path relative to the working directory, and thus not valid after
changing the working directory. This previously worked, when configure
was run from the project root, because @srcdir@ was the same as the
build working directory.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Glenn Washburn 
Closes #18744

* Linux: fix zfs_write() infinite loop on unfaultable buffer

On Linux, zfs_write() copies from the source buffer with page faults
disabled while the transaction is open, relying on
zfs_uio_prefaultpages() to make the pages resident beforehand.  When
dmu_write_uio_dbuf() returns EFAULT, the retry path only subtracts
the bytes consumed so far from the prefault accounting.  On zero
progress pfbytes does not change, the "pfbytes < nbytes" check never
triggers another prefault, and the loop retries the same failing
copy.  If the buffer can never be faulted in again, e.g. the owning
process was torn down while a thread was in pwrite(2), that thread
spins unkillably at 100% CPU holding the file range lock, blocking
every other accessor of the file.

This is a regression from commit b0cbc1aa9a ("Use big transactions
for small recordsize writes."), which dropped the unconditional
re-prefault the EFAULT path had carried since commit 779a6c0bf6
("deadlock between mm_sem and tx assign in zfs_write() and page
fault").  Restore those semantics by resetting pfbytes on EFAULT, as
suggested in the issue analysis: the next iteration then faults the
pages back in before retrying, and for a permanently inaccessible
buffer zfs_uio_prefaultpages() fails, breaking the loop with EFAULT,
which is already propagated to userspace.  Transient faults, such as
mmap'ed source pages evicted under memory pressure, retry as before.

Built and tested on Linux aarch64: the ZTS mmap and write-path groups
pass and a munmap-versus-write stress run leaves no stuck writers.  The
teardown race itself is not reproducible on demand, so the retry paths
were also checked by inspection against the reproducers in the issue.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #17129
Closes #18740

* linux: batch DMU reads of non-resident pages in mappedread()

When a range being read has at least one page in the page cache,
zfs_read() routes the whole chunk through mappedread(), which falls
back to a separate dmu_read_uio_dbuf() call for every non-resident
PAGE_SIZE piece.  Since cached pages outlive munmap(), a file which
was mapped at some point may sit mostly outside the page cache and
still pay this cost: one DMU call per 4K page instead of one per
chunk, measured in #16031 as a 4-10x sequential read slowdown.
Commit 39be46f43 ("Linux 5.18+ compat: Detect filemap_range_has_page")
fixed the detection side so fully uncached chunks bypass mappedread()
again, but a chunk holding even one resident page still degrades to
page-sized DMU reads for everything else.

Instead of issuing one DMU read per absent page, probe the page cache
with find_get_page() and extend the read over the whole run of
non-resident pages which follows, restoring chunk-sized DMU reads for
the uncached parts of a mapped file.

A page can be faulted in after it was observed absent and before the
DMU read covering it completes, but this is safe for the same reason
the existing single-page window is.  zfs_read() holds the znode
rangelock as reader across mappedread(), so the DMU contents of the
range are stable: zfs_write(), zfs_putpage() and truncation all
require the writer lock, and zfs_getpage() fills concurrently faulted
pages from those same contents.  A faulted page can only diverge from
the DMU once dirtied through a writable mapping, making that store
concurrent with this read, for which returning the pre-store data is
a valid outcome.  Stores which completed before the read began cannot
be missed: a dirty page cannot be cleaned and reclaimed while the
reader lock is held (writeback takes the writer lock), so it is still
found resident, or zfs_putpage() already copied its data into the
DMU.

FreeBSD's mappedread() has the same per-page fallback and could be
batched the same way in a follow-up.

Measured in a VM with a 1 GiB file held in the ARC and read
sequentially with dd: one resident page per 1 MiB read request
degrades throughput from ~11.4 GB/s (no resident pages) to ~5.7 GB/s
on the baseline, and this change restores ~11.4 GB/s; one resident
page per 32 MiB chunk, read in 32 MiB requests, improves from
~5.4 GB/s to ~7.3 GB/s.  Reads of a fully resident file are
unaffected (~20 GB/s before and after).  All tests in the ZTS mmap
group pass, including the mmap_read and mmap_seek cases.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #16031 
Closes #18741

* Add SECURITY.md policy file

Add a basic SECURITY.md file to establish the repository's security
reporting policy.  Includes guidance for reporting security issues
and what to expect.

Reviewed-by: Allan Jude 
Reviewed-by: George Melikov 
Signed-off-by: Brian Behlendorf 
Closes #18766

* Fix receive -x according to comment

The old condition skipped the -x for ANY property whose source wasn't 
explicitly ZPROP_SOURCE_VAL_RECVD - which caught inherited/default 
properties too, not just locally-set ones. The new condition correctly 
skips only when the property is locally-set on the destination 
(source == fsname), which is the documented intent.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Richard Kojedzinszky 
Closes #18737
Closes #18738

* ZTS: migration/setup: clear stale zfs_member label before new_fs

During a full ZTS run functional/migration/setup fails intermittently
when it mounts the non-ZFS device.  That device is often one an earlier
test used as a pool vdev.  'zpool destroy' leaves the vdev labels in
place and new_fs only overwrites the front of the device, so the
trailing labels can survive.  libblkid then probes the device as
ambiguous (both the new filesystem and zfs_member) and the
auto-detecting mount refuses, which setup reports as a spurious failure.

Wipe any residual signatures with wipefs before laying down the new
filesystem so the device carries a single, unambiguous type, and let
udev settle before the mount.  Skip the wipe in the single-disk case,
where the non-ZFS device is the same one the test pool was just created
on, so the live pool is left untouched.

Verified on Linux: after a pool create and destroy the scratch device
still carries a zfs_member label (blkid -p reports zfs_member); a
wipefs -a removes it so the following new_fs is the only signature and
the mount succeeds.  The functional/migration group passes with the
change.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #18492
Closes #18753

* zpl_inode: remove zpl_rename no-flags variants

Removed in 4.9.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18769

* CI: Fix race caused by shared ctr file updates

CTR is shared between the VMs and used as a global counter.  This 
uncoordinated shared access can result in a CI failure due to the
racing updates.  From the log:

`qemu-6-tests.sh: line 27: 1`
`6: syntax error in expression
(error token is "6")`

Resolve the issue by using separate ctr files by appending the ID.
The output now prints the total test cases count along with each VMs
individual count.

Reviewed-by: Tino Reichardt 
Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18778

* Remove libuutil from the pull request template and fix headings

- libuutil was removed in adb316f41.
- Normalize section headers to title case.
- Drop the trailing colon on "Checklist".

Reviewed-by: Brian Behlendorf 
Reviewed-by: George Melikov 
Signed-off-by: Alexander Moch 
Closes #18791

* CI: Update Alpine Linux runner to 3.24.1

Update the Alpine Linux CI runner from 3.23.2 to 3.24.1.

This refreshes the runner to the latest Alpine release while keeping
the existing CI configuration unchanged.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Moch 
Closes #18790

* Rate limit Direct I/O verify zevents

Each vdev initializes a vdev_dio_verify_rl rate limiter (governed by
zfs_dio_write_verify_events_per_second), but
zio_dio_chksum_verify_error_report() never consults it, so
dio_verify_rd and dio_verify_wr zevents are posted with no rate
limiting.  A workload that repeatedly trips the Direct I/O verify can
therefore produce an unbounded flood of zevents.

Gate both ereport posts through zfs_ratelimit(&vd->vdev_dio_verify_rl),
as is already done for the other per-vdev ereports (checksum, delay,
deadman).  The vs_dio_verify_errors vdev stat still increments on every
event, so the true count remains observable via zpool status -d.

The dio_write_verify test checks on every iteration that a
dio_verify_wr zevent was posted.  With rate limiting now in effect the
shared limiter window is exhausted after the first iteration, so later
iterations observe zero events and the test fails.  Raise
zfs_dio_write_verify_events_per_second for the duration of that test
(restored in cleanup), adding the matching tunables.cfg alias, so the
events stay observable.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18795

* Fix reads for blocks freed after being cloned

PR #18421 fixed a case when reads for blocks cloned after being
freed could return zeroes.  But it created an opposite problem,
when reads for blocks freed after being cloned could return non-
zero content from the cloning.

This patch fixes the problem by creating a more specialized
form of dnode_block_freed(), taking into account the TXG when
the cloning has happened and checking frees only in TXGs after.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Gary Guo 
Signed-off-by: Alexander Motin 
Closes #18421
Closes #18724

* libzfs: fallback VDEV_UPATH to VDEV_PATH for non-DM devices

When zfs_get_underlying_path() returns NULL for a non-DM device (e.g.
NVMe), the zfs_prepare_disk script was getting an empty VDEV_UPATH.  Per
the man page, VDEV_UPATH should fall back to VDEV_PATH when there is no
underlying path.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MISAPOR LAB 
Closes #18439
Closes #18802

* L2ARC: bound the rebuild by the write hand on a first sweep

l2arc_log_blkptr_valid() ends with (!evicted || dev->l2ad_first), which
disables the eviction-overlap test entirely on a first sweep. That test
is meaningless in that state, since l2arc_evict() returns immediately
and l2ad_evict never advances off l2ad_start, but dropping it leaves the
log block bounded only by device geometry. On a first sweep only the
region below the write hand has been written, so a block at or beyond
l2ad_hand describes data this incarnation never wrote.
l2arc_hdr_restore() performs no per-entry validation, so every such
entry inflates arcstat_l2_psize and, via vdev_space_update(), the cache
vdev's vs_alloc. Nothing reconciles it because l2arc_evict() never runs
on a first sweep, and once vs_alloc exceeds vs_space the unclamped
subtraction in zpool(8) wraps and the device reports 16.0E free. Stale
entries also let the L2ARC read offsets that were never written, which
then fail checksum verification.

Removing and re-adding a cache vdev is enough to set this up:
l2arc_add_vdev() resets l2ad_hand to l2ad_start while the previous
incarnation's log blocks remain higher up the device, and the backward
walk then crosses l2ad_start and wraps into them.

Bound the first-sweep case by the write hand instead of disabling the
check, and state the geometry conditions once rather than duplicating
them across both branches.

The 16.0E symptom has been reported since 2015. #10224 carries the only
prior analysis, which suspected a leak in l2arc_evict() and was closed
incidentally by #9789 rather than by a fix. It is also open on FreeBSD
as PR 250323.

External-issue: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=250323
Reviewed-by: Ameer Hamza 
Reviewed-by: Alexander Motin 
Signed-off-by: Nick Price 
Closes #3114
Closes #3400
Closes #5583
Closes #10224
Closes #12779
Closes #18827

* zdb: output refcounts from verify_spacemap_refcounts()

Output information about refcounts when there is refcount mismatch.
Use plain uint64_t even as the values are expected to not be large.
Also use unsigned as we should never get negative refcounts there.

Reviewed-by: Allan Jude 
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Toomas Soome 
Closes #18809

* mmp: skip non-writeable vdevs during activity check

The import-time MMP activity check added by c710f8792 writes an
uberblock to each top-level vdev and requires a matching number of
good writes before the pool is claimed.  Two vdev types that carry
no writeable device break this:

  - A hole vdev (left by removing a log) and an indirect vdev (left
    by removing a data device) are counted in the required-write
    total but can never be written, so good_writes never reaches
    req_writes.  The activity check then spuriously fails and the
    pool is reported as held by another host with hostid 0.

  - mmp_claim_uberblock() also issues a zio_flush() to the root vdev
    after the writes.  zio_flush() recurses to every leaf, and an
    indirect vdev is a childless top-level vdev, so it is issued a
    ZIO_TYPE_FLUSH.  That trips the ZIO_TYPE_WRITE assertion in
    vdev_indirect_io_start() and panics.

Skip hole and indirect vdevs when counting required writes, and skip
non-concrete vdevs in zio_flush() as they have no device to flush.

Add hole, indirect, log, cache, and spare vdevs to the pool used by
the mmp_inactive_import, mmp_exported_import, and mmp_concurrent_import
tests so the activity check exercises these vdev types.  The enriched
layout is opt-in, leaving multihost_history on the simple two-device
pool it relies on.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18823
Closes #18835

* libspl: Implement VERIFY_IMPLY and VERIFY_EQUIV

The libspl debug header is missing VERIFY_IMPLY and VERIFY_EQUIV macros
and instead directly implements IMPLY and EQUIV.  Break out the VERIFY
definitions to match the kernel macros and facilitate code sharing
between kernel and userland.

Sponsored-by: Cybersecure Pty Ltd
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Ryan Moeller 
Closes #18822

* cstyle: better tolerance for struct literals

`cstyle.pl` currently doesn't have much patience for code such as:

```
*myvar = (mystruct_t) {
	.ms_field = 42,
	.ms_other_field = "chow time"
};
```

The first line is a Catch-22. If there's a space before the curly brace,
then it's an illegal cast because of the trailing space. If there isn't
a space, then it's an illegal curly brace without a preceding space.

Either way, tagging the first line as `/* CSTYLED */` gets you nowhere
because `cstyle.pl` doesn't understand the structure. It sees the
fields as continuation lines and complains about the indentation.

This PR makes three changes:

- It allows the first line with a space between the type and the brace.

- It adds first lines of this type to the same category as structs,
  enums, and unions. Indentation is tracked, but no particular
  style of indentation is enforced.

- It modifies a few clauses in `lib/libefi/rdwr_efi.c` that used to
  squeak through `cstyle.pl` but are now (correctly?) detected. These
  are of the form `(int) sizeof (type_t)`. I've changed these to
  `(int)(sizeof (type_t))`. Easy to reverse if the original form
  is in fact preferred.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Garth Snyder 
Closes: #18839

* DDT: Fix several bugs in pruning

- Fix variables types to avoid overflows after 2B entries.
 - Make ddt_prune_walk() code some more symmetrical.
 - Fix zero oldest on exact target to histogram value match.
 - Make bin 0 properly start from 0, not 1 hour.
 - Take as a cutoff base a time of histogram build start.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #18838

* dmu_recv: Avoid potential null deref

Compilers are smart enough to deref only if the first && operand is
true, so this is mostly to avoid false positives from sanitizers.

Sponsored-by: Klara, Inc.
Sponsored-by: Wasabi Technology, Inc.
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Igor Ostapenko 
Closes #18848

* ZTS: make file_check actually compare the resume test results

file_check guards every comparison with a check that the snapshot
directory exists on both sides, and the resume tests receive with -u,
so the receive side is never mounted and the .zfs snapshot paths never
exist.  The function has been quietly comparing nothing in
rsend_019-022, rsend_024, rsend_030 and send-c_resume, so the resume
test family verified that receives succeed but not that the received
data matches.

Mount both sides before diffing (some tests also unmount the send
side), still compare only the snapshots both sides carry since several
tests send just one of them, and fail loudly when nothing at all was
compared so the check cannot rot back into a no-op.  Two callers
needed their expectations fixed once the checks came alive: rsend_024
streams from the head rather than a snapshot, so it now diffs the
mounted heads directly, and the first file_check in
send_partial_dataset pointed at a partial dataset with no snapshots,
so it now compares against the dataset the stream came from.

Tests: rsend group passes with the comparisons active, twice in a row
on one module load.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #18834

* zed: let autoexpand see capacity changes on partitioned disks

Growing a disk under a whole-disk vdev never triggers autoexpand
(#12505).  The kernel reports a capacity change on the disk itself
and nothing for the partitions, whose sizes did not change.  But
since zfs owns the whole disk it carries a partition table, and
zed_udev_monitor() drops any disk-with-partitions event on the
assumption that a partition event will follow.  For a resize none
ever does, so the ESC_DEV_DLE event that zfsdle_vdev_online() needs
is never generated and the pool stays at the old size until someone
runs zpool online -e by hand.  This is the common case for cloud
disks grown online.

Pass change events through when udev marks them RESIZE=1.  On the
matching side a disk-level event has no vdev guid to search by (the
label lives on the partition), and udev provides no ID_PATH on some
buses, so the physical path lookup can also come up empty.  When
that happens, read the ZFS label off the whole-disk partition and
match by the pool and vdev guids stored in it.  Unlike matching the
config path textually, this works no matter which name the pool was
imported with (by-id, by-path or a bare device node), and a stale
device path in an unrelated pool's config cannot steal the event,
since the label names the owning pool.  The fallback only runs for
guid-less disk events and only accepts a whole-disk vdev that is
not a spare or l2cache device.

The new zpool_expand_006_pos test covers this end to end on a
scsi_debug disk.  zpool_expand_001_pos already grows a scsi_debug
disk the same way but keeps passing on an unpatched zed, because
block_device_wait issues a bare udevadm trigger, which re-sends
change events for the zfs_member partitions and hands zed the vdev
guid the resize itself never delivered.  The new test drains the
pool-creation udev traffic and then only settles, so zed sees what
a production resize generates: one RESIZE=1 change event on the
disk.

Multipath maps take a different path through zed_udev_monitor() and
still need the manual online; that is unchanged here, and the same
goes for other device-mapper vdevs, whose partitions live on
separate dm nodes the fallback cannot derive from the map's name.
Disks with no devid source at all (virtio-blk, Xen) also stay out
of scope: their events are dropped earlier for lack of any device
identifier, and widening that is its own discussion.

Tests: zpool_expand_006_pos fails against unpatched zed and passes
with the fix; a pool imported by /dev/disk/by-id expands from the
bare disk event with "matched vdev ... by the label" in the zed
log; the zpool_expand group passes.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #12505
Closes #18833

* zfs: fix stale POSIX ACL cache after rollback

An online zfs rollback rezgets live znodes and clears the OpenZFS ACL
cache, but leaves the Linux VFS inode POSIX ACL cache intact. A later
non-root permission check can use an ACL added after the snapshot.

Reproduce by snapshotting a POSIX ACL file whose group mode bits require
a VFS ACL check, granting a named user read access, holding the inode
active, and rolling the mounted filesystem back. The named user remains
able to read until cache eviction.

Invalidate both the access and default VFS POSIX ACL caches from
zfs_rezget(), alongside the existing private cache invalidation, so
subsequent permission checks reload the recovered on-disk ACL.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Wang Zhaolong 
Closes #18837

* Add missing checks to zfs_clone_range_replay()

zfs_clone_range() does a few error checks that we are missing in
zfs_clone_range_replay().

Reported-by: Grok 4.5 Build Beta
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Richard Yao 
Closes #18867

* libzfs: Do not call munmap() when mmap() fails

Reported-by: Grok 4.5 Build Beta
Reviewed-by: Igor Kozhukhov 
Reviewed-by: Rob Norris 
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Richard Yao 
Closes #18869

* libzfs: don't truncate a resolved vdev path in zpool_vdev_name()

zpool_vdev_name() copied the result of realpath() into a 64 byte stack
buffer shared with the short formatted names, so a vdev whose resolved
path was longer than 63 bytes was reported cut short by zpool status -L
and anything else asking for VDEV_NAME_FOLLOW_LINKS.  The cut is made at
a byte boundary, so a multi-byte character straddling it is left as
invalid UTF-8, which also makes zpool status -j -L emit JSON a parser
rejects.

Resolve straight into a buffer of the right size.  realpath() fills a
caller supplied buffer of at least PATH_MAX, as it is used elsewhere in
libzutil, which also removes the intermediate allocation.

The other three users of that buffer format a guid, a raidz name, or a
draid name, and none of them can exceed its length.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18851
Closes #18871

* libzfs: String trimming should not operate out of bounds

Forward slashes are trimmed from ZPOOL_IMPORT_PATH, but if someone sets
a ZPOOL_IMPORT_PATH that is only forward slashes, our trim code will
underflow the string, causing an out of bounds operation.

Similarly, the SMB code could potentially trim a string consisting of
only new line characters until it experiences the same bug. The same fix
is applied to it.

These are memory bugs, but I suspect that it is very unlikely that they
would cause a problem, since the probability that an out of bounds write
would follow the out of bounds read should be low. That said, going out
of bounds is undefined behavior, which could cause incorrect code
generation should a compiler look at it wrong, so let us fix this.

Reported-by: Grok 4.5 Build Beta

Signed-off-by: Richard Yao 
Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Closes #18868

* mmp: do not require writes to mirror legs the config marks absent

The MMP uberblock claim requires one good write per configured leaf of
each top-level vdev.  For a mirror it required two writes
unconditionally (MIN(MAX(children, 1), 2)), so a mirror with a leg that
is persistently offline, faulted, or removed could produce only one good
write and the activity-check claim failed with EIO.  A degraded mirror
could therefore not be imported with multihost=on, blocking HA failover.

Count only the legs the pool config still expects to be present, and
require a write to every one of them.  A leg taken out of service is
recorded persistently in the config and is seen the same way by every
host, so it is not required.  A leg merely unreachable from the
importing host keeps none of those states and stays required, so a host
that can see only some of the legs of an otherwise healthy mirror still
fails the claim and cannot split the pool.

The previous cap of two writes was a compromise made because requiring
every child was too strict for wide mirrors, in particular where a leg
is left offline for long periods as part of a backup strategy.
Consulting the config covers that case directly, so the cap is no longer
needed: on a three-way mirror with one leg unreachable and not marked
absent, a cap of two would accept the claim while another host holding
the third leg could accept it as well.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18855

* ZTS: add coverage for the MMP claim on a degraded mirror

Add mmp_degraded_import, which verifies the uberblock claim requires a
write only to those mirror legs the pool configuration still expects to
be present.  A healthy mirror is claimed, a mirror with an offlined leg
is claimed and imports degraded, and a mirror with a leg this host
cannot open, and which the configuration does not mark absent, is
refused.  The degraded and unreachable cases repeat on a three-way
mirror, where the number of legs the configuration expects and the
number this host can reach come apart.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18855

* libzfs: don't read a dataset handle after closing it in resume send

zfs_send_resume_impl_cb_impl() closes the dataset handle before its
error switch and then reads zhp->zfs_name again in the ESRCH case.
zfs_name is an array declared inside struct zfs_handle, so the free()
at the end of zfs_close() releases it along with the handle, and
lzc_exists() copies out of the freed block.

The close dates from the original resume send.  The ESRCH case was added
three years later with redacted send, below a handle that was no longer
live.  The path is reachable: dsl_bookmark_lookup() returns ESRCH when
the incremental source is a bookmark that has gone away, which a resume
can lose a race with.

Keep a copy of the name alongside the error message that is already
formatted before the close, and test that instead.

Reported-by: RigelYoung <43904538+RigelYoung@users.noreply.github.com>
Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18870
Closes #18883

* man: zvol_request_sync is not ignored under blk-mq

The zfs.4 entry for zvol_request_sync claims it "is ignored when
running on a kernel that supports block multiqueue (blk-mq)". This
has never been true. The sentence and the blk-mq support it describes
landed in the same commit (6f73d0216 "zvol: Support blk-mq for better
performance"), which added

	if (zvol_request_sync)
		force_sync = 1;

to zvol_request_impl() -- the function shared by both submission
paths, reached from zvol_submit_bio() and from zvol_mq_queue_rq()
alike. No blk-mq guard was added there then, and none exists now.

Taken literally the sentence is worse than inaccurate: HAVE_BLK_MQ
was removed in 9601eeea1 because every supported kernel has blk-mq,
so the documented condition is always satisfied and the parameter
would never do anything.

Verified on 6.12.101 by counting call sites with kprobes during an
fio run. zvol_write() is reached either via the taskq, through
zvol_write_task(), or directly, so the two are a clean discriminator:

  config             zvol_write  zvol_write_task  zvol_tq CPU
  bio, default          1477413          1477413  7902 jiffies
  bio, sync=1           3873491                0            0
  blk-mq, default       2787187          2787187  7948 jiffies
  blk-mq, sync=1        3816106                0            0

With zvol_request_sync=1 no task is ever dispatched and the zvol
taskq threads get no CPU, on either path.

Replace the sentence with an accurate one.

Reviewed-by: Brian Behlendorf 
Signed-off-by: George Melikov 
Closes #18887

* CI: publish the per-VM test counter atomically

Each VM's log reader keeps a running test count in /tmp/ctr-vm$ID and
reads every other VM's counter to print the combined progress figure.
The counter is published with a plain redirect, which truncates the file
before it writes, so a reader can see it empty.  `read` then returns 1,
and because the reader inherits set -eu that ends the reader subshell.

The VM keeps running and its results are recovered later from the
artifact, but its output stops being prefixed into the live log from
that point on, with nothing said about why.

Seen on fedora44 in
https://github.com/openzfs/zfs/actions/runs/30772421272.  vm2's last
prefixed line is refreserv/cleanup at 01:57, carrying counter 745, while
vm1 keeps reporting vm2 frozen at 746 for the remaining 38 minutes: the
reader incremented the counter and died before printing that line.  vm2
itself ran on until 02:14 and its 899/11/6 summary never reached the
live log.

Publish through a temporary file and rename instead.  A reader then sees
either the old value or the new one.  Racing a writer against a reader
20000 times reproduces 10805 failed reads with the redirect and none
with the rename.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18885

* CI: don't fail a passing job when a log reader has already exited

Once a VM's tests finish the runner kills that VM's log reader.  If the
reader is already gone, kill reports ESRCH, and since the script runs
under set -eu that ends it and the job is reported failed after the
tests have already passed.

That is what turned the fedora44 run in
https://github.com/openzfs/zfs/actions/runs/30772421272 red:

  vm1: Results Summary
  vm1: PASS  1151
  vm1: FAIL     2
  vm1: SKIP     6
  ...
  qemu-6-tests.sh: line 143: kill: (20700) - No such process
  ##[error]Process completed with exit code 1.

All 13 failures in that run were on the expected list and neither VM
reported an unexpected one, so no test result was affected.  Only the
exit status was wrong.

The preceding commit removes the race that killed the reader, so this
should no longer be reachable, but a reader can still die for reasons
this script does not control and losing a whole run to the cleanup step
is a poor trade.  The message is left on stderr rather than discarded,
because a reader exiting early means live output was lost and that is
worth seeing.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18885

* build: Fix release detection when build dir is not source dir

When building outside of the root source directory, configure fails to
detect that the source for the build is a git repository because the
build directory is checked if it is a git repository. Instead check
source directory and use the source directory for generating the
release. Check for the .nogitrelease file in the source directory too.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Glenn Washburn 
Closes #18891

* ZTS: don't read a command's exit status as a missing binary

log_neg_expect() treats an exit status of 127 as a missing binary and
fails before it looks at what the command printed.  That is only a
convention of the shell, and a command is free to return 127 for its
own reasons.  fio returns the number of jobs which failed, so a run of
127 failing jobs is reported as though fio were not installed.

no_space/enospc_rm fills a pool with 200 fio jobs and requires them to
fail with ENOSPC.  How many of them get that far varies with timing,
and on the occasions it comes to exactly 127 the test fails with

  fio ... unexpectedly exited 127 (File not found)

even though the output holds the expected message.  A dozen runs here
landed between 183 and 199, and the failure seen in CI reported 127.

Only read 127 as a missing binary when the expected output is absent.
A command which printed what was asked of it plainly ran, so nothing
which passed before can start failing.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18726
Closes #18882

* nvpair: Fix operator precedence

59dc88602e23a436440e4164c6d9401da8f0dff2 made a mistake when doing a
check, which can cause us to continue processing when we should return
EFAULT.

Reported-by: Grok 4.5 Build Beta
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Richard Yao 
Closes #18874

* Linux 6.18 compat: vfs_parse_fs_string() takes 3 args

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18847

* Linux 6.3: follow_down() gains flags arg

We need the flags arg to trigger the snapshot mount. For earlier
kernels, we can emulate it with vfs_path_lookup()

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18847

* Linux 5.19/6.17: handle differences in how to flush delay workqueue

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18847

* CodeQL: Flag implicit compare-then-assign in branch conditions

Implicit compare-then-assign in branch conditions is buggy since
developers often mean assign-then-compare, but sometimes actually mean
compare-then-assign. GCC's -Wparentheses was originally meant to catch
assignment in place of comparison, requiring an extra set of parentheses
to turn this off. This had the happy coincidence of making developers
explicit about assign-then-compare vs compare-then-assign.

An outer level of extra parentheses will inhibit -Wparentheses warnings.
This often results in assign-then-compare being made explicit, but
instead of turning `if (x = foo() < 0)` into `if ((x = foo()) < 0)`, a
developer might write `if ((x = foo() < 0))`, which turns off the
warning, without fixing the problem. This happened in openzfs/zfs#18874.
There are other potential variations, such as `if ((x = (foo()) < 0))`,
which also suppresses GCC's warning, but fails to actually do anything
since the intended explicit parentheses to specify compare-then-assign
are around the right operand of the boolean operator, rather than around
the boolean operator, yet we have the additional parentheses needed to
silence GCC's -Wparentheses. In the `if ((x = (foo()) < 0))` case, the
intent was to make compare-then-assign explicit, and a typo caused it to
fail to become explicit. That is not a bug, but it makes it unclear what
the developer intended, which is problematic in itself.

This probably merits a bug report to GCC requesting a more intelligent
diagnostic that will treat compare-then-assign differently from
assignment in a branch condition. However, that is a slow process, this
has already bitten us once and with CodeQL, we can add our own check to
the PR process so that we catch other instances of this issue during
review, rather than some time later.

Given that assign-then-compare in branch conditions requires that
parentheses be added in such a way that the compiler AST no longer
contains an implicit compare-then-assign, we only need to check for an
implicit compare-then-assign in order to implement this check. Although
the likelihood of compound assignment being present in this bug pattern
is low, the same logic follows, so the check also will catch this
pattern on compound assignment. This check handles conditions in if,
while, do, for, ?:, && and ||. switch statements are intentionally
ignored, since using assign-then-compare in a switch statement would
turn the switch statement into a if-else. That is pointless, so allowing
an implicit compare-then-assign in switch statements is problem-free.
Coincidentally, GCC's -Wparentheses does not apply to switch statements
either. Finally, this considers all comparison operators, rather than
just the < operator used in the examples in this commit message.

The CodeQL check was written by Grok 4.5 Build Beta after several
iterations of prompt engineering and follow-up prompts to give it
corrections. It has also been subjected to a test suite of 19 true
positives and 17 true negatives to verify its behavior. It successfully
detected all true positives and fails to detect any true negatives. It
has also been applied not only to the OpenZFS codebase, but also the
Linux kernel and curl codebases, where it had zero detections. Related
queries in CodeQL were also run against the test suite, but had zero
detections. The query appears to be a well made query that has a very
high signal-to-noise ratio. It might be worth submitting to upstream
CodeQL for inclusion, but I would rather add it to our own repository so
we can begin benefiting from it today.

Assisted-by: Grok 4.5 Build Beta
Reviewed-by: Brian Behlendorf 
Signed-off-by: Richard Yao 
Closes #18899

* nvpair: Improve native handling of unterminated strings

This continues the work done in 59dc88602e23a436440e4164c6d9401da8f0dff2
and parallels what is already done for XDR encoding.

Reported-by: Grok 4.5 Build Beta
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alek Pinchuk 
Signed-off-by: Richard Yao 
Closes #18876

* nvpair: i_get_value_size() string array handling tweak

The strnlen() function needs to be given the length of the remaining
region to behave as intended, but it was given the length of the total
region on packed strings.

Reported-by: Grok 4.5 Build Beta
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alek Pinchuk 
Signed-off-by: Richard Yao 
Closes #18877

* Linux 7.2 compat: META

Update the META file to reflect compatibility with the 7.2
kernel.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18941

* Fix ddtprune causing space leak

In zio_ddt_free, if a pruned dde is still in ddt, it would do nothing
and cause space leak.

Reviewed-by: Rob Norris 
Reviewed-by: Brian Behlendorf 
Reviewed-by: Allan Jude 
Signed-off-by: Chunwei Chen 
Closes #17982
Closes #17983

* Make systemd-udev-settle optional for the import units

systemd-udev-settle.service has been deprecated for years, recent
systemd releases warn about it at boot, and distributions have begun
shipping without it, which turns the hard Requires= in
zfs-import-cache and zfs-import-scan into a broken import: a
Requires= on a masked or removed unit keeps the service from ever
starting.  Issue #10891.

Demote the dependency to Wants= and keep the After= ordering.  Where
the settle unit exists and completes, the boot is what it always
was: Wants= pulls settle in, the import waits for it, and the
device-symlink guarantee it provided is intact.  Where it is masked
or gone the wish is quietly dropped and the import runs anyway,
which beats not importing at all.  One deliberate behavior change:
if settle itself fails, on a system so large that enumeration
overruns its timeout, the old Requires= cancelled the import while
the new units go ahead at the timeout mark with whatever has been
enumerated by then.

Settle-less boots lose the wait for the udev queue, so the import
units gain two ordering edges in its place.
After=systemd-udev-trigger.service makes sure the coldplug events
are at least queued.  After=systemd-modules-load.service closes a
condition race the settle wait used to hide: both import units gate
on ConditionPathIsDirectory=/sys/module/zfs, and without the
multi-second settle delay that condition could be evaluated before
modules-load.d had finished loading zfs.ko, silently skipping the
import on an otherwise healthy boot.  Beyond that, a device whose
symlink appears a moment too late is only covered by the short
zfs_vdev_open_timeout_ms open-retry window; waiting for exactly the
devices a pool needs is what the per-pool import work (#18486) is
shaped to solve, and this stays the minimal step that keeps
settle-less systems importing today.

The stray After=systemd-udev-settle lines in zfs-mount, zfs-mount@
and zfs-volume-wait were only ordering hints against a unit that may
not exist, so they simply go away.

Tests: on a systemd 259 VM with a cachefile pool, rebooted with the
rendered unit: settle available, the boot pulls it in and the pool
imports as before; settle masked, the boot comes up with no failed
units and the pool still imports, where a masked settle previously
kept zfs-import-cache from starting at all.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Issue #10891
Closes #18832

* zhack: add "mmp reclaim" to recover a pool stranded by MMP

When a host fails together with the mirror legs attached to it, the
surviving labels still describe those legs as present, so the MMP
uberblock claim keeps demanding a write to every one of them and no
later import can satisfy it.  The pool cannot be imported by any host
again.

Add "zhack mmp reclaim", which imports once with the claim's required
write count relaxed for the mirror legs this host cannot open, marks
those leaves offline so that the ordinary imports which follow
succeed, and exports.

The relaxation is confined to userspace.  mmp_claim_relaxed is
declared under #ifndef _KERNEL, and module/Kbuild.in builds the module
with -D_KERNEL, so the flag cannot exist in a kernel module.  libzpool
does not define _KERNEL and so gets the check.  This follows the
zfeature_checks_disable pattern zhack already uses around the same
import, and is stronger, since that flag does exist in the kernel.

Only the number of required writes changes.  The write, the wait and
the re-read of the activity check are untouched, so a competing host
which shares any leg with this one is still detected and the import is
refused.  A live host whose legs are all invisible from here cannot be
detected by any write-and-read scheme, so this stays a manual
operation which assumes the peer has been fenced.

Legs are forgiven only under a top-level mirror, which is where the
relaxation lives, and exactly those legs are marked offline.  A raidz
or draid member is required as parity+1 in aggregate and never
demanded individually, so an absent one does not raise the requirement
and is left alone.  Offline is used rather than removed because it
persists unconditionally, is already excluded from the claim, and has
"zpool online" as its inverse when the hardware returns.

Reviewed-by: Brian Behlendorf 
Suggested-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18892

* ZTS: add coverage for zhack mmp reclaim

Six scenarios: a stranded pool is recovered and the claim then accepts
it, a live host sharing a leg is still refused, both top-level vdevs
are counted after a recovery, a pool without multihost is left alone,
a log vdev leg is not touched, and a raidz member is not touched.

Every recovery assertion re-imports as a third hostid.  zhack exports
cleanly under its own hostid, so importing again as the same host
takes the exported-and-matching-hostid path, skips the activity check
entirely, and would leave the claim unexercised and the test vacuous.

The assertions read req_writes and good_writes from the claim's own
dbgmsg line, which 20176224e added.  That is the only observable of
the claim arithmetic, at the cost of coupling the test to a debug
message this change does not control.

mmp_pool_destroy() used a bare "pgrep zhack", which matches any
process whose name merely contains zhack.  A ksh script named
mmp_zhack_reclaim.ksh has comm "mmp_zhack_recla", so the helper found
the running test and killed it.  Match the process name exactly.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18892

* mmp: tell a failed uberblock claim apart from remote activity

When the claim could not write to every device the config expects
present, spa_activity_check_claim() replaced the error from
mmp_claim_uberblock() with EREMOTEIO, so an operator whose peer died
together with its mirror legs was told another host holds the pool,
which sends them looking for a host that is not there.

Report the cause instead.  A shortfall has two causes worth telling
apart, so mmp_claim_uberblock() now counts the writes it issues
alongside the ones that succeed.  A leaf the config expects present but
which cannot be written is never issued one, so too few issued means a
device is absent, which persists across retries and is what
"zhack mmp reclaim" recovers; that returns ENODEV.  Enough issued but
too few good means the writes reached present devices and failed, which
a retry may clear; that stays EIO.  The issued count is gated exactly as
the good count is so the two describe the same set of leaves.

Both get a case in spa_ld_activity_result() and both still return
EREMOTEIO to userspace, as the ENXIO case already does, and the cause
travels to userspace in ZPOOL_CONFIG_MMP_RESULT so zpool(8) can say
which one it was and, for ENODEV, name the recovery.

ZPOOL_CONFIG_MMP_STATE stays MMP_STATE_ACTIVE for both even though
nothing is active.  An older zpool(8) knows only the two existing
states and reaches zfs_error_aux() with an uninitialized buffer for
anything else, so the state is kept as one it understands.  An older
zpool(8) against this kernel therefore prints what it prints today, and
a newer zpool(8) against an older kernel finds no cause reported and
falls back to the same text.

The paths where the claim genuinely detects another host still return
EREMOTEIO and are unaffected.

Also correct the comment above the write count, which still described
the fixed two writes per mirror that 8cdd9b2b7 replaced with one write
per leg the config expects present.

mmp_degraded_import.ksh asserted on the old message in the two cases
which are now ENODEV, and is updated with them.  The zhack case asserts
both directions, since a message that stops being emitted fails
silently: the shortfall must be reported, and it must not be reported
as another host holding the pool.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Michael Heller 
Closes #18892

* config: detect idmap method via generic_permission test

Since the switch from implicit to explicit userns, and to idmap,
happened right across the kernel in major releases, so it is enough to
use a single test and apply the results everywhere.

generic_permission() is a nice simple function with a simple interface,
so useful for an unambiguous test.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18769

* ZTS: test secpolicy_sys_config correctly limits namespace access

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18959

* ZTS: test secpolicy_zinject correctly limits namespace access

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18959

* secpolicy_nfs: remove, not used

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18959

* secpolicy_zinject: only permit a global zone credential

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18959

* secpolicy_sys_config: only permit a global zone credential

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18959

* secpolicy_zfs: add a note about the power of CAP_SYS_ADMIN

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18959

* vdev_open: pass credential to check for permission to open device

This commit adds a cred_t parameter to vdev_open() and threads it
through to all the vdev_op_open callbacks. The default is CRED(), ie the
credential of the calling task, which is usually some userspace control
process calling ioctl().

To handle the parallel vdev open case, we take additional and additional
reference to the cred for each task, and pass it down to vdev_open().

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18960

* zfs_file_open: add cred arg, use it to check access

If we're opening a file on behalf of the user, we need to ensure that
that user actually has access to it. Add a credential parameter to
zfs_file_open() and use it when opening the file.

Existing callers use kcred for now to get the same behaviour as before.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18960

* vdev_file: use calling cred to check for device access

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18960

* vdev_disk: use calling cred to check for device access

bdev_file_open_by_path() does not do any kind of credential check on the
given device path, so we need to do our own. We temporarily swap in the
passed in credential as the task credential, then call kern_path() and
inode_permission(), which together will ensure the credential can both
see and access the given path.

For the reopening case, we use the kernel credential. The idea here is
that since the device was already open, we shouldn't fail to reopen just
because the calling user can't see it, which would prevent device
removal, offline, online, etc.

Include some light reorganising in the error paths, since we might not
always have a device handle to carry the current error.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18960

* ZTS: device access tests

Tests that zpool create, add, attach and import all properly enforce the
restrictions on the calling user to access device nodes.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18960

* [zfs-2.3.9] Add 'capsh' to commands.cfg

Add missing 'capsh' to commands.cfg.  It was included in
master with 7839c4b5e1 but that was not backported to this branch.

Signed-off-by: Tony Hutter 

* [zfs-2.3.9] Add workaround for device_access ZTS test

Add workaround to get the device_access ZTS tests working on 2.3.9.

Signed-off-by: Tony Hutter 

* Tag zfs-2.3.9

META file and changelog updated.

Signed-off-by: Tony Hutter 

* vde…
snajpa pushed a commit to vpsfreecz/zfs that referenced this pull request Sep 23, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and did not log the change. Pass log=TRUE to zfs_freesp()
on the non-KEEP_SIZE extend path so the timestamps advance and the
size change lands in the ZIL; punch-hole and zero-range already use
that path. Matches zfs_space().

Upstream: 1d98dfe
Porting-notes: applies unchanged; includes the upstream
fallocate_extend_timestamps ZTS test.
Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
pull Bot pushed a commit to A-Archives-and-Forks/openzfs that referenced this pull request Sep 26, 2026
Commit 312bdab advertises STATX_ATTR_CHANGE_MONOTONIC and builds
the NFSv4 change_cookie from (ctime.tv_sec << 32) | zp->z_seq.
zp->z_seq is reset to a magic constant in zfs_znode_alloc(), so any
event that drops the znode from cache (memory pressure, remount,
reboot) regresses the lower bits of the cookie, a backward step
within the same second.

NFSv4 clients that trust this contract treat a regressed cookie as
evidence that the file's metadata cannot be relied on. VMware ESXi
over NFSv4.1 surfaces this as "The file specified is not a virtual
disk", and a VM stored on the affected NFS-exported ZFS dataset
fails to power on.

Widen z_seq to 64 bit and present it directly as the change_cookie,
dropping the ctime packing, so the cookie is a single monotonic
counter that no longer depends on the clock. FreeBSD's va_filerev
consumer also takes the wider value.

Persist z_seq via a new SA attribute SA_ZPL_SEQ. An in-core marker
zp->z_has_seq records whether the file already carries SA_ZPL_SEQ in
its layout; it is derived at load time and never stored on disk, so
no global pflag bit is consumed. ZFS_SEQ_MAY_GROW() keys off the
marker to grow the SA layout only on the first add per file;
ZFS_PERSIST_SEQ() then sets the marker and adds SEQ to the caller's
bulk alongside the file's other SA attributes. zfs_znode_alloc()
restores z_seq from SA_ZPL_SEQ when present and sets the marker;
zfs_rezget() recomputes the marker in place on rollback/recv without
disturbing the in-core z_seq, keeping the cookie monotonic.

A file written before this change carries no SA_ZPL_SEQ; on Linux it
is seeded with (ctime.tv_sec + 1) << 32 so the counter starts above
any pre-change cookie and stays monotonic across the upgrade. A
missing attribute is simply treated as not-yet-migrated, not an
error. FreeBSD never folded ctime into va_filerev, so it needs no
seed.

No feature flag or on-disk format change is needed: the new SA
attribute is keyed by name, so an implementation that does not know
it preserves it opaquely, and the first modify lazily migrates each
file. Covers both the Linux and FreeBSD ZPL.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
pull Bot pushed a commit to A-Archives-and-Forks/openzfs that referenced this pull request Sep 26, 2026
Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
pull Bot pushed a commit to A-Archives-and-Forks/openzfs that referenced this pull request Sep 26, 2026
The SA bulk arrays are filled by series of conditional SA_ADD_BULK_ATTR
calls whose worst-case count is easy to miscount, so a future attribute
add could silently overrun the fixed-size array.
Add ASSERT3S(count, <=, ARRAY_SIZE(bulk)) before each sa_bulk_update so
an overrun trips in debug builds, and use ARRAY_SIZE so the bound stays
tied to the declaration. Linux zfs_setattr allocates its arrays on the
heap sized by 'bulks', so it asserts against that instead.
Also tighten FreeBSD zfs_setattr's xattr_bulk from [7] to [6], its
actual worst case.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes openzfs#18573
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Status: Accepted Ready to integrate (reviewed, tested)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants