Skip to content

zfs: suppress reclaim lockdep in zfs_inactive - #18505

Closed
Gality369 wants to merge 1 commit into
openzfs:masterfrom
Gality369:linux-zinactive-reclaim-lockdep
Closed

Gality369 wants to merge 1 commit into
openzfs:masterfrom
Gality369:linux-zinactive-reclaim-lockdep

Conversation

@Gality369

@Gality369 Gality369 commented May 7, 2026 •

Copy link
Copy Markdown
Contributor

Motivation and Context

A possible circular locking dependency involving kswapd was reported
in the zfs_inactive/zfs_zinactive path. The warning occurs when reclaim
context reaches the synchronous unlinked-deletion path.

WARNING: possible circular locking dependency detected
------------------------------------------------------
kswapd0/55 is trying to acquire lock:
ffff88801d089f18 (&zfsvfs->z_teardown_inactive_lock){.+.+}-{4:4}, at: zfs_inactive+0x194/0x1300 fs/zfs/os/linux/zfs/zfs_vnops_os.c:4084

but task is already holding lock:
ffffffff87abcaa0 (fs_reclaim){+.+.}-{0:0}, at: balance_pgdat+0xc06/0x18b0 mm/vmscan.c:7137

which lock already depends on the new lock.

the existing dependency chain (in reverse order) is:

-> #3 (fs_reclaim){+.+.}-{0:0}:
       __fs_reclaim_acquire mm/page_alloc.c:4264 [inline]
       fs_reclaim_acquire+0xe6/0x120 mm/page_alloc.c:4278
       might_alloc include/linux/sched/mm.h:318 [inline]
       slab_pre_alloc_hook mm/slub.c:4929 [inline]
       slab_alloc_node mm/slub.c:5264 [inline]
       __do_kmalloc_node mm/slub.c:5649 [inline]
       __kvmalloc_node_noprof+0xef/0xa20 mm/slub.c:7112
       spl_kvmalloc+0x102/0x140 fs/zfs/os/linux/spl/spl-kmem.c:146
       spl_kmem_alloc_impl fs/zfs/os/linux/spl/spl-kmem.c:260 [inline]
       spl_kmem_alloc+0x240/0x260 fs/zfs/os/linux/spl/spl-kmem.c:467
       rrn_add fs/zfs/zfs/rrwlock.c:107 [inline]
       rrw_enter_read_impl+0x2c8/0x970 fs/zfs/zfs/rrwlock.c:186
       rrw_enter_read fs/zfs/zfs/rrwlock.c:198 [inline]
       rrw_enter+0x57/0x70 fs/zfs/zfs/rrwlock.c:235
       dsl_pool_config_enter fs/zfs/zfs/dsl_pool.c:1431 [inline]
       dsl_pool_hold+0x105/0x130 fs/zfs/zfs/dsl_pool.c:1403
       dmu_objset_hold_flags+0xae/0x1b0 fs/zfs/zfs/dmu_objset.c:751
       dmu_objset_hold+0x2c/0x40 fs/zfs/zfs/dmu_objset.c:772
       zpl_mount_impl fs/zfs/os/linux/zfs/zpl_super.c:394 [inline]
       zpl_mount+0xb6/0x770 fs/zfs/os/linux/zfs/zpl_super.c:465
       legacy_get_tree+0x112/0x220 fs/fs_context.c:663
       vfs_get_tree+0x9a/0x370 fs/super.c:1758
       fc_mount fs/namespace.c:1199 [inline]
       do_new_mount_fc fs/namespace.c:3642 [inline]
       do_new_mount fs/namespace.c:3718 [inline]
       path_mount+0x5b8/0x1ea0 fs/namespace.c:4028
       do_mount fs/namespace.c:4041 [inline]
       __do_sys_mount fs/namespace.c:4229 [inline]
       __se_sys_mount fs/namespace.c:4206 [inline]
       __x64_sys_mount+0x282/0x320 fs/namespace.c:4206
       x64_sys_call+0x1a7d/0x26a0 arch/x86/include/generated/asm/syscalls_64.h:166
       do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
       do_syscall_64+0x93/0xf80 arch/x86/entry/syscall_64.c:94
       entry_SYSCALL_64_after_hwframe+0x76/0x7e

-> #2 (&rrl->rr_lock){+.+.}-{4:4}:
       __mutex_lock_common kernel/locking/mutex.c:598 [inline]
       __mutex_lock+0x1a6/0x1d80 kernel/locking/mutex.c:760
       mutex_lock_nested+0x1b/0x30 kernel/locking/mutex.c:812
       rrw_enter_read_impl+0x77/0x970 fs/zfs/zfs/rrwlock.c:166
       rrw_enter_read fs/zfs/zfs/rrwlock.c:198 [inline]
       rrw_enter+0x57/0x70 fs/zfs/zfs/rrwlock.c:235
       dmu_buf_lock_parent+0x22c/0x380 fs/zfs/zfs/dbuf.c:1350
       dbuf_read+0xb88/0x1940 fs/zfs/zfs/dbuf.c:1834
       dmu_buf_hold_by_dnode+0x8e/0xe0 fs/zfs/zfs/dmu.c:232
       mzap_create_impl+0xb5/0x470 fs/zfs/zfs/zap_micro.c:842
       zap_create_impl fs/zfs/zfs/zap_micro.c:879 [inline]
       zap_create_norm_dnsize fs/zfs/zfs/zap_micro.c:968 [inline]
       zap_create_dnsize+0xd1/0x120 fs/zfs/zfs/zap_micro.c:952
       zap_create_link_dnsize+0xae/0x1b0 fs/zfs/zfs/zap.c:1104
       zap_create_link+0x3d/0x60 fs/zfs/zfs/zap.c:1095
       sa_attr_register_sync+0x8f1/0xd30 fs/zfs/zfs/sa.c:1859
       sa_replace_all_by_template_locked fs/zfs/zfs/sa.c:1892 [inline]
       sa_replace_all_by_template+0x4cd/0x640 fs/zfs/zfs/sa.c:1903
       zfs_mknode+0x162f/0x3e20 fs/zfs/os/linux/zfs/zfs_znode_os.c:905
       zfs_create_fs+0xc46/0x11f0 fs/zfs/os/linux/zfs/zfs_znode_os.c:1952
       dsl_pool_create+0x3f8/0x840 fs/zfs/zfs/dsl_pool.c:556
       spa_create+0xd2e/0x1a00 fs/zfs/zfs/spa.c:7207
       zfs_ioc_pool_create+0x4fb/0x650 fs/zfs/zfs/zfs_ioctl.c:1522
       zfsdev_ioctl_common+0x128c/0x1600 fs/zfs/zfs/zfs_ioctl.c:8239
       zfsdev_ioctl+0x65/0x120 fs/zfs/os/linux/zfs/zfs_ioctl_os.c:145
       vfs_ioctl fs/ioctl.c:51 [inline]
       __do_sys_ioctl fs/ioctl.c:597 [inline]
       __se_sys_ioctl fs/ioctl.c:583 [inline]
       __x64_sys_ioctl+0x197/0x1e0 fs/ioctl.c:583
       x64_sys_call+0x1144/0x26a0 arch/x86/include/generated/asm/syscalls_64.h:17
       do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
       do_syscall_64+0x93/0xf80 arch/x86/entry/syscall_64.c:94
       entry_SYSCALL_64_after_hwframe+0x76/0x7e

-> #1 (&hdl->sa_lock){+.+.}-{4:4}:
       __mutex_lock_common kernel/locking/mutex.c:598 [inline]
       __mutex_lock+0x1a6/0x1d80 kernel/locking/mutex.c:760
       mutex_lock_nested+0x1b/0x30 kernel/locking/mutex.c:812
       sa_lookup+0xf8/0x560 fs/zfs/zfs/sa.c:1509
       zfs_rmnode+0x245/0x14a0 fs/zfs/os/linux/zfs/zfs_dir.c:710
       zfs_zinactive+0x612/0x7e0 fs/zfs/os/linux/zfs/zfs_znode_os.c:1354
       zfs_inactive+0x258/0x1300 fs/zfs/os/linux/zfs/zfs_vnops_os.c:4113
       zpl_evict_inode+0x7d/0xe0 fs/zfs/os/linux/zfs/zpl_super.c:144
       evict+0x38e/0x8f0 fs/inode.c:810
       iput_final fs/inode.c:1914 [inline]
       iput fs/inode.c:1966 [inline]
       iput+0x55b/0x8b0 fs/inode.c:1926
       do_unlinkat+0x4cb/0x680 fs/namei.c:4744
       __do_sys_unlinkat fs/namei.c:4778 [inline]
       __se_sys_unlinkat fs/namei.c:4771 [inline]
       __x64_sys_unlinkat+0xae/0x130 fs/namei.c:4771
       x64_sys_call+0x1066/0x26a0 arch/x86/include/generated/asm/syscalls_64.h:264
       do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
       do_syscall_64+0x93/0xf80 arch/x86/entry/syscall_64.c:94
       entry_SYSCALL_64_after_hwframe+0x76/0x7e

-> #0 (&zfsvfs->z_teardown_inactive_lock){.+.+}-{4:4}:
       check_prev_add kernel/locking/lockdep.c:3165 [inline]
       check_prevs_add kernel/locking/lockdep.c:3284 [inline]
       validate_chain kernel/locking/lockdep.c:3908 [inline]
       __lock_acquire+0x14ae/0x21e0 kernel/locking/lockdep.c:5237
       lock_acquire kernel/locking/lockdep.c:5868 [inline]
       lock_acquire+0x169/0x2f0 kernel/locking/lockdep.c:5825
       down_read+0x9c/0x4a0 kernel/locking/rwsem.c:1537
       zfs_inactive+0x194/0x1300 fs/zfs/os/linux/zfs/zfs_vnops_os.c:4084
       zpl_evict_inode+0x7d/0xe0 fs/zfs/os/linux/zfs/zpl_super.c:144
       evict+0x38e/0x8f0 fs/inode.c:810
       dispose_list+0xe4/0x1d0 fs/inode.c:852
       prune_icache_sb+0xef/0x170 fs/inode.c:1000
       super_cache_scan+0x354/0x530 fs/super.c:224
       do_shrink_slab+0x37f/0xe00 mm/shrinker.c:437
       shrink_slab+0x173/0x1000 mm/shrinker.c:664
       shrink_one+0x44f/0x8a0 mm/vmscan.c:4955
       shrink_many mm/vmscan.c:5016 [inline]
       lru_gen_shrink_node mm/vmscan.c:5094 [inline]
       shrink_node+0x2428/0x3870 mm/vmscan.c:6081
       kswapd_shrink_node mm/vmscan.c:6941 [inline]
       balance_pgdat+0xb3d/0x18b0 mm/vmscan.c:7124
       kswapd+0x527/0xa20 mm/vmscan.c:7389
       kthread+0x3f0/0x850 kernel/kthread.c:463
       ret_from_fork+0x50f/0x610 arch/x86/kernel/process.c:158
       ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245

other info that might help us debug this:

Chain exists of:
  &zfsvfs->z_teardown_inactive_lock --> &rrl->rr_lock --> fs_reclaim

 Possible unsafe locking scenario:

       CPU0                    CPU1
       ----                    ----
  lock(fs_reclaim);
                               lock(&rrl->rr_lock);
                               lock(fs_reclaim);
  rlock(&zfsvfs->z_teardown_inactive_lock);

 *** DEADLOCK ***

2 locks held by kswapd0/55:
 #0: ffffffff87abcaa0 (fs_reclaim){+.+.}-{0:0}, at: balance_pgdat+0xc06/0x18b0 mm/vmscan.c:7137
 #1: ffff88800afbc0e0 (&type->s_umount_key#60){.+.+}-{4:4}, at: super_trylock_shared fs/super.c:562 [inline]
 #1: ffff88800afbc0e0 (&type->s_umount_key#60){.+.+}-{4:4}, at: super_cache_scan+0x8c/0x530 fs/super.c:197

Call Trace:
 __dump_stack lib/dump_stack.c:94 [inline]
 dump_stack_lvl+0xbe/0x130 lib/dump_stack.c:120
 dump_stack+0x15/0x20 lib/dump_stack.c:129
 print_circular_bug+0x285/0x360 kernel/locking/lockdep.c:2043
 check_noncircular+0x14e/0x170 kernel/locking/lockdep.c:2175
 check_prev_add kernel/locking/lockdep.c:3165 [inline]
 check_prevs_add kernel/locking/lockdep.c:3284 [inline]
 validate_chain kernel/locking/lockdep.c:3908 [inline]
 __lock_acquire+0x14ae/0x21e0 kernel/locking/lockdep.c:5237
 lock_acquire kernel/locking/lockdep.c:5868 [inline]
 lock_acquire+0x169/0x2f0 kernel/locking/lockdep.c:5825
 down_read+0x9c/0x4a0 kernel/locking/rwsem.c:1537
 zfs_inactive+0x194/0x1300 fs/zfs/os/linux/zfs/zfs_vnops_os.c:4084
 zpl_evict_inode+0x7d/0xe0 fs/zfs/os/linux/zfs/zpl_super.c:144
 evict+0x38e/0x8f0 fs/inode.c:810
 dispose_list+0xe4/0x1d0 fs/inode.c:852
 prune_icache_sb+0xef/0x170 fs/inode.c:1000
 super_cache_scan+0x354/0x530 fs/super.c:224
 do_shrink_slab+0x37f/0xe00 mm/shrinker.c:437
 shrink_slab+0x173/0x1000 mm/shrinker.c:664
 shrink_one+0x44f/0x8a0 mm/vmscan.c:4955
 shrink_many mm/vmscan.c:5016 [inline]
 lru_gen_shrink_node mm/vmscan.c:5094 [inline]
 shrink_node+0x2428/0x3870 mm/vmscan.c:6081
 kswapd_shrink_node mm/vmscan.c:6941 [inline]
 balance_pgdat+0xb3d/0x18b0 mm/vmscan.c:7124
 kswapd+0x527/0xa20 mm/vmscan.c:7389
 kthread+0x3f0/0x850 kernel/kthread.c:463
 ret_from_fork+0x50f/0x610 arch/x86/kernel/process.c:158
 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245

Description

This keeps the existing zfs_zinactive() and unlinked cleanup semantics.

For reclaim threads only, zfs_inactive() temporarily disables lockdep
around the acquire and release of z_teardown_inactive_lock. The real
rwsem is still taken normally, so teardown serialization is unchanged.

The patch also removes the extra zfs_unlinked_drain_try() path and the
additional z_drain_cv/TASKQID_INITIAL drain state added by the previous
version.

How Has This Been Tested?

Tested on Linux.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • Quality assurance (non-breaking change which makes the code more robust against bugs)

Checklist:

Copilot AI review requested due to automatic review settings May 7, 2026 08:23

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses a reported Linux lockdep circular dependency involving kswapd entering the zfs_inactive() / zfs_zinactive() path by avoiding synchronous zfs_rmnode() work in reclaim-thread eviction, and instead attempting to defer final deletion to the unlinked-drain taskq using a non-blocking dispatch.

Changes:

  • Detect reclaim-thread eviction in zfs_zinactive() and defer inline zfs_rmnode() for unlinked znodes.
  • Introduce zfs_unlinked_drain_try() to attempt scheduling unlinked draining using TQ_NOSLEEP.
  • Export the new helper via the Linux zfs_dir.h header.

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 3 comments.

File Description
module/os/linux/zfs/zfs_znode_os.c Defers inline zfs_rmnode() during reclaim-thread eviction and triggers best-effort unlinked draining.
module/os/linux/zfs/zfs_dir.c Adds zfs_unlinked_drain_try() which dispatches the drain task using TQ_NOSLEEP.
include/os/linux/zfs/sys/zfs_dir.h Declares zfs_unlinked_drain_try() for Linux consumers.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread module/os/linux/zfs/zfs_znode_os.c Outdated
Comment thread module/os/linux/zfs/zfs_dir.c Outdated
Comment thread module/os/linux/zfs/zfs_dir.c Outdated
@Gality369
Gality369 force-pushed the linux-zinactive-reclaim-lockdep branch from 00c773a to feaef24 Compare May 7, 2026 12:45
@behlendorf behlendorf added the Status: Code Review Needed Ready for review and testing label May 7, 2026
@behlendorf

Copy link
Copy Markdown
Contributor

I haven't looked closely yet, but this change appears to have introduced a new deadlock. See the "vm1: console logs" link under the "Test Summary" step for the failed builders. For example, for Centos Stream 9:

https://github.com/openzfs/zfs/actions/runs/25496566454/job/74817848886?pr=18505

@Gality369

Copy link
Copy Markdown
Contributor Author

Thanks for taking a look and pointing this out. The lockdep fix turned out to be more complicated than I expected, and I’m sorry for the trouble caused by the new deadlock. I’ll continue investigating this and update the PR once I have a better solution.

Copilot AI review requested due to automatic review settings May 11, 2026 07:06
@Gality369
Gality369 force-pushed the linux-zinactive-reclaim-lockdep branch from feaef24 to a11e2aa Compare May 11, 2026 07:06

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment thread module/os/linux/zfs/zfs_vnops_os.c Outdated
@Gality369 Gality369 changed the title linux: defer zfs_rmnode from reclaim in zfs_zinactive zfs: suppress reclaim lockdep in zfs_inactive May 11, 2026
Comment thread module/os/linux/zfs/zfs_vnops_os.c Outdated
@Gality369
Gality369 force-pushed the linux-zinactive-reclaim-lockdep branch from a11e2aa to b0bb945 Compare May 14, 2026 02:15
Copilot AI review requested due to automatic review settings May 14, 2026 02:44
@Gality369
Gality369 force-pushed the linux-zinactive-reclaim-lockdep branch from b0bb945 to b72eeb3 Compare May 14, 2026 02:44

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated 3 comments.

Comments suppressed due to low confidence (2)

module/os/linux/zfs/zfs_vnops_os.c:4095

  • current_is_reclaim_thread() on Linux returns true only for kswapd (current_is_kswapd()), but the same lockdep chain can be hit from any task performing direct reclaim that ends up in super_cache_scan → prune_icache_sb → evict → zpl_evict_inode → zfs_inactive. Direct-reclaim tasks hold fs_reclaim just like kswapd, so the splat (and the underlying ordering concern) will still be reported for them. If the intent really is to suppress the warning, the gating predicate should also cover direct reclaim (e.g., current->flags & PF_MEMALLOC or check current->reclaim_state).
		no_lockdep = current_is_reclaim_thread();

include/os/linux/spl/sys/rwlock.h:264

  • rw_tryenter_nolockdep and rw_downgrade_nolockdep are added but never used in this PR (verified via search). If they are not needed by the immediate fix, please drop them; otherwise add the actual call sites. Adding unused API surface invites future misuse of lockdep_off().
#define	rw_tryenter_nolockdep(rwp, rw) /* CSTYLED */		\
({									\
	spl_rw_lockdep_off();						\
	int _rc_ = spl_rw_tryenter_impl(rwp, rw);			\
	spl_rw_lockdep_on();						\
	_rc_;								\
})

#define	rw_enter(rwp, rw) /* CSTYLED */				\
({									\
	spl_rw_lockdep_off_maybe(rwp);					\
	spl_rw_enter_impl(rwp, rw);					\
	spl_rw_lockdep_on_maybe(rwp);					\
})

#define	rw_enter_nolockdep(rwp, rw) /* CSTYLED */		\
({									\
	spl_rw_lockdep_off();						\
	spl_rw_enter_impl(rwp, rw);					\
	spl_rw_lockdep_on();						\
})

#define	rw_exit(rwp) /* CSTYLED */				\
({									\
	spl_rw_lockdep_off_maybe(rwp);					\
	spl_rw_exit_impl(rwp);						\
	spl_rw_lockdep_on_maybe(rwp);					\
})

#define	rw_exit_nolockdep(rwp) /* CSTYLED */			\
({									\
	spl_rw_lockdep_off();						\
	spl_rw_exit_impl(rwp);						\
	spl_rw_lockdep_on();						\
})

#define	rw_downgrade(rwp) /* CSTYLED */			\
({									\
	spl_rw_lockdep_off_maybe(rwp);					\
	spl_rw_downgrade_impl(rwp);					\
	spl_rw_lockdep_on_maybe(rwp);					\
})

#define	rw_downgrade_nolockdep(rwp) /* CSTYLED */		\
({									\
	spl_rw_lockdep_off();						\
	spl_rw_downgrade_impl(rwp);					\
	spl_rw_lockdep_on();						\
})

Comment thread module/os/linux/zfs/zfs_vnops_os.c Outdated
Comment thread include/os/linux/spl/sys/rwlock.h Outdated
Comment thread include/os/linux/spl/sys/rwlock.h
Comment thread module/os/linux/zfs/zfs_vnops_os.c
Comment thread module/os/linux/zfs/zfs_vnops_os.c Outdated
Comment thread include/os/linux/spl/sys/rwlock.h Outdated
kswapd can enter zfs_inactive() from inode reclaim while holding
fs_reclaim. The z_teardown_inactive_lock still serializes teardown,
but the reclaim-thread acquire/release pair can produce a lockdep
cycle through zfs_zinactive() and zfs_rmnode().

Add Linux rwlock nolockdep wrappers alongside the existing rwlock
macros and use them only for the reclaim-thread
z_teardown_inactive_lock acquire/release in zfs_inactive(). Keep
the real rwsem semantics unchanged and leave CONFIG_LOCKDEP
handling in the platform rwlock layer.

Signed-off-by: ZhengYuan Huang 
@Gality369
Gality369 force-pushed the linux-zinactive-reclaim-lockdep branch from b72eeb3 to 80c05d1 Compare May 14, 2026 04:54

@behlendorf behlendorf left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Much nicer. Thanks for reworking this.

@behlendorf behlendorf added Status: Accepted Ready to integrate (reviewed, tested) and removed Status: Code Review Needed Ready for review and testing labels May 14, 2026
@satmandu satmandu mentioned this pull request Jun 9, 2026
14 tasks
tonyhutter pushed a commit that referenced this pull request Jun 11, 2026
kswapd can enter zfs_inactive() from inode reclaim while holding
fs_reclaim. The z_teardown_inactive_lock still serializes teardown,
but the reclaim-thread acquire/release pair can produce a lockdep
cycle through zfs_zinactive() and zfs_rmnode().

Add Linux rwlock nolockdep wrappers alongside the existing rwlock
macros and use them only for the reclaim-thread
z_teardown_inactive_lock acquire/release in zfs_inactive(). Keep
the real rwsem semantics unchanged and leave CONFIG_LOCKDEP
handling in the platform rwlock layer.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes #18505
tonyhutter pushed a commit that referenced this pull request Jun 11, 2026
kswapd can enter zfs_inactive() from inode reclaim while holding
fs_reclaim. The z_teardown_inactive_lock still serializes teardown,
but the reclaim-thread acquire/release pair can produce a lockdep
cycle through zfs_zinactive() and zfs_rmnode().

Add Linux rwlock nolockdep wrappers alongside the existing rwlock
macros and use them only for the reclaim-thread
z_teardown_inactive_lock acquire/release in zfs_inactive(). Keep
the real rwsem semantics unchanged and leave CONFIG_LOCKDEP
handling in the platform rwlock layer.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes #18505
bugclerk pushed a commit to truenas/zfs that referenced this pull request Jun 15, 2026
kswapd can enter zfs_inactive() from inode reclaim while holding
fs_reclaim. The z_teardown_inactive_lock still serializes teardown,
but the reclaim-thread acquire/release pair can produce a lockdep
cycle through zfs_zinactive() and zfs_rmnode().

Add Linux rwlock nolockdep wrappers alongside the existing rwlock
macros and use them only for the reclaim-thread
z_teardown_inactive_lock acquire/release in zfs_inactive(). Keep
the real rwsem semantics unchanged and leave CONFIG_LOCKDEP
handling in the platform rwlock layer.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes openzfs#18505
(cherry picked from commit f8d6096)
creatorcary pushed a commit to truenas/zfs that referenced this pull request Jun 15, 2026
…xhamza) (#406)

* CI: FreeBSD 15.1 STABLE

Update the freebsd15-1s builder to the released STABLE image.

Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18524

* CI: Fix 99.99 META version

We have an option in zfs-qemu-packages to test against a specific kernel
version.  However, qemu-3-deps.sh was incorrectly hard coded to look
at $2 for a kernel version argument (which could come in $2 or $3
depending on if --poweroff was also passed).  This caused the CI
to incorrectly edit META with a max supported kernel version of 99.99
when we didn't want that.

Fix this by looking at all the arguments for something that looks
like a kernel version and set that as the kernel max in META.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes openzfs#18526
Closes openzfs#18531

* CI: Remove deprecated Fedora 42

Fedora 42 was deprecated on May 13 2026.  Remove it from CI tests.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18545

* CI: Allow testing with a newer GCC on ARM builder

Add a text box to specify a custom GCC version (like '16') when
running the zfs-arm builder.  This allows you to test with a newer
GCC than the Ubuntu default.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes openzfs#18540

* CI: remove FreeBSD 13.5 (EOL April 30, 2026)

FreeBSD 13.5 and stable/13 reached End-of-Life on April 30, 2026 and no
longer receive security support, so they fall outside README.md's stated
support policy.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Christos Longros 
Closes openzfs#18553

* CI: Add Ubuntu 26.04 builder

The Ubuntu 26.04 LTS, named "Resolute Raccoon, was released on
April 23, 2026.  Add to the supported releases in README.md and
add a CI builder for it.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18547

* CI: Fix qemu-guest-agent systemd enable

The qemu-guest-agent.service for Debian and Ubuntu does
not contain an install section which prevents it from
being enabled.  Add a drop-in override file so it can
be enabled and the service started on boot.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18547

* ZTS: zfs_unshare_006_pos.ksh enable usershares

Ensure samba usershares are enabled in the CI test environment for
the zfs_unshare_006_pos test case.  By default they are disabled
in the Ubuntu 26.04 LTS and must be enabled.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18547

* CI: Build custom branch from zfs-qemu-packages

The zfs-qemu-packages workflow allows us to easily build RPMs for the
current branch.  However, there can be cases where we want to use the
current CI environment to build older releases.  This can happen when
the VM or runner environment changes, and the older CI doesn't have
the updates needed to run with it anymore.

This commit adds in a text box to specify a specific branch/tag to build
using the current CI environment.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes openzfs#18569

* CI: enable FreeBSD 15.0-RELEASE in matrix

Add freebsd15-0r to the FreeBSD presets

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Christos Longros 
Closes openzfs#18561

* CI: skip full CI runs on push events

Full CI runs for proposed changes always occur in the PR where the
review is done and patch approved.  Once merged the full CI is run
again using the merged commit.  This is somewhat overkill.  In the
interest of reducing the CI load only run the zloop and checkstyle
workflows which are enough to verify the build on the master branch.
Push events to forks will continue to trigger a full CI run.

Reviewed-by: George Melikov 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18571

* CI: run full CI when a workflow YAML changes

FULL_RUN_REGEX in generate-ci-type.py covered .github/workflows/scripts/
but not the workflow YAML files, so a PR that only edited zfs-qemu.yml
got "quick" CI and never tested its own matrix change. Add the YAML
files to the list.

Reviewed-by: George Melikov 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes openzfs#18577

* .github: update workflows README

Describe the current zfs-qemu pipeline, ci_type selection, supported
guests, and the code-checking and other auxiliary workflows.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes openzfs#18590

* CI: Update checkstyle checkout action to v6

The checkstyle workflow was the only one still pinned to
actions/checkout@v4; the other workflows already use v6.
Bump it to match.

Reviewed-by: Tony Hutter 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes openzfs#18600

* CI: Lustre 6.16 kernel compatibility fix (openzfs#18602)

Almalinux 9,10 kernels now include a backport of Linux commit
v6.15-13744-g41cb08555c41 which renames the from_timer() function
to timer_container_of().  Apply the upstream Lustre compatibility
patch to our builds.  This patch should be included in the next
Lustre release and can be dropped then.

ZFS-CI-Type: quick

Signed-off-by: Brian Behlendorf 
Reviewed-by: Tony Hutter 

* CI: skip smatch, zloop, and zfs-arm for documentation-only changes

Follow-up to openzfs#18518, which skipped the qemu matrix on doc-only PRs.
zloop, zfs-arm, and smatch are irrelevant to doc-only changes.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Tony Hutter 
Signed-off-by: Christos Longros 
Closes openzfs#18601

* build: add ZFS_DEBUG Kconfig for copy-builtin

... so we can toggle ZFS debug assertions from the
Linux kernel build without having to regenerate the
ZFS patch.

Update the qemu test script to also set this kernel
config.

Reviewed-by: Tony Hutter 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Timothy Day 
Co-authored-by: Timothy Day 
Closes openzfs#18595

* CI: apt-get update before purging host packages

The package removal ran against a stale package index and failed to
fetch a package that had been removed from the repository. Refresh
the index first.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes openzfs#18607
Closes openzfs#18609

* CI: add concurrency support to zfs-arm

The zfs-arm workflow was the only build/test workflow without a
concurrency block, so superseded runs were not cancelled.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes openzfs#18608

* nvpair: Check for un-terminated strings in packed nvlist

Add additional checks to verify a packed string or string array nvpair
is terminated.  Or more specifically, verify doing a strlen() on the
prospective string does not overrun the packed nvlist buffer.

Also add additional checks in the libzfs_input_checks test case to
verify un-terminated strings, and add in a nvlist ioctl payload
fuzz test for good measure.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes openzfs#18604

* Fix the integer type in zfs_ioc_userspace_many()

Fix the mismatched type in zfs_ioc_userspace_many() and limit the
number of entries returned to 1000.  When a size larger than this
is requested the response is truncated, zfs_userspace() already
correctly handles short responses.  Historically, zfs_userspace()
has requested 100 entries at a time, this cap allows for 10x larger
batch sizes if needed in the future.

Reported-by: Yuxiang Yang, Yizhou Zhao, Ao Wang, Xuewei Feng, Qi Li,
Reported-by: and Ke Xu from Tsinghua University using GLM-5.1 from Z.ai
Reviewed-by: Alexander Motin 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18615

* sharenfs: Check for invalid characters

Check for invalid characters in sharenfs/sharesmb dataset props.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes openzfs#18613

* Fix uninitialized variable warning in vdev_prop_get()

Update vdev_prop_get_objid() to set objid on error as the comment
in vdev_prop_get() describes.

    "objid is set to 0 when absent and the few cases that call
    zap_lookup directly guard against this below."

This resolves the following possible uninitialized variable warning.

    module/zfs/vdev.c: In function ‘vdev_prop_get’:
    module/zfs/vdev.c:6913:12: error: ‘objid’ may be used uninitialized
    in this function [-Werror=maybe-uninitialized]

Reviewed-by: Alexander Motin 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18616

* Extend dataset zfs_ioc_set_prop() secpolicy

When zc->zc_cookie is set this indicates to zfs_ioc_set_prop() that
these are received properties and ZPROP_HAS_RECVD will be set on the
dataset.  This is only done as part of a `zfs receive` so additionally
apply the zfs_secpolicy_recv() policy.  Individual property checks
continue to be handled by zfs_check_settable().

Reviewed-by: Tony Hutter 
Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18617

* pam: use open fd instead of path

Instead of performing multiple operations on the path name in
zfs_key_config_modify_session_counter() open the file once and
perform the fchown, fchmod, and openat on the open file handle.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18618

* Remove /etc/sudoers.d/zfs

The smartctl exception in /etc/sudoers.d/zfs doesn't cover devices
like NVMe or symlinked devices.  Just get rid of it rather than
keep maintaining it.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Tony Hutter 
Closes openzfs#18626

* CI: Re-enable CodeQL workflows on push

This workflow was disabled 'on push' recently in commit 1916c2c
to reduce redundant CI runs.  However, this check is fairly quick
and we want it run regularly against the branches.  Enable it.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18627

* CI: Update CodeQL actions to v4

CodeQL Action v3 has been deprecated and will be retired
December 2026.  Update codeql.yml to use CodeQL Action v4
and update the runner to ubuntu-24.04.

Signed-off-by: Brian Behlendorf 
Closes openzfs#18629

* CI: Increase default RCU stall timeout on Linux

When CONFIG_RCU_CPU_STALL_TIMEOUT is configured an RCU stall which
exceeds the default timeout will trigger an NMI and panic the VM.
Given the heavily virtualized nature of the CI environment we want
to make sure to only trigger this due to a real deadlock and not
due to over-subscription of the systems resources.  This timeout
normally defaults to 20-30 seconds and this change increases it
to 120 seconds.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18624

* CI: Add alternative URLs for CentOS stream

Fallback to trying the "CentOS Strean Composes" repo for the qcow2
images if the regular URLs fail.  The Composes repo contains the daily
autobuilt Stream images.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes openzfs#18628

* Add additional verification of size fields and strings (openzfs#18623)

- Check for size fields that convert to smaller integers.
- Explicitly terminate bootenv string.
- Initialize variables that could be returned in an error case.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Reviewed-by: Alexander Motin 
Signed-off-by: Tony Hutter 
Closes openzfs#18623

* Fix uninitialized variable warning in zil_parse()

This resolves the following possible uninitialized variable warning
when building with --enable-code-coverage and gcc 8.5.0.

    module/zfs/zil.c: In function ‘zil_parse’:
    module/zfs/zil.c:549:47: warning: ‘end’ may be used uninitialized
    in this function [-Wmaybe-uninitialized]

Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18633

* ZTS: relax zpool_import_parallel_pos.ksh timing

Occasionally in the CI this test will fail because the parallel import
took longer than half of the serial time (but still less than the full
serial time).  Increase the cutoff to 3/4 of the serial time to preserve
the intent yet try and avoid these false positive failures.

Reviewed-by: Chris Longros 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18634

* linux: verify stale znodes in legacy fallocate

The mode=0 and FALLOC_FL_KEEP_SIZE preallocation path can reach
zfs_freesp() directly and call zfs_statvfs() before going through the
normal zpl_enter_verify_zp() boundary.

When zfs_rezget() tears down a failed SA reload, a stale inode may
remain alive in the VFS with z_sa_hdl cleared. The unchecked
fallocate path can then reach sa_lookup(zp->z_sa_hdl, ...) through
zfs_statvfs() or zfs_freesp() and crash on a NULL SA handle.

Use zfs_enter_verify_zp() in zfs_statvfs() so stale znodes are
rejected under the teardown lock for both fallocate and statfs.
Also wrap the direct zfs_freesp() call in
zpl_enter_verify_zp()/zfs_exit() so this path follows the same
validation rules as the other Linux ZPL file operations.

Fixes: f734301
("linux: add basic fallocate(mode=0/2) compatibility")

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Co-authored-by: gality369 
Closes openzfs#18458

* Linux: annotate nested xattr setattr znode locks

zfs_setattr() updates both the target znode and its hidden xattr
directory when ownership, mode, or project ID changes. The xattr
directory uses the same z_acl_lock and z_lock classes as the
parent znode, so lockdep reports recursive locking when the
second znode's mutexes are acquired.

This is a lockdep false positive rather than a real deadlock.
attrzp is the target file's hidden xattr directory, and the code
does not acquire these znode mutexes in the reverse order.
Acquire the attrzp mutexes with mutex_enter_nested() so lockdep
treats them as nested.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Co-authored-by: gality369 
Closes openzfs#18506

* Linux: avoid znode list lock inversion during resume

Lockdep reports a circular locking dependency during mounted filesystem
rollback.  zfs_resume_fs() walks z_all_znodes under z_znodes_lock and
calls zfs_rezget(), which takes the per-object znode hold lock via
zfs_znode_hold_enter().

The normal zget path takes these locks in the opposite order.
zfs_zget() takes the per-object hold lock before zfs_znode_alloc()
inserts the znode on z_all_znodes under z_znodes_lock.  Resume can
therefore establish z_znodes_lock -> zh_lock while normal lookup
creates zh_lock -> z_znodes_lock.

Pin the current and next znodes with igrab() while holding the list
lock, then drop the list lock before reloading the znode.  Existing
stale inode handling is preserved, and both the suspended reference
and temporary walk reference are released asynchronously.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes openzfs#18517

* linux/zpl_super: handle 'source' option directly

vfs_parse_fs_param_source() didn't appear until 5.14, and was not
backported to kernel.org LTS kernels. It's simple enough that it's
easier to just handle it ourselves rather than use a configure check.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes openzfs#18529

* linux: suppress reclaim lockdep in zfs_inactive via rwlock wrappers

kswapd can enter zfs_inactive() from inode reclaim while holding
fs_reclaim. The z_teardown_inactive_lock still serializes teardown,
but the reclaim-thread acquire/release pair can produce a lockdep
cycle through zfs_zinactive() and zfs_rmnode().

Add Linux rwlock nolockdep wrappers alongside the existing rwlock
macros and use them only for the reclaim-thread
z_teardown_inactive_lock acquire/release in zfs_inactive(). Keep
the real rwsem semantics unchanged and leave CONFIG_LOCKDEP
handling in the platform rwlock layer.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes openzfs#18505

* config: show progress output for kernel API checks

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes openzfs#18554

* linux/super: properly apply ro/rw mount option to superblock

f5a9e3a changed how SB_RDONLY was applied to the new mount in a way
that was too simplistic - it only sets readonly on the filesystem if the
mount was 'ro', but it never clears it if the mount was 'rw'. This
causes the 'rw' option to effectively be ignored, and so the readonly=
property wins out.

This fixes it by doing it the right way: checking the flags mask to see
if it was actually provided as an option at all, and then setting or
clearing it as appropriate.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes openzfs#18557
Closes openzfs#18563

* Linux 5.6 compat: fix fs_parse API mismatch

Added m4 macro to check fs_parse API signature and wrappers.  Before 
5.6, fs_parse() took a struct fs_parameter_description which wraps
the parameter specs with name and enum pointers. From 5.6, the 
description struct was removed and fs_parse() accepts the 
fs_parameter_spec directly.

Reviewed-by: Rob Norris 
Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes openzfs#18585

* Fix aarch64 build failure by removing earlyclobber (openzfs#18532)

The UVR macros used "+&w" (read-write + earlyclobber) as the
constraint for NEON register operands that are declared as explicit
hard-register variables via:

register unsigned char wN asm("vN") __attribute__((vector_size(16)));

The + modifier implicitly makes the operand also an input (reading the
register before the asm runs). The & (earlyclobber) modifier says "this
output may be written before all inputs are consumed." Having an
earlyclobber output on the same hard-register that is simultaneously
an input is a contradiction — GCC 16 now strictly diagnoses this.

The fix removes the & from "+&w", yielding "+w". The earlyclobber
was both incorrect (contradicts the implicit input) and unnecessary
(the physical registers are already hard-bound, so the compiler has no
freedom to assign conflicting registers anyway).



Issue openzfs#18525

Signed-off-by: Brian Behlendorf 
Co-authored-by: Claude Sonnet 4.6 
Reviewed-by: Tony Hutter 

* Fix "panic: cache_vop_rename: lingering negative entry"

A FreeBSD ZFS filesystem with properties "utf8only=on" and
"normalization=formD" consistently produces this panic when
building the lang/perl-5.42.0 port.

A ZFS file system with "utf8only=off" and "normalization=none"
works fine.

The cause of the panic seems to be incorrectly using the FreeBSD
namecache when normalisation is present. This commit adds a
predicate to prevent that.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Jan Martin Mikkelsen 
Closes openzfs#18430

* key lookup failure should always return EACCES

spa_do_crypt_abd() already maps a missing key to EACCES. However
spa_do_crypt_mac_abd(), spa_do_crypt_objset_mac_abd(), and
spa_crypt_get_salt() still return the raw
spa_keystore_lookup_key() error (ENOENT). This is inconsistent
As we want to treat all “no key” failures as a permission
failure. Standardize on EACCES for the unloaded-key case.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alek Pinchuk 
Closes openzfs#18448

* Fix off-by-one in PREVIOUSLY_REDACTED handler that drops last block

In send_reader_thread(), the PREVIOUSLY_REDACTED handler computed
file_max as MIN(dn->dn_maxblkid, range->end_blkid).  dn_maxblkid is
an inclusive maximum block ID while range->end_blkid is exclusive (one
past the last block).  The resulting file_max was then used as an
exclusive loop bound, causing the last block of any file (at index
dn_maxblkid) to be silently skipped when a PREVIOUSLY_REDACTED range
covered the end of the file.

The block was never written to the send stream so the receiver kept
zeros there.  ZFS reported no error because the stream itself was
valid; the data was simply absent.

Fix: use dn_maxblkid + 1 so file_max is consistently exclusive.

Add a regression test (redacted_max_blkid.ksh) that modifies only the
last block of a file in one clone, creates a redaction bookmark from
it, then sends an unmodified clone incrementally from that bookmark.
The PREVIOUSLY_REDACTED path must fill in the last block; the test
verifies it is not zeros and matches the original.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Paul Dagnelie 
Reviewed-by: Reviewed-by: Tony Hutter 
Signed-off-by: Manoj Joseph 
Closes openzfs#18477

* Avoid flushing unrelated NFS exports on snapshot unmount

zfsctl_snapshot_unmount() called exportfs_flush() before every umount
attempt to drop NFS export cache references that pin the snapshot
mountpoint.  The flush has global effect on the host's NFS exports and
clients, so paying it on every snapshot unmount (including auto-expire
rounds for snapshots that were never NFS-accessed) impacts unrelated
snapshots and clients.

ZFS cannot invalidate individual export cache entries because the
relevant sunrpc cache APIs are exported GPL-only.  Defer the global
flush so it runs only when the umount has actually failed, then retry
once.  Snapshots that are not NFS-pinned succeed on the first attempt
and never trigger the flush.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Youzhong Yang 
Signed-off-by: Ameer Hamza 
Closes openzfs#18476

* zfs: annotate nested dd_lock in reservation sync accounting

When reservation sync updates a child's reserved space, it rolls the
delta into ancestor space accounting while still holding the child's
dd_lock.  That locking order is intentional, but Linux lockdep sees
the ancestor acquisition as recursive because it lacks a nested lock
subclass annotation.

Teach the reservation-sync space-accounting path to acquire ancestor
dd_lock instances with a nested subclass.  Keep the existing public
interfaces and accounting behavior unchanged by routing only the
ancestor rollup through local helpers.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Signed-off-by: gality369 
Closes openzfs#18497

* sa: fix sa_add_projid lock ordering

sa_add_projid() currently acquires hdl->sa_lock before zp->z_lock.
Several same-znode update paths take zp->z_lock and then call
sa_update() or sa_bulk_update() on the same SA handle.

On Linux, FS_IOC_FSSETXATTR reaches zfs_setattr() through
zpl_ioctl_setxattr() without outer inode serialization. This makes
the reversed lock order a real ABBA deadlock rather than a lockdep
false positive when projid is added to an old-format inode while
another thread updates the same znode.

Acquire zp->z_lock before hdl->sa_lock in sa_add_projid() to match
the existing znode update ordering.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Co-authored-by: gality369 
Closes openzfs#18503

* zdb: detect BRT and DDT leaks during block traversal

During -b traversal, track BRT and DDT reference counts and report
blocks claimed more times than their reference tables account for
if it causes claim errors, instead of just asserting it.  Also
report entries with references not fully consumed by the traversal.

Add zdb leaks checks to cloning and dedup tests. This should make
sure the pools are in a sane state after completing the functional
tests.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes openzfs#18494

* zarcstat: detect attached L2ARC device with no data

zarcstat and zarcsummary detected L2ARC presence using the l2_size
kstat, which is data held in L2ARC, not whether a cache device is
attached. When a cache device was attached but empty (freshly added,
or fully evicted):

  - zarcstat rejected "-f l2*" with "Incompatible field specified!"
  - zarcsummary printed "L2ARC not detected, skipping section",
    hiding cumulative I/O history and health counters

Expose the existing l2arc_ndev counter as a new kstat l2_dev_count.
It is maintained by l2arc_add_vdev() and l2arc_remove_vdev(), so it
tracks attachment in real time. Use it in both tools, falling back to
l2_size for compatibility with older kernel modules.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Ameer Hamza 
Closes openzfs#18499

* Fix double free for blocks cloned after DDT prune

Before this change, for blocks marked with D flag but absent in DDT
(pruned from it), zio_ddt_free() fell back to ZIO_STAGE_DVA_FREE
without trying ZIO_STAGE_BRT_FREE first.  Same time such blocks
might be present in BRT, and not handling that would result in
double/multiple free.

This change makes ZIO_DDT_FREE_PIPELINE include ZIO_FREE_PIPELINE,
just adding required ZIO_STAGE_ISSUE_ASYNC and ZIO_STAGE_DDT_FREE,
and moves DDT stages before BRT.  This way, if the block is found
in DDT by zio_ddt_free(), the pipeline is short-circuited to
ZIO_INTERLOCK_PIPELINE, similar to what zio_brt_free() does.  If
not, then BRT is checked, and if also no match, the block is freed.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Rob Norris 
Signed-off-by: Alexander Motin 
Closes openzfs#18520

* arc: export additional required symbols

External consumers of arc_read() need to be able to destroy the
returned arc_buf_t.  Add the arc_buf_destroy() interface as an
exported symbol.

Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes openzfs#18533

* zap_impl: use flex array field for mzap_phys_t.mz_chunks

mz_phys_t is always a full-block allocation, with mz_chunks[] as an
array over the rest of the block past the header.

Recent Linux compiled with CONFIG_UBSAN will complain about this:

    UBSAN: array-index-out-of-bounds in module/zfs/zap.c:1236:28
    index 2 is out of range for type 'mzap_ent_phys_t [1]'

The fix is straightforward; simply convert this field to a flex member.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes openzfs#18550

* spl_kvmalloc: remove __GFP_COMP before calling vmalloc()

In cb18330 we stopped using it for KM_VMEM allocations, since its not
a valid flag for vmalloc(). However, there's a fallback path for
non-KM_VMEM allocations to use vmalloc(), and we need to remove
__GFP_COMP there too to avoid a warning.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes openzfs#18558

* FreeBSD: Make it possible to build openzfs.ko with sanitizers

Add make options which let one respectively compile the kernel modules
with the address sanitizer, memory sanitizer, and undefined behaviour
sanitizer enabled.  This makes it much easier to run the ZTS with those
sanitizers enabled.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Signed-off-by: Mark Johnston 
Closes openzfs#18596

* enforce exact decompressed length for lz4, gzip, and zstd

Decompressors must expand a ZFS block to exactly the expected number
of bytes. Treat decompression to an unexpected length as failure, so
truncated or short output is not accepted as valid decompression. This
makes our handling of decompress return values consistent with the
decompression functions' APIs.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Alek Pinchuk 
Closes openzfs#18599

* dsl_scan: close errorscrub cursor on pause

If the cursor were ever to actively hold resources, not finalising it
would mean leaking those resources whenever the scrub is paused.

The cursor is already reinitialized from the stored serialized form
if/when it is resumed, so there's nothing we need from the old one, just
to release it.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes openzfs#18603

* When reading a vdev label skip libzfs_core_init()

There's no need to call libzfs_core_init() when `zdb -l` is used to
read a vdev label.

Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes openzfs#18606

* Simplify dnode_level_is_l2cacheable()

We should not dereference through dn_handle->dnh_dnode once we
already have a dnode pointer.  The result will be the same.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes openzfs#18212

* Remove parent ZIO from dbuf_prefetch()

I am not sure why it was added there 10 years ago, but it seems not
needed now.  According to my tests removing it improves sequential
read performance with recordsize=4K by 5-10% by reducing the CPU
overhead in prefetcher.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Rob Norris 
Reviewed-by: Ameer Hamza 
Reviewed-by: Akash B 
Signed-off-by: Alexander Motin 
Closes openzfs#18214

* Fix log vdev removal issues

When we clear the log, we should clear all the fields, not only
zh_log.  Otherwise remaining ZIL_REPLAY_NEEDED will prevent the
vdev removal.  Handle it also from the other side, when zh_log
is already cleared, while zh_flags is not.

spa_vdev_remove_log() asserts that allocated space on removed log
device is zero.  While it should be so in perfect world, it might
be not if space leaked at any point.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes openzfs#18277

* ZVOL: Add encryption key check for block cloning

Somehow during block cloning porting from file systems was missed
the check for identical encryption keys.  As result, blocks cloned
between unrelated ZVOLs produced authentication errors on later
reads.  Having same or different encryption root does not matter.

This patch copies dmu_objset_crypto_key_equal() call from FS side.

Reviewed-by: Ameer Hamza 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes openzfs#18315

* abd: Fix stats asymmetry in case of Direct I/O

abd_alloc_from_pages() does not call abd_update_scatter_stats(),
since memory is not really allocated there.  But abd_free_scatter()
called by abd_free() does.  It causes negative overflow of some
ABD and possibly ARC counters.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Rob Norris 
Signed-off-by: Alexander Motin 
Closes openzfs#18390

* Tag zfs-2.4.3

META file and changelog updated.

Signed-off-by: Tony Hutter 

---------

Signed-off-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Signed-off-by: Christos Longros 
Signed-off-by: Timothy Day 
Signed-off-by: ZhengYuan Huang 
Signed-off-by: Rob Norris 
Signed-off-by: tiehexue 
Signed-off-by: Jan Martin Mikkelsen 
Signed-off-by: Alek Pinchuk 
Signed-off-by: Manoj Joseph 
Signed-off-by: Ameer Hamza 
Signed-off-by: gality369 
Signed-off-by: Alexander Motin 
Signed-off-by: Mark Johnston 
Signed-off-by: Alek Pinchuk 
Co-authored-by: Brian Behlendorf 
Co-authored-by: Tony Hutter 
Co-authored-by: Christos Longros <98426896+chrislongros@users.noreply.github.com>
Co-authored-by: Timothy Day 
Co-authored-by: Timothy Day 
Co-authored-by: Gality <68463495+Gality369@users.noreply.github.com>
Co-authored-by: gality369 
Co-authored-by: Rob Norris 
Co-authored-by: ZhengYuan Huang 
Co-authored-by: tiehexue 
Co-authored-by: Claude Sonnet 4.6 
Co-authored-by: Jan Martin Mikkelsen 
Co-authored-by: Alek P 
Co-authored-by: Manoj Joseph 
Co-authored-by: Ameer Hamza 
Co-authored-by: Alexander Motin 
Co-authored-by: Mark Johnston 
lundman pushed a commit to openzfsonosx/openzfs-fork that referenced this pull request Jul 30, 2026
kswapd can enter zfs_inactive() from inode reclaim while holding
fs_reclaim. The z_teardown_inactive_lock still serializes teardown,
but the reclaim-thread acquire/release pair can produce a lockdep
cycle through zfs_zinactive() and zfs_rmnode().

Add Linux rwlock nolockdep wrappers alongside the existing rwlock
macros and use them only for the reclaim-thread
z_teardown_inactive_lock acquire/release in zfs_inactive(). Keep
the real rwsem semantics unchanged and leave CONFIG_LOCKDEP
handling in the platform rwlock layer.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes openzfs#18505
creatorcary pushed a commit to truenas/zfs that referenced this pull request Aug 26, 2026
…431)

* CI: FreeBSD 15.1 STABLE

Update the freebsd15-1s builder to the released STABLE image.

Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18524

* CI: Fix 99.99 META version

We have an option in zfs-qemu-packages to test against a specific kernel
version.  However, qemu-3-deps.sh was incorrectly hard coded to look
at $2 for a kernel version argument (which could come in $2 or $3
depending on if --poweroff was also passed).  This caused the CI
to incorrectly edit META with a max supported kernel version of 99.99
when we didn't want that.

Fix this by looking at all the arguments for something that looks
like a kernel version and set that as the kernel max in META.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18526
Closes #18531

* CI: Remove deprecated Fedora 42

Fedora 42 was deprecated on May 13 2026.  Remove it from CI tests.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18545

* CI: Allow testing with a newer GCC on ARM builder

Add a text box to specify a custom GCC version (like '16') when
running the zfs-arm builder.  This allows you to test with a newer
GCC than the Ubuntu default.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18540

* CI: remove FreeBSD 13.5 (EOL April 30, 2026)

FreeBSD 13.5 and stable/13 reached End-of-Life on April 30, 2026 and no
longer receive security support, so they fall outside README.md's stated
support policy.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Christos Longros 
Closes #18553

* CI: Add Ubuntu 26.04 builder

The Ubuntu 26.04 LTS, named "Resolute Raccoon, was released on
April 23, 2026.  Add to the supported releases in README.md and
add a CI builder for it.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* CI: Fix qemu-guest-agent systemd enable

The qemu-guest-agent.service for Debian and Ubuntu does
not contain an install section which prevents it from
being enabled.  Add a drop-in override file so it can
be enabled and the service started on boot.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* ZTS: zfs_unshare_006_pos.ksh enable usershares

Ensure samba usershares are enabled in the CI test environment for
the zfs_unshare_006_pos test case.  By default they are disabled
in the Ubuntu 26.04 LTS and must be enabled.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* CI: Build custom branch from zfs-qemu-packages

The zfs-qemu-packages workflow allows us to easily build RPMs for the
current branch.  However, there can be cases where we want to use the
current CI environment to build older releases.  This can happen when
the VM or runner environment changes, and the older CI doesn't have
the updates needed to run with it anymore.

This commit adds in a text box to specify a specific branch/tag to build
using the current CI environment.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18569

* CI: enable FreeBSD 15.0-RELEASE in matrix

Add freebsd15-0r to the FreeBSD presets

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Christos Longros 
Closes #18561

* CI: skip full CI runs on push events

Full CI runs for proposed changes always occur in the PR where the
review is done and patch approved.  Once merged the full CI is run
again using the merged commit.  This is somewhat overkill.  In the
interest of reducing the CI load only run the zloop and checkstyle
workflows which are enough to verify the build on the master branch.
Push events to forks will continue to trigger a full CI run.

Reviewed-by: George Melikov 
Signed-off-by: Brian Behlendorf 
Closes #18571

* CI: run full CI when a workflow YAML changes

FULL_RUN_REGEX in generate-ci-type.py covered .github/workflows/scripts/
but not the workflow YAML files, so a PR that only edited zfs-qemu.yml
got "quick" CI and never tested its own matrix change. Add the YAML
files to the list.

Reviewed-by: George Melikov 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18577

* .github: update workflows README

Describe the current zfs-qemu pipeline, ci_type selection, supported
guests, and the code-checking and other auxiliary workflows.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18590

* CI: Update checkstyle checkout action to v6

The checkstyle workflow was the only one still pinned to
actions/checkout@v4; the other workflows already use v6.
Bump it to match.

Reviewed-by: Tony Hutter 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18600

* CI: Lustre 6.16 kernel compatibility fix (#18602)

Almalinux 9,10 kernels now include a backport of Linux commit
v6.15-13744-g41cb08555c41 which renames the from_timer() function
to timer_container_of().  Apply the upstream Lustre compatibility
patch to our builds.  This patch should be included in the next
Lustre release and can be dropped then.

ZFS-CI-Type: quick

Signed-off-by: Brian Behlendorf 
Reviewed-by: Tony Hutter 

* CI: skip smatch, zloop, and zfs-arm for documentation-only changes

Follow-up to #18518, which skipped the qemu matrix on doc-only PRs.
zloop, zfs-arm, and smatch are irrelevant to doc-only changes.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Tony Hutter 
Signed-off-by: Christos Longros 
Closes #18601

* build: add ZFS_DEBUG Kconfig for copy-builtin

... so we can toggle ZFS debug assertions from the
Linux kernel build without having to regenerate the
ZFS patch.

Update the qemu test script to also set this kernel
config.

Reviewed-by: Tony Hutter 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Timothy Day 
Co-authored-by: Timothy Day 
Closes #18595

* CI: apt-get update before purging host packages

The package removal ran against a stale package index and failed to
fetch a package that had been removed from the repository. Refresh
the index first.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18607
Closes #18609

* CI: add concurrency support to zfs-arm

The zfs-arm workflow was the only build/test workflow without a
concurrency block, so superseded runs were not cancelled.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Christos Longros 
Closes #18608

* nvpair: Check for un-terminated strings in packed nvlist

Add additional checks to verify a packed string or string array nvpair
is terminated.  Or more specifically, verify doing a strlen() on the
prospective string does not overrun the packed nvlist buffer.

Also add additional checks in the libzfs_input_checks test case to
verify un-terminated strings, and add in a nvlist ioctl payload
fuzz test for good measure.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18604

* Fix the integer type in zfs_ioc_userspace_many()

Fix the mismatched type in zfs_ioc_userspace_many() and limit the
number of entries returned to 1000.  When a size larger than this
is requested the response is truncated, zfs_userspace() already
correctly handles short responses.  Historically, zfs_userspace()
has requested 100 entries at a time, this cap allows for 10x larger
batch sizes if needed in the future.

Reported-by: Yuxiang Yang, Yizhou Zhao, Ao Wang, Xuewei Feng, Qi Li,
Reported-by: and Ke Xu from Tsinghua University using GLM-5.1 from Z.ai
Reviewed-by: Alexander Motin 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18615

* sharenfs: Check for invalid characters

Check for invalid characters in sharenfs/sharesmb dataset props.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18613

* Fix uninitialized variable warning in vdev_prop_get()

Update vdev_prop_get_objid() to set objid on error as the comment
in vdev_prop_get() describes.

    "objid is set to 0 when absent and the few cases that call
    zap_lookup directly guard against this below."

This resolves the following possible uninitialized variable warning.

    module/zfs/vdev.c: In function ‘vdev_prop_get’:
    module/zfs/vdev.c:6913:12: error: ‘objid’ may be used uninitialized
    in this function [-Werror=maybe-uninitialized]

Reviewed-by: Alexander Motin 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18616

* Extend dataset zfs_ioc_set_prop() secpolicy

When zc->zc_cookie is set this indicates to zfs_ioc_set_prop() that
these are received properties and ZPROP_HAS_RECVD will be set on the
dataset.  This is only done as part of a `zfs receive` so additionally
apply the zfs_secpolicy_recv() policy.  Individual property checks
continue to be handled by zfs_check_settable().

Reviewed-by: Tony Hutter 
Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18617

* pam: use open fd instead of path

Instead of performing multiple operations on the path name in
zfs_key_config_modify_session_counter() open the file once and
perform the fchown, fchmod, and openat on the open file handle.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18618

* Remove /etc/sudoers.d/zfs

The smartctl exception in /etc/sudoers.d/zfs doesn't cover devices
like NVMe or symlinked devices.  Just get rid of it rather than
keep maintaining it.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Tony Hutter 
Closes #18626

* CI: Re-enable CodeQL workflows on push

This workflow was disabled 'on push' recently in commit 1916c2c5
to reduce redundant CI runs.  However, this check is fairly quick
and we want it run regularly against the branches.  Enable it.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18627

* CI: Update CodeQL actions to v4

CodeQL Action v3 has been deprecated and will be retired
December 2026.  Update codeql.yml to use CodeQL Action v4
and update the runner to ubuntu-24.04.

Signed-off-by: Brian Behlendorf 
Closes #18629

* CI: Increase default RCU stall timeout on Linux

When CONFIG_RCU_CPU_STALL_TIMEOUT is configured an RCU stall which
exceeds the default timeout will trigger an NMI and panic the VM.
Given the heavily virtualized nature of the CI environment we want
to make sure to only trigger this due to a real deadlock and not
due to over-subscription of the systems resources.  This timeout
normally defaults to 20-30 seconds and this change increases it
to 120 seconds.

Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18624

* CI: Add alternative URLs for CentOS stream

Fallback to trying the "CentOS Strean Composes" repo for the qcow2
images if the regular URLs fail.  The Composes repo contains the daily
autobuilt Stream images.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18628

* Add additional verification of size fields and strings (#18623)

- Check for size fields that convert to smaller integers.
- Explicitly terminate bootenv string.
- Initialize variables that could be returned in an error case.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Reviewed-by: Alexander Motin 
Signed-off-by: Tony Hutter 
Closes #18623

* Fix uninitialized variable warning in zil_parse()

This resolves the following possible uninitialized variable warning
when building with --enable-code-coverage and gcc 8.5.0.

    module/zfs/zil.c: In function ‘zil_parse’:
    module/zfs/zil.c:549:47: warning: ‘end’ may be used uninitialized
    in this function [-Wmaybe-uninitialized]

Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18633

* ZTS: relax zpool_import_parallel_pos.ksh timing

Occasionally in the CI this test will fail because the parallel import
took longer than half of the serial time (but still less than the full
serial time).  Increase the cutoff to 3/4 of the serial time to preserve
the intent yet try and avoid these false positive failures.

Reviewed-by: Chris Longros 
Signed-off-by: Brian Behlendorf 
Closes #18634

* linux: verify stale znodes in legacy fallocate

The mode=0 and FALLOC_FL_KEEP_SIZE preallocation path can reach
zfs_freesp() directly and call zfs_statvfs() before going through the
normal zpl_enter_verify_zp() boundary.

When zfs_rezget() tears down a failed SA reload, a stale inode may
remain alive in the VFS with z_sa_hdl cleared. The unchecked
fallocate path can then reach sa_lookup(zp->z_sa_hdl, ...) through
zfs_statvfs() or zfs_freesp() and crash on a NULL SA handle.

Use zfs_enter_verify_zp() in zfs_statvfs() so stale znodes are
rejected under the teardown lock for both fallocate and statfs.
Also wrap the direct zfs_freesp() call in
zpl_enter_verify_zp()/zfs_exit() so this path follows the same
validation rules as the other Linux ZPL file operations.

Fixes: f734301d2267
("linux: add basic fallocate(mode=0/2) compatibility")

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Co-authored-by: gality369 
Closes #18458

* Linux: annotate nested xattr setattr znode locks

zfs_setattr() updates both the target znode and its hidden xattr
directory when ownership, mode, or project ID changes. The xattr
directory uses the same z_acl_lock and z_lock classes as the
parent znode, so lockdep reports recursive locking when the
second znode's mutexes are acquired.

This is a lockdep false positive rather than a real deadlock.
attrzp is the target file's hidden xattr directory, and the code
does not acquire these znode mutexes in the reverse order.
Acquire the attrzp mutexes with mutex_enter_nested() so lockdep
treats them as nested.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Co-authored-by: gality369 
Closes #18506

* Linux: avoid znode list lock inversion during resume

Lockdep reports a circular locking dependency during mounted filesystem
rollback.  zfs_resume_fs() walks z_all_znodes under z_znodes_lock and
calls zfs_rezget(), which takes the per-object znode hold lock via
zfs_znode_hold_enter().

The normal zget path takes these locks in the opposite order.
zfs_zget() takes the per-object hold lock before zfs_znode_alloc()
inserts the znode on z_all_znodes under z_znodes_lock.  Resume can
therefore establish z_znodes_lock -> zh_lock while normal lookup
creates zh_lock -> z_znodes_lock.

Pin the current and next znodes with igrab() while holding the list
lock, then drop the list lock before reloading the znode.  Existing
stale inode handling is preserved, and both the suspended reference
and temporary walk reference are released asynchronously.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes #18517

* linux/zpl_super: handle 'source' option directly

vfs_parse_fs_param_source() didn't appear until 5.14, and was not
backported to kernel.org LTS kernels. It's simple enough that it's
easier to just handle it ourselves rather than use a configure check.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18529

* linux: suppress reclaim lockdep in zfs_inactive via rwlock wrappers

kswapd can enter zfs_inactive() from inode reclaim while holding
fs_reclaim. The z_teardown_inactive_lock still serializes teardown,
but the reclaim-thread acquire/release pair can produce a lockdep
cycle through zfs_zinactive() and zfs_rmnode().

Add Linux rwlock nolockdep wrappers alongside the existing rwlock
macros and use them only for the reclaim-thread
z_teardown_inactive_lock acquire/release in zfs_inactive(). Keep
the real rwsem semantics unchanged and leave CONFIG_LOCKDEP
handling in the platform rwlock layer.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes #18505

* config: show progress output for kernel API checks

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18554

* linux/super: properly apply ro/rw mount option to superblock

f5a9e3a622 changed how SB_RDONLY was applied to the new mount in a way
that was too simplistic - it only sets readonly on the filesystem if the
mount was 'ro', but it never clears it if the mount was 'rw'. This
causes the 'rw' option to effectively be ignored, and so the readonly=
property wins out.

This fixes it by doing it the right way: checking the flags mask to see
if it was actually provided as an option at all, and then setting or
clearing it as appropriate.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18557
Closes #18563

* Linux 5.6 compat: fix fs_parse API mismatch

Added m4 macro to check fs_parse API signature and wrappers.  Before 
5.6, fs_parse() took a struct fs_parameter_description which wraps
the parameter specs with name and enum pointers. From 5.6, the 
description struct was removed and fs_parse() accepts the 
fs_parameter_spec directly.

Reviewed-by: Rob Norris 
Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18585

* Fix aarch64 build failure by removing earlyclobber (#18532)

The UVR macros used "+&w" (read-write + earlyclobber) as the
constraint for NEON register operands that are declared as explicit
hard-register variables via:

register unsigned char wN asm("vN") __attribute__((vector_size(16)));

The + modifier implicitly makes the operand also an input (reading the
register before the asm runs). The & (earlyclobber) modifier says "this
output may be written before all inputs are consumed." Having an
earlyclobber output on the same hard-register that is simultaneously
an input is a contradiction — GCC 16 now strictly diagnoses this.

The fix removes the & from "+&w", yielding "+w". The earlyclobber
was both incorrect (contradicts the implicit input) and unnecessary
(the physical registers are already hard-bound, so the compiler has no
freedom to assign conflicting registers anyway).



Issue #18525

Signed-off-by: Brian Behlendorf 
Co-authored-by: Claude Sonnet 4.6 
Reviewed-by: Tony Hutter 

* Fix "panic: cache_vop_rename: lingering negative entry"

A FreeBSD ZFS filesystem with properties "utf8only=on" and
"normalization=formD" consistently produces this panic when
building the lang/perl-5.42.0 port.

A ZFS file system with "utf8only=off" and "normalization=none"
works fine.

The cause of the panic seems to be incorrectly using the FreeBSD
namecache when normalisation is present. This commit adds a
predicate to prevent that.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Jan Martin Mikkelsen 
Closes #18430

* key lookup failure should always return EACCES

spa_do_crypt_abd() already maps a missing key to EACCES. However
spa_do_crypt_mac_abd(), spa_do_crypt_objset_mac_abd(), and
spa_crypt_get_salt() still return the raw
spa_keystore_lookup_key() error (ENOENT). This is inconsistent
As we want to treat all “no key” failures as a permission
failure. Standardize on EACCES for the unloaded-key case.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alek Pinchuk 
Closes #18448

* Fix off-by-one in PREVIOUSLY_REDACTED handler that drops last block

In send_reader_thread(), the PREVIOUSLY_REDACTED handler computed
file_max as MIN(dn->dn_maxblkid, range->end_blkid).  dn_maxblkid is
an inclusive maximum block ID while range->end_blkid is exclusive (one
past the last block).  The resulting file_max was then used as an
exclusive loop bound, causing the last block of any file (at index
dn_maxblkid) to be silently skipped when a PREVIOUSLY_REDACTED range
covered the end of the file.

The block was never written to the send stream so the receiver kept
zeros there.  ZFS reported no error because the stream itself was
valid; the data was simply absent.

Fix: use dn_maxblkid + 1 so file_max is consistently exclusive.

Add a regression test (redacted_max_blkid.ksh) that modifies only the
last block of a file in one clone, creates a redaction bookmark from
it, then sends an unmodified clone incrementally from that bookmark.
The PREVIOUSLY_REDACTED path must fill in the last block; the test
verifies it is not zeros and matches the original.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Paul Dagnelie 
Reviewed-by: Reviewed-by: Tony Hutter 
Signed-off-by: Manoj Joseph 
Closes #18477

* Avoid flushing unrelated NFS exports on snapshot unmount

zfsctl_snapshot_unmount() called exportfs_flush() before every umount
attempt to drop NFS export cache references that pin the snapshot
mountpoint.  The flush has global effect on the host's NFS exports and
clients, so paying it on every snapshot unmount (including auto-expire
rounds for snapshots that were never NFS-accessed) impacts unrelated
snapshots and clients.

ZFS cannot invalidate individual export cache entries because the
relevant sunrpc cache APIs are exported GPL-only.  Defer the global
flush so it runs only when the umount has actually failed, then retry
once.  Snapshots that are not NFS-pinned succeed on the first attempt
and never trigger the flush.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Youzhong Yang 
Signed-off-by: Ameer Hamza 
Closes #18476

* zfs: annotate nested dd_lock in reservation sync accounting

When reservation sync updates a child's reserved space, it rolls the
delta into ancestor space accounting while still holding the child's
dd_lock.  That locking order is intentional, but Linux lockdep sees
the ancestor acquisition as recursive because it lacks a nested lock
subclass annotation.

Teach the reservation-sync space-accounting path to acquire ancestor
dd_lock instances with a nested subclass.  Keep the existing public
interfaces and accounting behavior unchanged by routing only the
ancestor rollup through local helpers.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Signed-off-by: gality369 
Closes #18497

* sa: fix sa_add_projid lock ordering

sa_add_projid() currently acquires hdl->sa_lock before zp->z_lock.
Several same-znode update paths take zp->z_lock and then call
sa_update() or sa_bulk_update() on the same SA handle.

On Linux, FS_IOC_FSSETXATTR reaches zfs_setattr() through
zpl_ioctl_setxattr() without outer inode serialization. This makes
the reversed lock order a real ABBA deadlock rather than a lockdep
false positive when projid is added to an old-format inode while
another thread updates the same znode.

Acquire zp->z_lock before hdl->sa_lock in sa_add_projid() to match
the existing znode update ordering.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Co-authored-by: gality369 
Closes #18503

* zdb: detect BRT and DDT leaks during block traversal

During -b traversal, track BRT and DDT reference counts and report
blocks claimed more times than their reference tables account for
if it causes claim errors, instead of just asserting it.  Also
report entries with references not fully consumed by the traversal.

Add zdb leaks checks to cloning and dedup tests. This should make
sure the pools are in a sane state after completing the functional
tests.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #18494

* zarcstat: detect attached L2ARC device with no data

zarcstat and zarcsummary detected L2ARC presence using the l2_size
kstat, which is data held in L2ARC, not whether a cache device is
attached. When a cache device was attached but empty (freshly added,
or fully evicted):

  - zarcstat rejected "-f l2*" with "Incompatible field specified!"
  - zarcsummary printed "L2ARC not detected, skipping section",
    hiding cumulative I/O history and health counters

Expose the existing l2arc_ndev counter as a new kstat l2_dev_count.
It is maintained by l2arc_add_vdev() and l2arc_remove_vdev(), so it
tracks attachment in real time. Use it in both tools, falling back to
l2_size for compatibility with older kernel modules.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Ameer Hamza 
Closes #18499

* Fix double free for blocks cloned after DDT prune

Before this change, for blocks marked with D flag but absent in DDT
(pruned from it), zio_ddt_free() fell back to ZIO_STAGE_DVA_FREE
without trying ZIO_STAGE_BRT_FREE first.  Same time such blocks
might be present in BRT, and not handling that would result in
double/multiple free.

This change makes ZIO_DDT_FREE_PIPELINE include ZIO_FREE_PIPELINE,
just adding required ZIO_STAGE_ISSUE_ASYNC and ZIO_STAGE_DDT_FREE,
and moves DDT stages before BRT.  This way, if the block is found
in DDT by zio_ddt_free(), the pipeline is short-circuited to
ZIO_INTERLOCK_PIPELINE, similar to what zio_brt_free() does.  If
not, then BRT is checked, and if also no match, the block is freed.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Rob Norris 
Signed-off-by: Alexander Motin 
Closes #18520

* arc: export additional required symbols

External consumers of arc_read() need to be able to destroy the
returned arc_buf_t.  Add the arc_buf_destroy() interface as an
exported symbol.

Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18533

* zap_impl: use flex array field for mzap_phys_t.mz_chunks

mz_phys_t is always a full-block allocation, with mz_chunks[] as an
array over the rest of the block past the header.

Recent Linux compiled with CONFIG_UBSAN will complain about this:

    UBSAN: array-index-out-of-bounds in module/zfs/zap.c:1236:28
    index 2 is out of range for type 'mzap_ent_phys_t [1]'

The fix is straightforward; simply convert this field to a flex member.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18550

* spl_kvmalloc: remove __GFP_COMP before calling vmalloc()

In cb1833023 we stopped using it for KM_VMEM allocations, since its not
a valid flag for vmalloc(). However, there's a fallback path for
non-KM_VMEM allocations to use vmalloc(), and we need to remove
__GFP_COMP there too to avoid a warning.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18558

* FreeBSD: Make it possible to build openzfs.ko with sanitizers

Add make options which let one respectively compile the kernel modules
with the address sanitizer, memory sanitizer, and undefined behaviour
sanitizer enabled.  This makes it much easier to run the ZTS with those
sanitizers enabled.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Signed-off-by: Mark Johnston 
Closes #18596

* enforce exact decompressed length for lz4, gzip, and zstd

Decompressors must expand a ZFS block to exactly the expected number
of bytes. Treat decompression to an unexpected length as failure, so
truncated or short output is not accepted as valid decompression. This
makes our handling of decompress return values consistent with the
decompression functions' APIs.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Alek Pinchuk 
Closes #18599

* dsl_scan: close errorscrub cursor on pause

If the cursor were ever to actively hold resources, not finalising it
would mean leaking those resources whenever the scrub is paused.

The cursor is already reinitialized from the stored serialized form
if/when it is resumed, so there's nothing we need from the old one, just
to release it.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Rob Norris 
Closes #18603

* When reading a vdev label skip libzfs_core_init()

There's no need to call libzfs_core_init() when `zdb -l` is used to
read a vdev label.

Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18606

* Simplify dnode_level_is_l2cacheable()

We should not dereference through dn_handle->dnh_dnode once we
already have a dnode pointer.  The result will be the same.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #18212

* Remove parent ZIO from dbuf_prefetch()

I am not sure why it was added there 10 years ago, but it seems not
needed now.  According to my tests removing it improves sequential
read performance with recordsize=4K by 5-10% by reducing the CPU
overhead in prefetcher.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Rob Norris 
Reviewed-by: Ameer Hamza 
Reviewed-by: Akash B 
Signed-off-by: Alexander Motin 
Closes #18214

* Fix log vdev removal issues

When we clear the log, we should clear all the fields, not only
zh_log.  Otherwise remaining ZIL_REPLAY_NEEDED will prevent the
vdev removal.  Handle it also from the other side, when zh_log
is already cleared, while zh_flags is not.

spa_vdev_remove_log() asserts that allocated space on removed log
device is zero.  While it should be so in perfect world, it might
be not if space leaked at any point.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #18277

* ZVOL: Add encryption key check for block cloning

Somehow during block cloning porting from file systems was missed
the check for identical encryption keys.  As result, blocks cloned
between unrelated ZVOLs produced authentication errors on later
reads.  Having same or different encryption root does not matter.

This patch copies dmu_objset_crypto_key_equal() call from FS side.

Reviewed-by: Ameer Hamza 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #18315

* abd: Fix stats asymmetry in case of Direct I/O

abd_alloc_from_pages() does not call abd_update_scatter_stats(),
since memory is not really allocated there.  But abd_free_scatter()
called by abd_free() does.  It causes negative overflow of some
ABD and possibly ARC counters.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Rob Norris 
Signed-off-by: Alexander Motin 
Closes #18390

* Tag zfs-2.4.3

META file and changelog updated.

Signed-off-by: Tony Hutter 

* Constify some rrd_*() functions

These don't modify the db, so just constify them while we're in the
area.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Signed-off-by: Kyle Evans 

* Add dbrrd_latest_time() to grab the latest timestamp in the db

Returns 0 if the database is empty, otherwise it returns the highest
value of the minutely db.  dbrrd_add() will already enforce the property
that these are monotonically increasing, so we won't try to second-guess
it.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Signed-off-by: Kyle Evans 

* freebsd: set mnt_time on the rootfs at mountroot time

FreeBSD's vfs_mountroot() will collect `mnt_time` from every filesystem
that we mounted and use the highest timestamp as a source for the system
time if we didn't get anything from an attached RTC.

Use the rrd mechanism added to gather up a notion of the latest time
and set it on mnt_time.  If the timestamp db is empty, we just fallback
to the uberblock timestamp and hope that that is in the right ballpark.

Relevant: FreeBSD PR254058[0] reporting the problem downstream

[0] https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=254058

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Reviewed-by: Chris Longros 
Signed-off-by: Kyle Evans 

* zio_ddt_write: compute have_dvas after taking dde_io_lock

In zio_ddt_write(), have_dvas and is_ganged were computed before
dde_io_lock was taken. A concurrent zio_ddt_child_write_done() error
path calls ddt_phys_unextend() under dde_io_lock, which can zero
DVA[0] while another thread is between computing have_dvas and taking
dde_io_lock. That thread then uses the stale have_dvas=1 to call
ddt_bp_fill(), copying the zeroed DVA into the BP. A zero DVA resolves
as a hole, producing blocks that read back as zeros with no checksum
error (silent data corruption).

Fix by moving have_dvas and is_ganged computation to after dde_io_lock
is taken, so they always reflect the current state of dde->dde_phys.

Regression introduced by a41ef36858 ("DDT: Reduce global DDT lock
scope during writes").

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Co-authored-by: Saju Palayur 
Signed-off-by: Saju Palayur 
Closes #18366
Closes #18544
(cherry picked from commit 6fb72fda0f60d9efb591e320f83f78b19ec451cc)
(cherry picked from commit 0a07efe3283df42db41a2060dc3ae0c455a7c056)

* ZTS: Pass dec instead of hex to mknod

On Ubuntu 26.04 the default mknod command returns an error when
provided the major and minor numbers in hex.  Switch to passing
decimal values.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* ZTS: statx_dioalign.ksh update to stride_dd

The uutils 0.8.0 version of dd appears to diverge from GNU behavior
and does not fail when an unaligned write O_DIRECT write is issued.
Update the test case to use stride_dd which is provided by the ZTS
so the expected syscall behavior can be verified.

Reviewed-by: Tino Reichardt 
Reviewed-by: George Melikov 
Reviewed-by: Tony Hutter 
Signed-off-by: Brian Behlendorf 
Closes #18547

* ZTS: delegate: add test for send sub-permissions

Regular send and raw send are actually separate operations with separate
permissions. This adds a test to test the combinations properly using
the existing permission test infrastructure.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18672
(cherry picked from commit 4d1d00f9fe0c87b2334613a0efaa758275f938b8)

* ZTS: remove send_delegation tests

These tests are doing the same tests as delegate/zfs_allow_send, and are
hard to follow and maintain. There's no need for them now, so drop them.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18672
(cherry picked from commit 562b96ced9f28b8517913e5e9f7eaf1686db7cd0)

* ZTS: delegate: add encryption option for test fixture datasets

The delegate test framework doesn't care about the encryption status of
the dataset under test, so by adding an option to create with encryption
the framework can be used to check encryption-related permissions
without any further fanfare.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit 8303a36488da79c13d0fcca4365d71d5180407c3)

* ZTS: delegate: check send permissions on encrypted datasets

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit bce9a8ef7d0b4c1b8e323ef2b045379177c20410)

* ZTS: delegate: test send:encrypted

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit 166a6672502c6398b3ef549d4a17f113f5cb2e8d)

* zfs_secpolicy_send: lift checks to common function for both

The permissions checks for send are a little involved because different
permissions grant different abilities, and there's two ways to initiate
a send.

This lifts the common permissions checks into a single function, and
ensures that we maintain a single dataset hold across all checks. This
will become important in the next commit when we need to check a
specific dataset property as part of the permission check.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit f0d69e6b16d93ab49f1914a29e64d5df6e4f7bc2)

* delegate: add 'send:encrypted' permission

send:encrypted is like send:raw, but only permits encrypted datasets to
be sent - raw send is not permitted for unencrypted datasets.

This commit creates the permission, wires it up, and adds the check for
it in zfs_secpolicy_send_impl(), if it is the last send permission
standing, the dataset is checked for its encryption state.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18673
(cherry picked from commit 97b9ba7a982e36d39a3cf7db271da2fab110769d)

* zbookmark_compare: handle "marker" bookmarks with negative levels

"Marker" bookmarks (those with zb_level == ZB_ROOT_LEVEL, ZB_ZIL_LEVEL
or ZB_DNODE_LEVEL) represent valid blocks, but are associated with a
dataset directly rather than with a specific object within it. They end
up on bookmark lists during scan prefetch, and so need to be sorted
ahead of any "true" object blocks.

The problem is that for negative levels, BP_SPANB produces a negative
shift, which is not legal C. Fortunately the results are used only for
comparison, so the worst possible behaviour in a forgiving compilation
environment is a mis-sort, which for the scan/traverse cases, means that
we haven't prefetched certain metadata before we actually need it. But
there _is_ UB in there, and UBSAN does rightly complain.

Here we fix all this by handling these bookmarks directly - sorting them
ahead of "true" object blocks, which is usually what scan/traverse will
prefer. And we don't do any interesting math on these bookmarks, so we
sidestep the whole UB thing.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #14777
Closes #18652

* arc: add a few invariant checks in release builds

Convert a couple ASSERTs invariants to VERIFYs to enforce them
in release builds to be able to root-case #18782 kernel panic,
whenever it happens again.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Andriy Tkachuk 
Closes #18840
(cherry picked from commit 023d44b9ef68f1eff6b97d98c6a16cdb4b93cf76)

* CI: Have zfs-build-packages workflow build tarballs on Alma (#18662)

Previously, zfs-build-packages would only build source tarballs
on Fedora due to problems with building them on RHEL 7.  That's
a relic of the past now, as we haven't supported RHEL 7 since
it went EOL in 2024.  With this change, we now build the tarballs
on both Alma and Fedora.

Signed-off-by: Tony Hutter 
Reviewed-by: Olaf Faaland 
Reviewed-by: Chris Longros 

* Update our CI runners to the newest FreeBSD 15.1 RELEASE (#18667)

Signed-off-by: Christos Longros 
Reviewed-by: Alexander Motin 
Reviewed-by: Tony Hutter 

* Linux 7.1 compat: META (#18682)

Update the META file to reflect compatibility with the 7.1
kernel.

Signed-off-by: Tony Hutter 
Signed-off-by: Rob Norris 
Reviewed-by: Chris Longros 

* Fix handling of _PC_HAS_HIDDENSYSTEM for FreeBSD

The hidden and system flags are only supported for
ZFS pools if the z_use_fuids is true.  Fix
zfs_freebsd_pathconf() to check this.

Reviewed-by: Alexander Motin 
Signed-off-by: Rick Macklem 
Closes #18688

* zfs_ioctl: fix EBUSY race between quota queries and mount

zfsvfs_hold() fell back to zfsvfs_create() -> dmu_objset_own()
(exclusive) for unmounted datasets.  A concurrent zfs_domount()
also calls dmu_objset_own(), causing EBUSY on the same dataset.

Introduce zfsvfs_create_hold() using dmu_objset_hold() (shared
hold) instead.  Shared holds do not conflict with exclusive owns,
eliminating the race.  The release path (zfsvfs_rele,
zfsvfs_create_impl error) uses dmu_objset_ds()->ds_owner to
determine whether to disown or rele, avoiding the need for an
extra flag in zfsvfs_t.

Added tests userspace_005, groupspace_005, projectspace_006
(50 iter race test).

Reviewed-by: Brian Behlendorf 
Reviewed-by: Tony Hutter 
Signed-off-by: HeonJe Lee 
Closes #18611

* CI: Re-allow workflow_dispatch on zfs-qemu

Allow zfs-qemu to be invoked from a workflow_dispatch event (a.k.a,
manually running a workflow).  This may have been accidentally disabled
in 1916c2c55.

Reviewed-by: Chris Longros 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Tony Hutter 
Closes #18680

* initramfs-zfs should not try to copy directories

We had find only return files from the beginning for libgcc.so, but not
libfetch/libcurl. This oversight affected a user when vmware installed
its own libcurl.so.4 in a directory called libcurl.so.4, since our code
then tried to copy a directory, which fails.

Reviewed-by: Chris Longros 
Reviewed-by: Brian Behlendorf 
Suggested-by: Carsten Härle 
Signed-off-by: Richard Yao 
Closes #18582
Closes #18686

* Clean up embedded slog metaslab across txgs

On a read-write import, metaslab_set_fragmentation() can dirty a
metaslab via vdev_dirty() while still in the txg==0 load path when its
space map has an unexpected bonus size (e.g. a makefs-created pool
whose space-map dnodes use the boot loader's 24-byte space_map_phys_t
with nblkptr=3, giving db_size=64). If that metaslab is then selected
as the embedded slog, vdev_metaslab_init() only removed it from
vdev_ms_list when txg != 0, so the txg==0 case left it queued and
metaslab_fini() tripped VERIFY(!txg_list_member(&vd->vdev_ms_list,
msp, t)).

Remove slog_ms from the dirty list for every TXG_SIZE slot before
metaslab_fini() so the cleanup is correct regardless of txg.

Reported on FreeBSD as PR 281520:
External-issue: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=281520

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Nick Price 
Closes #18693

* README: update supported FreeBSD release to 15.1

Our CI runners moved to FreeBSD 15.1 in 0a4b59765 (#18667), but the
README still lists 15.0. Update it to match the CI version.

Reviewed-by: Brian Behlendorf 
Reviewed-by: Alexander Motin 
Signed-off-by: Christos Longros 
Closes #18696

* honor file argument in file_wait_event

grep the log path passed by the caller instead of always using
ZED_DEBUG_LOG.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Alek Pinchuk 
Closes #18700

* Update mtime/ctime when fallocate grows a file

Growing a file with fallocate updated its size but left mtime/ctime
unchanged and didn't log the change. A fallocate that changes the file
size should update mtime/ctime, and the change should be logged so it
survives a crash.
Pass log=TRUE to zfs_freesp() on the extend path so it updates the
timestamps and logs the size change, matching zfs_space(). Punch-hole
and zero-range already use this path and are unaffected.

Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes #18573

* ddt_log: Fix refcount tagging for begin/commit

Sponsored-by: Klara, Inc.
Sponsored-by: Wasabi Technology, Inc.
Reviewed-by: Rob Norris 
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Igor Ostapenko 
Closes #18706

* CI: Increase default watchdog NMI timeout on Linux

When the watchdog driver is configured and enabled an NMI will be
generated when the watchdogd process fails to regularly reset the
watchdog timer.  Given the heavily virtualized and potentially
over-subscribed nature of the CI environment increase the default
timeout to 120 seconds (normally defaults to 30 seconds).

Reviewed-by: Christos Longros 
Signed-off-by: Brian Behlendorf 
Closes #18704

* Fix race between device removal completion and pool export

vdev_remove_complete() finalizes a device removal in two phases under
the spa lock framework. Between the two phases it called
spa_vdev_exit(), which drops both the config locks (SCL_ALL) and
spa_namespace_lock and blocks on a txg sync. By that point
vdev_remove_replace_with_indirect() has already set svr->svr_thread =
NULL, and that is the only thing the export path (spa_export_common()
-> spa_async_suspend() -> spa_vdev_remove_suspend()) waits on. Once the
namespace lock is dropped, a concurrent export or destroy can acquire
it and set spa->spa_export_thread. When the removal thread re-enters
for its second phase via spa_vdev_enter(), it trips the
ASSERT0P(spa->spa_export_thread) assertion.

Hold spa_namespace_lock across both phases instead of dropping and
re-taking it: the intermediate spa_vdev_exit() becomes
spa_vdev_config_exit(), which drops only SCL_ALL and syncs the txg
while keeping the namespace lock held, and the second spa_vdev_enter()
becomes spa_vdev_config_enter(). Because the namespace lock is never
dropped between the phases, a concurrent export blocks in
spa_namespace_enter() and cannot set spa_export_thread until removal
finalization is done. This mirrors the multi-phase locking pattern
already used by the attach/detach and split paths in spa.c.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Prakash Surya 
Closes #18657

* linux: handle mmap read beyond file size

When performing a mmap read past the end of a file there is no data to
read, so simply zero-fill the page and return success.  zfs_getpage()
limits the range lock appropriately to cover the offset being read.

Reported-by: Iliya Polihronov (@vnsavage) (Automattic)
Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18715

* Using net/cloud-init to unpin specific python

These days freebsd 15/16 fail when fetching py311-cloud-init.
Switch to net/cloud-init to avoid python version pinning.

Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18717

* linux/abd: remove BIO support functions

Not used since the "classic" vdev_disk implementation was removed in
5764e218ba.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18719

* Linux 7.2: zpl_super: convert to sget_fc()

The old sget() superblock matcher has been removed in favour of the
fscontext-based sget_fc(). This converts to it. Its largely a signature
change, no functional change.

sget_fc() has existed since fscontext was introduced, so there's no need
for separate feature tests.

Sponsored-by: TrueNAS
Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18677

* zpl_ctldir: remove comments describing ancient kernel behaviour

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18722

* Fix insufficient locking in dedup verify

Introduction of dde_io_lock removed global DDT lock acquisition
from write completion.  As result, white ZIO ABD could be freed
while zio_ddt_collision() is comparing against it.  Taking there
dde_io_lock should fix the issue.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Alexander Motin 
Closes #17960
Closes #18712
Closes #18720

* libzfs: fix MS_CRYPT/MS_OVERLAY collision with umount2(2) flags

MS_CRYPT and MS_OVERLAY are libzfs-internal mount flags, but their
values (0x8 and 0x4) aliased the umount2(2) flags UMOUNT_NOFOLLOW and
MNT_EXPIRE. A consumer that legitimately set UMOUNT_NOFOLLOW therefore
had that bit read as MS_CRYPT, so libzfs unloaded the dataset's
encryption key as a side effect.

Move both flags to high bits unused by umount2(2) and strip them before
the unmount syscall in do_unmount() (umount2(2) on Linux, unmount(2) on
FreeBSD) and in cmd/zfs. MS_CRYPT is a compile-time macro, so consumers
that set it (for example truenas_pylibzfs) must be rebuilt against the
new header.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Ameer Hamza 
Closes #18713

* ZTS: ctime_001_pos increase tolerance

The ctime_001_pos test checks that timestamp updates occur for a
file after performing certain operations (read, write, chown, etc).
The test case allowed for a +4 second tolerance in the timestamp
value which is generous but up to +7 second discrepencies have been
seen in the CI.  Bump the tolerance to +10 seconds to prevent these
false positives.  As long as the value increases and is reasonably
close to the expected value consider that to be sufficient.

Reviewed-by: Christos Longros 
Reviewed-by: Alexander Motin 
Signed-off-by: Brian Behlendorf 
Closes #18733

* build: Fix for building dist target outside of project root

Do not change the working directory when doing copying because $distdir
is a path relative to the working directory, and thus not valid after
changing the working directory. This previously worked, when configure
was run from the project root, because @srcdir@ was the same as the
build working directory.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Glenn Washburn 
Closes #18744

* Linux: fix zfs_write() infinite loop on unfaultable buffer

On Linux, zfs_write() copies from the source buffer with page faults
disabled while the transaction is open, relying on
zfs_uio_prefaultpages() to make the pages resident beforehand.  When
dmu_write_uio_dbuf() returns EFAULT, the retry path only subtracts
the bytes consumed so far from the prefault accounting.  On zero
progress pfbytes does not change, the "pfbytes < nbytes" check never
triggers another prefault, and the loop retries the same failing
copy.  If the buffer can never be faulted in again, e.g. the owning
process was torn down while a thread was in pwrite(2), that thread
spins unkillably at 100% CPU holding the file range lock, blocking
every other accessor of the file.

This is a regression from commit b0cbc1aa9a ("Use big transactions
for small recordsize writes."), which dropped the unconditional
re-prefault the EFAULT path had carried since commit 779a6c0bf6
("deadlock between mm_sem and tx assign in zfs_write() and page
fault").  Restore those semantics by resetting pfbytes on EFAULT, as
suggested in the issue analysis: the next iteration then faults the
pages back in before retrying, and for a permanently inaccessible
buffer zfs_uio_prefaultpages() fails, breaking the loop with EFAULT,
which is already propagated to userspace.  Transient faults, such as
mmap'ed source pages evicted under memory pressure, retry as before.

Built and tested on Linux aarch64: the ZTS mmap and write-path groups
pass and a munmap-versus-write stress run leaves no stuck writers.  The
teardown race itself is not reproducible on demand, so the retry paths
were also checked by inspection against the reproducers in the issue.

Reviewed-by: Alexander Motin 
Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #17129
Closes #18740

* linux: batch DMU reads of non-resident pages in mappedread()

When a range being read has at least one page in the page cache,
zfs_read() routes the whole chunk through mappedread(), which falls
back to a separate dmu_read_uio_dbuf() call for every non-resident
PAGE_SIZE piece.  Since cached pages outlive munmap(), a file which
was mapped at some point may sit mostly outside the page cache and
still pay this cost: one DMU call per 4K page instead of one per
chunk, measured in #16031 as a 4-10x sequential read slowdown.
Commit 39be46f43 ("Linux 5.18+ compat: Detect filemap_range_has_page")
fixed the detection side so fully uncached chunks bypass mappedread()
again, but a chunk holding even one resident page still degrades to
page-sized DMU reads for everything else.

Instead of issuing one DMU read per absent page, probe the page cache
with find_get_page() and extend the read over the whole run of
non-resident pages which follows, restoring chunk-sized DMU reads for
the uncached parts of a mapped file.

A page can be faulted in after it was observed absent and before the
DMU read covering it completes, but this is safe for the same reason
the existing single-page window is.  zfs_read() holds the znode
rangelock as reader across mappedread(), so the DMU contents of the
range are stable: zfs_write(), zfs_putpage() and truncation all
require the writer lock, and zfs_getpage() fills concurrently faulted
pages from those same contents.  A faulted page can only diverge from
the DMU once dirtied through a writable mapping, making that store
concurrent with this read, for which returning the pre-store data is
a valid outcome.  Stores which completed before the read began cannot
be missed: a dirty page cannot be cleaned and reclaimed while the
reader lock is held (writeback takes the writer lock), so it is still
found resident, or zfs_putpage() already copied its data into the
DMU.

FreeBSD's mappedread() has the same per-page fallback and could be
batched the same way in a follow-up.

Measured in a VM with a 1 GiB file held in the ARC and read
sequentially with dd: one resident page per 1 MiB read request
degrades throughput from ~11.4 GB/s (no resident pages) to ~5.7 GB/s
on the baseline, and this change restores ~11.4 GB/s; one resident
page per 32 MiB chunk, read in 32 MiB requests, improves from
~5.4 GB/s to ~7.3 GB/s.  Reads of a fully resident file are
unaffected (~20 GB/s before and after).  All tests in the ZTS mmap
group pass, including the mmap_read and mmap_seek cases.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #16031 
Closes #18741

* Add SECURITY.md policy file

Add a basic SECURITY.md file to establish the repository's security
reporting policy.  Includes guidance for reporting security issues
and what to expect.

Reviewed-by: Allan Jude 
Reviewed-by: George Melikov 
Signed-off-by: Brian Behlendorf 
Closes #18766

* Fix receive -x according to comment

The old condition skipped the -x for ANY property whose source wasn't 
explicitly ZPROP_SOURCE_VAL_RECVD - which caught inherited/default 
properties too, not just locally-set ones. The new condition correctly 
skips only when the property is locally-set on the destination 
(source == fsname), which is the documented intent.

Reviewed-by: Brian Behlendorf 
Signed-off-by: Richard Kojedzinszky 
Closes #18737
Closes #18738

* Fix receive of split large blocks with a short trailing chunk

A dataset with a large recordsize can store a single-block file whose
block size is not a power of two.  When such a block is sent without
large blocks (no -L), the sender splits it into SPA_OLD_MAXBLOCKSIZE
(128K) chunks, and the final chunk is smaller than the block size.

flush_write_batch_impl() already handles any WRITE record whose size
differs from the object's block size by doing a normal dmu_write(), but
it first asserted the record was always larger than the block size.  The
shorter trailing chunk violates that assertion, so receiving such a
stream panicked the receive_writer thread on debug builds; production
builds took the correct dmu_write() path and were unaffected.

Drop the assertion and describe both size-mismatch cases in the comment;
the dmu_write() path already handles a record smaller than the block
size.  Add an rsend test that sends such a block without -L (initial and
incremental, plain and compressed) and verifies the received file
matches.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Issue #17829
Closes #18749

* ZTS: migration/setup: clear stale zfs_member label before new_fs

During a full ZTS run functional/migration/setup fails intermittently
when it mounts the non-ZFS device.  That device is often one an earlier
test used as a pool vdev.  'zpool destroy' leaves the vdev labels in
place and new_fs only overwrites the front of the device, so the
trailing labels can survive.  libblkid then probes the device as
ambiguous (both the new filesystem and zfs_member) and the
auto-detecting mount refuses, which setup reports as a spurious failure.

Wipe any residual signatures with wipefs before laying down the new
filesystem so the device carries a single, unambiguous type, and let
udev settle before the mount.  Skip the wipe in the single-disk case,
where the non-ZFS device is the same one the test pool was just created
on, so the live pool is left untouched.

Verified on Linux: after a pool create and destroy the scratch device
still carries a zfs_member label (blkid -p reports zfs_member); a
wipefs -a removes it so the following new_fs is the only signature and
the mount succeeds.  The functional/migration group passes with the
change.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #18492
Closes #18753

* ZTS: stop zpool_initialize tests racing initialize to completion

zpool_initialize_import_export and zpool_initialize_suspend_resume start
initializing a one-disk pool, wait a fixed couple of seconds, and then
expect initializing to still be running so it can be observed across an
export/import and suspended.  On a small or fast vdev the default 1 MiB
initialize chunk lets the whole disk finish within that window, after
which "zpool initialize -s" fails with "there is no active
initialization" and the test fails.

Throttle initializing with zfs_initialize_chunk_size, exactly as the
zpool_wait_initialize_* tests already do, so it stays active long enough
to observe regardless of vdev size or speed.  The tunable is saved and
restored per test.  Cleanup destroys the pool before restoring the chunk
size: the initialize thread rereads zfs_initialize_chunk_size on every
write but allocates its fill buffer once at the smaller size, so raising
it back while the thread is still running would issue a write larger
than that buffer.

With this both tests pass reliably, so drop their entries from the
zts-report.py.in maybe list (#11948, #17311).

Tests: ran both tests against the reproduced race in a VM -- with the
default chunk size the suspend step fails with "there is no active
initialization"; with the smaller chunk size initializing is still
running across the export/import and suspend, and the tunable is
restored on exit.

Reviewed-by: Brian Behlendorf 
Signed-off-by: MorganaFuture <103630661+MorganaFuture@users.noreply.github.com>
Closes #11948
Closes #17311
Closes #18771

* zpl_inode: remove zpl_rename no-flags variants

Removed in 4.9.

Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf 
Signed-off-by: Rob Norris 
Closes #18769

* CI: Fix race caused by shared ctr file updates

CTR is shared between the VMs and used as a global counter.  This 
uncoordinated shared access can result in a CI failure due to the
racing updates.  From the log:

`qemu-6-tests.sh: line 27: 1`
`6: syntax error in expression
(error token is "6")`

Resolve the issue by using separate ctr files by appending the ID.
The output now prints the total test cases count along with each VMs
individual count.

Reviewed-by: Tino Reichardt 
Reviewed-by: Brian Behlendorf 
Signed-off-by: tiehexue 
Closes #18778

* ZTS: fix zpool_initialize_multiple_pools suspend race

zpool_initialize_multiple_pools verifies that "zpool initialize -a -s"
suspends initializing on all four pools. It starts initializing and,
without slowing it down, expects every pool to still be initializing
when the suspend runs.

On fast storage a pool can finish initializing before the suspend, so
"zpool initialize -a -s" reports "there is no active initialization"
and the pool's status is "completed" rather than "suspended". The
suspend command then fails the test intermittently.

Throttle zfs_initialize_chunk_size before the suspend phase, the same
way the zpool_wait_initialize_* tests do, so initializing stays active
through the suspend and cancel checks. The earlier phases that wait for
initializing to finish ("zpool wait" and "-w -a") are left at the
default chunk size so they still complete promptly.

The chunk size is saved and restored, and the pools are destroyed
before it is restored so a running initialize thread never sees a
larger value than the buffer it allocated at the smaller size.
…
pull Bot pushed a commit to A-Archives-and-Forks/openzfs that referenced this pull request Sep 26, 2026
kswapd can enter zfs_inactive() from inode reclaim while holding
fs_reclaim. The z_teardown_inactive_lock still serializes teardown,
but the reclaim-thread acquire/release pair can produce a lockdep
cycle through zfs_zinactive() and zfs_rmnode().

Add Linux rwlock nolockdep wrappers alongside the existing rwlock
macros and use them only for the reclaim-thread
z_teardown_inactive_lock acquire/release in zfs_inactive(). Keep
the real rwsem semantics unchanged and leave CONFIG_LOCKDEP
handling in the platform rwlock layer.

Reviewed-by: Brian Behlendorf 
Signed-off-by: ZhengYuan Huang 
Closes openzfs#18505
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Status: Accepted Ready to integrate (reviewed, tested)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants