| SYSTEMD-VMSPAWN(1) | systemd-vmspawn | SYSTEMD-VMSPAWN(1) |
NAME
systemd-vmspawn - Spawn an OS in a virtual machine
SYNOPSIS
systemd-vmspawn [OPTIONS...] [ARGS...]
DESCRIPTION
systemd-vmspawn may be used to start a virtual machine from an OS image. In many ways it is similar to systemd-nspawn(1), but launches a full virtual machine instead of using namespaces.
File descriptors for /dev/kvm and /dev/vhost-vsock can be passed to systemd-vmspawn via systemd's native socket passing interface (see sd_listen_fds(3) for details about the precise protocol used and the order in which the file descriptors are passed), these file descriptors must be passed with the names "kvm" and "vhost-vsock" respectively.
Note: on Ubuntu/Debian derivatives systemd-vmspawn requires the user to be in the "kvm" group to use the VSOCK options.
OPTIONS
The excess arguments are passed as extra kernel command line arguments using SMBIOS.
The following options are understood:
-q, --quiet
Added in version 256.
--system, --user
Added in version 260.
Image Options
-D, --directory=
One of either --directory= or --image= must be specified. If neither are specified --directory=. is assumed.
Note: If mounting a non-root owned directory you may require --private-users= to map into the user's subuid namespace. An example of how to use /etc/subuid for this is given later.
Added in version 256.
-x, --ephemeral
Added in version 260.
-i, --image=
Added in version 255.
--image-format=FORMAT
Added in version 260.
--image-disk-type=TYPE
Added in version 261.
--discard-disk=BOOL
Added in version 256.
--grow-image=BYTES, -G BYTES
If --ephemeral is specified, the original image file is left untouched and the requested size is applied to the ephemeral copy instead. In that case qcow2 images are supported too.
Added in version 258.
Host Configuration
--cpus=CPUS
Added in version 255.
--ram=BYTES[:MAXBYTES[:SLOTS]]
Added in version 255.
--kvm=BOOL
Added in version 255.
--vsock=BOOL
Added in version 255.
--vsock-cid=CID
Added in version 255.
--tpm=BOOL
Added in version 256.
--tpm-state=PATH|auto|off
If --ephemeral is specified, "auto" behaves like "off".
Added in version 258.
--efi-nvram-template=PATH
Added in version 261.
--efi-nvram-state=PATH|auto|off
If --ephemeral is specified, "auto" behaves like "off".
Added in version 261.
--secure-boot=BOOL
With --coco=, setting this to "yes" additionally requires the "enrolled-keys" firmware feature, overriding its default exclusion: confidential computing firmware is stateless, so keys cannot be enrolled at runtime and Secure Boot is only operative with keys baked into the image at build time, which in turn makes the firmware refuse to boot unsigned images from the very first boot. The default exclusion of "enrolled-keys" hence keeps unsigned images bootable.
Added in version 255.
--firmware=PATH
Added in version 256.
--firmware-features=FEATURE[,FEATURE...]
Added in version 261.
--coco=
Configures whether to run the guest as a confidential VM. Takes one of "no", "sev-snp" or "tdx". Defaults to "no".
"sev-snp" enables AMD SEV-SNP. This requires KVM on an x86_64 host with SNP-capable hardware and firmware. A suitable SNP-built OVMF firmware is picked automatically from the installed QEMU firmware descriptors, by requiring the "amd-sev-snp" firmware feature; use --firmware= with a path to a firmware descriptor file to select a specific one. Secure Boot is effectively off by default (see --secure-boot=). Direct kernel boot via --linux= is required so that the kernel, initrd and command line are hashed into the launch measurement ("kernel-hashes=on"); booting the kernel off the disk image via the firmware would leave it outside the measurement. Credentials passed via --set-credential= or --load-credential= are bundled into a cpio archive appended to the initrd (mirroring what systemd-stub does for ESP credentials), so they enter the launch measurement via "kernel-hashes=on"; the SMBIOS and fw_cfg channels normally used to deliver credentials are not used because they are unmeasured and would be discarded by PID1 in confidential guests. This channel is measured but not confidential with respect to the host or VMM: the initrd (and thus the credentials it carries) is supplied to QEMU as plaintext and only its hash enters the launch measurement, which guarantees integrity but does not keep the credentials secret from the host. This requires the guest to run a sufficiently recent version of systemd (supporting /.extra/system_credentials/). A vTPM, if attached via --tpm=, must be treated as untrusted by the guest.
"tdx" enables Intel TDX. This requires KVM on an x86_64 host with TDX-capable hardware and a TDX-enabled host kernel. As with "sev-snp", a TDX-built OVMF (TDVF) firmware is picked automatically from the installed QEMU firmware descriptors, by requiring the "intel-tdx" firmware feature; use --firmware= with a path to a firmware descriptor file to select a specific one. The CPU model is fixed to "host". Firmware is measured into MRTD when the TD is built. Secure Boot cannot be enrolled at runtime (there is no writable NVRAM); its state is fixed by the selected TDVF image and is part of the measured firmware (see --secure-boot=). When booting a UKI, the whole UKI PE is measured into RTMR 1, and the loaded sections are measured individually by systemd-stub into RTMR 2. For direct linux boot, firmware measures the kernel PE into RTMR 1, and the Linux EFI stub measures initrd and command line into RTMR 2. Credentials passed via --set-credential= or --load-credential= are delivered through SMBIOS Type 11 OEM strings, which the firmware (TDVF) measures into RTMR 0. This channel is measured but not confidential with respect to the host or VMM, since the host assembles the SMBIOS table. A vTPM, if attached via --tpm=, must be treated as untrusted by the guest. To obtain TD Quotes for remote attestation, the guest is wired to the host's local TDX Quote Generation Service automatically: the unix socket /run/tdx-qgs/qgs.socket is used if it exists, otherwise vsock port 4050 on the host (cid 2). If the QGS is listening on neither channel, the guest's quote requests will fail. Added in version 262.
Added in version 261.
Networking Options
-n, --network-tap
Note: root privileges are required to use TAP networking. Additionally, systemd-networkd(8) must be running and correctly set up on the host to provision the host interface. The relevant ".network" file can be found at /usr/lib/systemd/network/80-vm-vt.network.
Added in version 255.
--network-user-mode
Added in version 255.
Execution Options
--linux=PATH
Added in version 256.
--initrd=PATH
--initrd= can be specified multiple times and vmspawn will merge them together.
Added in version 256.
--smbios11=STRING, -s STRING
Added in version 258.
--notify-ready=
Defaults to true. (Note that this is unlike the option of the same name to systemd-nspawn(1) that defaults to false.)
Added in version 258.
System Identity Options
-M, --machine=
Added in version 255.
--uuid=
Added in version 256.
Property Options
-S, --slice=
Added in version 258.
--property=
Added in version 258.
--register=
Added in version 256.
User Namespacing Options
--private-users=UID_SHIFT[:UID_RANGE]
If one or two colon-separated numbers are specified, user namespacing is turned on. UID_SHIFT specifies the first host UID/GID to map, UID_RANGE is optional and specifies number of host UIDs/GIDs to assign to the virtual machine. If UID_RANGE is omitted, 65536 UIDs/GIDs are assigned.
When user namespaces are used, the GID range assigned to each virtual machine is always chosen identical to the UID range.
Added in version 256.
Mount Options
--bind=PATH, --bind-ro=PATH
Added in version 256.
--extra-drive=[FORMAT:][DISKTYPE:]PATH
Added in version 256.
--bind-volume=PROVIDER:VOLUME[:CONFIG][:K=V,...]
The trailing comma-separated K=V list passes parameters to io.systemd.StorageProvider.Acquire(): template=, create= (one of "any", "new", "open"), read-only= (or ro=; takes a boolean or "auto"), size= / create-size= (size for created volumes), request-as= (one of "blk", "reg", "dir"; "dir" is rejected by vmspawn).
Each attached volume is identified by the name "PROVIDER:VOLUME". Volumes attached at startup via this option cannot be detached at runtime via machinectl unbind-volume; only volumes added at runtime via machinectl bind-volume are removable.
The provider is looked up under /run/systemd/io.systemd.StorageProvider/ for system mode (or $XDG_RUNTIME_DIR/systemd/io.systemd.StorageProvider/ for user mode), matching the runtime scope chosen via --user / --system.
Added in version 261.
--bind-user=
The combination of the two operations above ensures that it is possible to log into the virtual machine using the same account information as on the host. The user is only mapped transiently, while the virtual machine is running, and the mapping itself does not result in persistent changes to the virtual machine (except maybe for log messages generated at login time, and similar). Note that in particular the UID/GID assignment in the virtual machine is not made persistently. If the user is mapped transiently, it is best to not allow the user to make persistent changes to the virtual machine. If the user leaves files or directories owned by the user, and those UIDs/GIDs are reused during later virtual machine invocations (possibly with a different --bind-user= mapping), those files and directories will be accessible to the "new" user.
The user/group record mapping only works if the virtual machine contains systemd 258 or newer, with nss-systemd properly configured in nsswitch.conf. See nss-systemd(8) for details.
Note that the user record propagated from the host into the virtual machine will contain the UNIX password hash of the user, so that seamless logins in the virtual machine are possible. If the virtual machine is less trusted than the host it is hence important to use a strong UNIX password hash function (e.g. yescrypt or similar, with the "$y$" hash prefix).
Added in version 259.
--bind-user-shell=
Note: This will not check whether the specified shells exist in the virtual machine.
This operation is only supported in combination with --bind-user=.
Added in version 259.
--bind-user-group=NAME
Note: This will not check whether the specified groups exist in the virtual machine.
This operation is only supported in combination with --bind-user=.
Added in version 259.
Logging Options
--forward-journal=FILE|DIR
Added in version 256.
--forward-journal-max-use=BYTES, --forward-journal-keep-free=BYTES, --forward-journal-max-file-size=BYTES, --forward-journal-max-files=N
Added in version 261.
SSH Options
--pass-ssh-key=BOOL
The generated keys are ephemeral. That is they are valid only for the current invocation of systemd-vmspawn, and are typically not persisted.
Added in version 256.
--ssh-key-type=TYPE
By default, "ed25519" keys are generated, however "rsa" keys may also be useful if the VM has a particularly old version of sshd(8).
Added in version 256.
Input/Output Options
--console=MODE
Added in version 256.
--console-transport=TRANSPORT
Added in version 261.
--background=COLOR
Added in version 256.
Credentials
--load-credential=ID:PATH, --set-credential=ID:VALUE
In order to embed binary data into the credential data for --set-credential=, use C-style escaping (i.e. "\n" to embed a newline, or "\x00" to embed a NUL byte). Note that the invoking shell might already apply unescaping once, hence this might require double escaping!
Credentials are preferably passed to the VM via SMBIOS Type 11 strings or QEMU fw_cfg files. If neither mechanism is available, credentials are passed on the kernel command line using systemd.set_credential_binary= which is not a confidential channel. Do not use this for passing secrets to the VM in that case.
Under --coco=sev-snp, SMBIOS and fw_cfg are not covered by the SNP launch measurement and are discarded by PID1 in confidential guests. Credentials are therefore packaged into a cpio archive containing .extra/system_credentials/ID.cred entries and appended to the initrd that QEMU loads, so they enter the launch measurement via "kernel-hashes=on". PID1 imports them from the initramfs at boot. As with the kernel command line, this is a measured but not a confidential channel: QEMU receives the initrd (and thus the embedded credentials) as plaintext from the host and only its hash is covered by the launch measurement, so a modified initrd produces a different launch measurement that a relying party can detect via remote attestation, but the credentials are not hidden from the host or VMM. This requires the guest to run a sufficiently recent version of systemd (supporting /.extra/system_credentials/).
Added in version 255.
Other
--no-pager
-h, --help
--version
--no-ask-password
ENVIRONMENT
$SYSTEMD_LOG_LEVEL
$SYSTEMD_LOG_COLOR
This setting is only useful when messages are written directly to the terminal, because journalctl(1) and other tools that display logs will color messages based on the log level on their own.
$SYSTEMD_LOG_TIME
This setting is only useful when messages are written directly to the terminal or a file, because journalctl(1) and other tools that display logs will attach timestamps based on the entry metadata on their own.
$SYSTEMD_LOG_LOCATION
Note that the log location is often attached as metadata to journal entries anyway. Including it directly in the message text can nevertheless be convenient when debugging programs.
$SYSTEMD_LOG_TID
Note that the this information is attached as metadata to journal entries anyway. Including it directly in the message text can nevertheless be convenient when debugging programs.
$SYSTEMD_LOG_TARGET
$SYSTEMD_LOG_RATELIMIT_KMSG
$SYSTEMD_PAGER, $PAGER
Note: if $SYSTEMD_PAGERSECURE is not set, $SYSTEMD_PAGER and $PAGER can only be used to disable the pager (with "cat" or ""), and are otherwise ignored.
$SYSTEMD_LESS
Users might want to change two options in particular:
K
If the value of $SYSTEMD_LESS does not include "K", and the pager that is invoked is less, Ctrl+C will be ignored by the executable, and needs to be handled by the pager.
X
Note that setting the regular $LESS environment variable has no effect for less invocations by systemd tools.
See less(1) for more discussion.
$SYSTEMD_LESSCHARSET
Note that setting the regular $LESSCHARSET environment variable has no effect for less invocations by systemd tools.
$SYSTEMD_PAGERSECURE
This option takes a boolean argument. When set to true, the "secure mode" of the pager is enabled. In "secure mode", LESSSECURE=1 will be set when invoking the pager, which instructs the pager to disable commands that open or create new files or start new subprocesses. Currently only less(1) is known to understand this variable and implement "secure mode".
When set to false, no limitation is placed on the pager. Setting SYSTEMD_PAGERSECURE=0 or not removing it from the inherited environment may allow the user to invoke arbitrary commands.
When $SYSTEMD_PAGERSECURE is not set, systemd tools attempt to automatically figure out if "secure mode" should be enabled and whether the pager supports it. "Secure mode" is enabled if the effective UID is not the same as the owner of the login session, see geteuid(2) and sd_pid_get_owner_uid(3), or when running under sudo(8) or similar tools ($SUDO_UID is set [3]). In those cases, SYSTEMD_PAGERSECURE=1 will be set and pagers which are not known to implement "secure mode" will not be used at all. Note that this autodetection only covers the most common mechanisms to elevate privileges and is intended as convenience. It is recommended to explicitly set $SYSTEMD_PAGERSECURE or disable the pager.
Note that if the $SYSTEMD_PAGER or $PAGER variables are to be honoured, other than to disable the pager, $SYSTEMD_PAGERSECURE must be set too.
$SYSTEMD_COLORS
true
false
"16", "256", "24bit"
"auto-16", "auto-256", "auto-24bit"
$SYSTEMD_URLIFY
EXAMPLES
Example 1. Run an Arch Linux VM image generated by mkosi
$ mkosi -d arch -p systemd -p linux --autologin -o image.raw -f build $ systemd-vmspawn --image=image.raw
Example 2. Import and run a Fedora 44 Cloud image using importctl
$ curl -L \
-O https://download.fedoraproject.org/pub/fedora/linux/releases/44/Cloud/x86_64/images/Fedora-Cloud-Base-Generic-44-1.7.x86_64.qcow2 \
-O https://download.fedoraproject.org/pub/fedora/linux/releases/44/Cloud/x86_64/images/Fedora-Cloud-44-1.7-x86_64-CHECKSUM \
-O https://fedoraproject.org/fedora.gpg
$ gpgv --keyring ./fedora.gpg Fedora-Cloud-44-1.7-x86_64-CHECKSUM
$ sha256sum -c Fedora-Cloud-44-1.7-x86_64-CHECKSUM
# importctl import-raw -m Fedora-Cloud-Base-Generic-44-1.7.x86_64.qcow2 fedora-44-cloud
# systemd-vmspawn -M fedora-44-cloud
Example 3. Build and run systemd's system image and forward the VM's journal to a local file
$ mkosi build
$ systemd-vmspawn \
-D mkosi.output/system \
--private-users $(grep $(whoami) /etc/subuid | cut -d: -f2) \
--linux mkosi.output/system.efi \
--forward-journal=vm.journal \
enforcing=0
Note: this example also uses a kernel command line argument to ensure SELinux is not started in enforcing mode.
Example 4. SSH into a running VM using systemd-ssh-proxy
$ mkosi build
$ my_vsock_cid=3735928559
$ systemd-vmspawn \
-D mkosi.output/system \
--private-users $(grep $(whoami) /etc/subuid | cut -d: -f2) \
--linux mkosi.output/system.efi \
--vsock-cid $my_vsock_cid \
enforcing=0
$ ssh root@vsock/$my_vsock_cid -i /run/user/$UID/systemd/vmspawn/machine-*-system-ed25519
EXIT STATUS
If an error occurred the value errno is propagated to the return code. If EXIT_STATUS is supplied by the running image that is returned. Otherwise, EXIT_SUCCESS is returned.
SEE ALSO
systemd(1), mkosi(1), machinectl(1), importctl(1), UAPI.1 Boot Loader Specification[1]
NOTES
- 1.
- UAPI.1 Boot Loader Specification
- 2.
- ANSI Escape Code (Wikipedia)
- 3.
- It is recommended for other tools to set and check $SUDO_UID as appropriate, treating it is a common interface.
| systemd 262 |