Known issues
This section lists MOSK known issues with workarounds for the MOSK release 26.1.2.
Bare metal
[62999] Master node replacement fails on MOSK clusters with BGP announcement enabled
On MOSK clusters with BGP announcement enabled (useBGPAnnouncement: true in
the Cluster object), master nodes are deployed one per rack, and the
Kubernetes API Virtual IP (VIP), for example, 10.0.30.100/32 is assigned to
the loopback (lo) interface of each master node. For details on BGP
announcement, see Configure BGP announcement for cluster API LB address.
During master node replacement, lcm-agent on the new node tries to reach
the API VIP on the lo interface for the first time but cannot connect to
it. This happens because cloud-init binds the API VIP to the lo
interface before the local containerized API proxy starts listening on it, so
the connection is refused. The lcm-agent cannot report back to the
management cluster. As a result, the Machine object remains stuck in the
PendingLCMAgent state.
As a workaround, temporarily remove the API VIP from the lo interface to
force API traffic through the physical network to a healthy master node. Once
the node finishes provisioning, the local VIP handling takes over
automatically.
Workaround:
From the management cluster, obtain the IP address of the new node from its
IpamHostobject:kubectl get ipamhost <NODE_NAME> -o jsonpath="{.status.serviceMap['ipam/SVC-k8s-lcm'][0].ipAddress}"
SSH to the new node:
ssh -i <path-to-ssh-key> mcc-user@<NEW_NODE_IP>
Wait for
cloud-initto complete on the new node:sudo cloud-init status --wait
Before proceeding, ensure that the output reports
status: done.Warning
Do not proceed with networking modifications until
cloud-inithas finished executing, so it does not overwrite your changes.Verify that the API VIP is bound to the
lointerface:ip addr show dev lo
Example of system response on the affected node:
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1000 inet 127.0.0.1/8 scope host lo inet 10.0.30.100/32 scope global loConfirm that
lcm-agentis failing with a connection refused error:sudo journalctl -u lcm-agent* -n 20 --no-pager
Remove the
/32API VIP from thelointerface:sudo ip addr del 10.0.30.100/32 dev lo
Substitute
10.0.30.100with the actual API VIP of your cluster.Verify outbound network connectivity to the API. It should now route across the network to one of the healthy master nodes:
curl -k -v https://10.0.30.100:443/api
An HTTP response such as
HTTP/2 401or403confirms that the traffic is reaching an active Kubernetes API server. For example:* Connected to 10.0.30.100 (10.0.30.100) port 443 < HTTP/2 401 { "kind": "Status", "message": "Unauthorized", "code": 401 }Reload
systemdto clear the warnings about unit files that have changed on disk:sudo systemctl daemon-reload
Restart the
lcm-agentdaemon to trigger an immediate check-in:sudo systemctl restart lcm-agent*.service
Verify that the node recovers:
On the new node, inspect the agent logs to confirm a successful tick execution:
sudo journalctl -u lcm-agent* -f
In the system response, find the
tick finishedmessages without connection errors.From the management or seed node, verify the machine phase transition:
kubectl get machine,lcmmachine -o wide
The
LCMPHASEfor the new node should transition fromPendingLCMAgenttoPrepare,Deploy, and then toReady.
Ceph
[54195] Ceph OSD experiencing slow operations in BlueStore during MOSK deployment
Note
Since MOSK 26.2, this issue is documented in Troubleshooting Guide: Ceph OSD experiencing slow operations in BlueStore.
The description and workaround below remain valid for this release.
During MOSK cluster deployment, the following false-positive example alert for Ceph may raise:
Failed to configure Ceph cluster: ceph cluster verification is failed:
[BLUESTORE_SLOW_OP_ALERT: 3 OSD(s) experiencing slow operations in BlueStore]
The issue occurs due to the following upstream Ceph issues:
To verify whether the cluster is affected:
Enter the
pelagia-ceph-toolspod:Verify the Ceph cluster status:
Verify Ceph health:
ceph -sExample of a positive system response in the affected cluster:
cluster: id: 6ae41eb3-262e-4da9-8847-25efed2fcaa2 health: HEALTH_WARN 2 OSD(s) experiencing slow operations in BlueStore services: mon: 3 daemons, quorum a,b,c (age 9h) mgr: a(active, since 9h), standbys: b osd: 4 osds: 4 up (since 9h), 4 in (since 9h) rgw: 2 daemons active (2 hosts, 1 zones) data: pools: 15 pools, 409 pgs objects: 1.67k objects, 4.6 GiB usage: 11 GiB used, 2.1 TiB / 2.1 TiB avail pgs: 409 active+clean io: client: 85 B/s rd, 500 KiB/s wr, 0 op/s rd, 27 op/s wr
Verify Ceph health details:
ceph health detail
Example of a positive system response in the affected cluster:
HEALTH_WARN 2 OSD(s) experiencing slow operations in BlueStore [WRN] BLUESTORE_SLOW_OP_ALERT: 2 OSD(s) experiencing slow operations in BlueStore osd.2 observed slow operation indications in BlueStore osd.3 observed slow operation indications in BlueStore
Exit the
pelagia-ceph-toolspod.
Workaround:
Configure the bluestore_slow_ops_warn options as follows:
kubectl -n ceph-lcm-mirantis edit cephdeployment
spec:
cephClusterSpec:
rookConfig:
osd|bluestore_slow_ops_warn_lifetime: "600"
osd|bluestore_slow_ops_warn_threshold: "10"
Wait for up to five minutes for the change to apply and the alert to disappear during cluster deployment.
This configuration triggers the alert only if at least 10 BlueStore slow operations occur during last 10 minutes. If triggered, it indicates a potential hardware disk issue on the BlueStore host that must be verified and reconfigured accordingly.
[58609] Ceph rebalancing gets stuck during disabled node removal
When disabling or removing a Ceph node during operations such as a rolling
reboot, Ceph may not finish rebalancing if only two of three OSD nodes remain
active. The CephDeployment object can remain in Maintenance, causing
the rebalance process to wait indefinitely for Ceph to become ready. The issue
may only affect environments with a small number of Ceph OSD nodes, pool
replica count set to one less than the number of storage nodes
(replicas=storage_nodes_count-1), and failure domain host.
As a workaround, run the following command for the affected Ceph OSD node:
ceph osd reweight <osdId> 0
Cluster update
[8106] Frequent node disconnections with mcc-keepalived forcing new election
Note
Since MOSK 26.2, this issue is documented in Troubleshooting Guide: Frequent node disconnections with mcc-keepalived forcing new election.
The description and workaround below remain valid for this release.
After cluster update, some nodes may remain in an unstable Ready state
with mcc-keepalived constantly reelecting the leader and failing to
acquire the VIP address, which produces forcing new election messages in
logs.
Workaround:
Identify the leader node that owns the VIP:
On any control plane node, run the following command:
cat /etc/keepalived/keepalived.confIn the system response, capture the VIP used for the cluster.
Using the VIP, identify the leader node:
ip a| grep <VIP>
If the VIP is not found, run the command on another control plane node until you find the leader.
Connect to the non-leader control plane nodes and change the priority on these nodes in
keepalived.conf:vi /etc/keepalived/keepalived.confFor example, change the priority on each node to 150 and 200 respectively:
vrrp_instance VRRP1 { state MASTER garp_master_delay 15 interface k8s-lcm virtual_router_id 154 priority 100 # Change it on one node to 150 and on the other node to 200 virtual_ipaddress { 10.205.88.181 }Restart the
mcc-keepalivedservice on the control plane nodes where the priority was changed:systemctl restart mcc-keepalived
In 10-15 minutes, verify the logs of the node identified in step 1:
journalctl -u mcc-keepalived -f | grep election
You should no longer see the
forcing new electionmessages, and the flapping node status should be resolved.
[64721] MOSK cluster upgrade gets stuck due to runc-ee package downgrade
Clusters deployed with MOSK 26.1 (Cluster release 21.1.0) starting from
2026-08-04 may have issues upgrading to more recent versions, such as MOSK
26.1.x (Cluster releases 21.1.x) patch releases or MOSK 26.2 (Cluster release
21.2.0). The issue occurs during lcm-ansible execution at the step that
installs the runc-ee package.
Symptom: the first control plane node gets stuck in the deploy phase
with the following error in the lcm-ansible logs:
TASK [containerd : Install containerd.io packages]
task path: .../lcm-ansible-<version>/roles/containerd/tasks/Debian.yml:62
The following packages will be upgraded:
containerd.io* → 1.7.31m1+fips-0ubuntu0.24.04.1
The following packages will be DOWNGRADED:
runc-ee → 1.4.2m1-0ubuntu0.24.04.1
E: Packages were downgraded and -y was used without --allow-downgrades.
The issue occurs because the runc-ee package version was not pinned and was
installed from a constantly updated repository in MOSK 26.1 (Cluster release
21.1.0). Starting from MOSK 26.1.x (Cluster releases 21.1.x) patch releases and
MOSK 26.2 (Cluster release 21.2.0), the package version is pinned.
Workaround:
SSH to any node of the affected cluster.
Install the pinned version of the
runc-eepackage:apt-get --allow-downgrades install runc-ee=1.4.2m1-0ubuntu0.24.04.1 -y --allow-change-held-packages
Hold the package to prevent future automatic updates:
apt-mark hold runc-ee
Repeat steps 1-3 for the remaining nodes of the affected cluster.
LCM
[42889] Graceful reboot gets stuck when Kubernetes and OpenStack control planes are drained simultaneously
When a GracefulRebootRequest targets both the Kubernetes and OpenStack
control plane machines, either by listing machines of both types in
spec.machines or by leaving the list empty to reboot all cluster nodes,
the rolling reboot may get stuck. This happens because both node groups are
drained in parallel, and the OpenStack workload manager running on the
Kubernetes control plane becomes unavailable while the OpenStack control plane
nodes are simultaneously being drained.
Workaround:
Identify the machines that have not yet been rebooted.
Delete the stuck
GracefulRebootRequest:kubectl -n <projectName> delete gracefulrebootrequest <gracefulRebootRequestName>
Recreate the reboot requests in two sequential steps as described in Perform a rolling reboot of a cluster using CLI: first for the Kubernetes control plane machines, then, once that request completes and is deleted, for the remaining machines that still require a reboot.
MOSK management console
[50168] Inability to use a new project right after creation
A newly created project does not display all available tabs in the MOSK
management console and contains different access denied errors during first
five minutes after creation.
To work around the issue, refresh the browser in five minutes after the project creation.
OpenSDN
[40032] tf-rabbitmq fails to start after rolling reboot
Occasionally, RabbitMQ instances in tf-rabbitmq pods fail to enable
the tracking_records_in_ets during the initialization process.
To work around the issue, restart the affected pods manually.
[51101] tf-config pods fail to process API calls
Note
Since MOSK 26.2, this issue is documented in Troubleshooting Guide: The OpenSDN tf-config pods fail to process API calls.
The description and workaround below remain valid for this release.
The OpenSDN tf-config pods may fail to process API calls when the uWSGI
listen queue is full. As a result, pods report Unhealthy and OpenSDN
deployments can fail. In the pod logs, repeated messages appear such as:
*** uWSGI listen queue of socket "10.10.0.155:8082" (fd: 3) full !!!
(101/100) ***
Workaround:
Delete all tf-config pods one by one so they are recreated.
List the
tf-configpods:kubectl get pods -l tungstenfabric=config -n tf
Delete one
tf-configpod:kubectl delete pod <POD_NAME> -n tf
Wait for the new pod to be created.
Verify that the new pod has status
Runningand the restart count does not increase:kubectl get pods -l tungstenfabric=config -n tf
Example of a positive system response:
tf-config-jcfrr 4/4 Running 0 2m
Repeat steps 2-3 for the remaining
tf-configpods one by one.
OpenStack
[31186,34132] Pods get stuck during MariaDB operations
Note
Since MOSK 26.2, this issue is documented in Troubleshooting Guide: Pods get stuck during MariaDB operations.
The description and workaround below remain valid for this release.
During MariaDB operations on a management cluster, Pods may get stuck in continuous restarts with the following example error:
[ERROR] WSREP: Corrupt buffer header: \
addr: 0x7faec6f8e518, \
seqno: 3185219421952815104, \
size: 909455917, \
ctx: 0x557094f65038, \
flags: 11577. store: 49, \
type: 49
Workaround:
Create a backup of the
/var/lib/mysqldirectory on themariadb-serverPod.Verify that other replicas are up and ready.
Remove the
galera.cachefile for the affectedmariadb-serverPod.Remove the affected
mariadb-serverPod or wait until it is automatically restarted.
After Kubernetes restarts the Pod, the Pod clones the database in 1-2 minutes and restores the quorum.
[53401] Credential rotation reports success without performing action
Occasionally, the password rotation procedure for admin or service
credentials may incorrectly report success without actually initiating
the rotation process. This can result in unchanged credentials despite
the procedure indicating completion.
To work around the issue, restart the rotation procedure and verify that the credentials have been successfully updated.
[54570] The rfs-openstack-redis pod gets stuck in the Completed state
Note
Since MOSK 26.2, this issue is documented in Troubleshooting Guide: The rfs-openstack-redis pod gets stuck in the Completed state.
The description and workaround below remain valid for this release.
After node reboot, the rfs-openstack-redis pod may get stuck in the
Completed state blocking synchronization of the Redis cluster.
As a workaround, delete the rfs-openstack-redis pod that remains in the
Completed state:
kubectl -n openstack-redis delete <pod-name>
[57473] OpenStack update fails due to neutron-ovs-agent-default start failure
During OpenStack update from Caracal to Epoxy, neutron-ovs-agent-default
may fail to start with the The DaemonSet neutron-ovs-agent-default is not
ready error in the Rockoon logs due to Kopf missing the OsDpl update
events.
As a workaround, recreate the rockoon pod of the affected MOSK cluster:
kubectl -n osh-system rollout restart deployment rockoon
[63801] Encrypted ephemeral VM disk becomes undecryptable
Rebooting a virtual machine (VM) with encrypted ephemeral storage after upgrading from OpenStack Caracal to Epoxy makes its data undecryptable.
Symptoms:
After upgrade from OpenStack Caracal to Epoxy, a VM can no longer access existing data on its encrypted ephemeral disk after a hard reboot or a stop/start. This issue affects VMs that meet both of the following conditions:
A VM uses an encrypted ephemeral storage that was created on OpenStack Caracal or earlier
A VM was created on OpenStack Caracal, and OpenStack was later upgraded to Epoxy
Warning
The data is not lost and remains recoverable as long as no application inside the affected VM writes to the ephemeral disk, for example, by reformatting it or re-creating partitions.
Cause:
OpenStack Caracal container images in MOSK are based on Ubuntu 22.04.
OpenStack Epoxy images are based on Ubuntu 24.04, which includes
cryptsetup 2.7. This cryptsetup version changes the default
hashing algorithm for disk encryption from ripemd160 to sha256.
For details, see Cryptsetup 2.7.0 Release Notes.
OpenStack Nova does not explicitly set the hashing algorithm, as per upstream known issue #1639221, when it configures disk encryption. As a result, after the upgrade to Epoxy, Nova sets up encryption with different parameters than those used when the disk was originally created, which breaks decryption. MOSK does not automatically re-encrypt affected disks with the original hashing algorithm.
Workaround:
The fix for this issue is included in the nova image delivered with
MOSK 26.2 for OpenStack Epoxy. The image includes fixes that let you
control the hashing algorithm that Nova uses for disk encryption:
Nova now explicitly specifies the hashing algorithm when it creates disk encryption. On Epoxy, the default is
sha256, matching thecryptsetupdefault in Epoxy-based MOSK images. This default applies to all VMs on a given compute node. You can override it in the configuration of anynova-computeservice.You can also override the hashing algorithm for an individual VM using the
dmcrypt_hashinstance metadata key.
To fix VMs that are already broken or at risk of breaking after the upgrade to Epoxy:
Pin the MOSK 26.2
novaimage for thenova-computeservice using the<OPENSTACKDEPLOYMENT-NAME>-artifactsConfigMap:apiVersion: v1 kind: ConfigMap metadata: labels: openstack.lcm.mirantis.com/watch: "true" name: <OPENSTACKDEPLOYMENT-NAME>-artifacts namespace: openstack data: epoxy: | nova_compute: mirantis.azurecr.io/openstack/nova:epoxy-noble-20260829052712
Caution
Remove the pinned image from the ConfigMap only after you update your cluster to MOSK 26.2 or newer. Otherwise, restarting an existing virtual machine breaks its ephemeral disk the same way.
Add the following metadata to each affected instance:
openstack server set --property dmcrypt_hash=ripemd160 <INSTANCE-ID>
If the instance is already broken (can not read its ephemeral disk), power-cycle the instance.
The disk becomes decryptable again.
To prevent existing VMs from breaking before you upgrade from Caracal to Epoxy:
While still on Caracal, pin the MOSK 26.2
novaimage for thenova-computeservice using the<OPENSTACKDEPLOYMENT-NAME>-artifactsConfigMap. The image is pinned per OpenStack release and takes effect when you upgrade to Epoxy:apiVersion: v1 kind: ConfigMap metadata: labels: openstack.lcm.mirantis.com/watch: "true" name: <OPENSTACKDEPLOYMENT-NAME>-artifacts namespace: openstack data: epoxy: | nova_compute: mirantis.azurecr.io/openstack/nova:epoxy-noble-20260829052712
Caution
Remove the pinned image from the ConfigMap only after you update your cluster to MOSK 26.2 or newer. Otherwise, restarting an existing VM breaks its ephemeral disk the same way.
While still on Caracal, set the following configuration option to
ripemd160. The setting has no effect until you upgrade to Epoxy.kind: OpenStackDeployment spec: services: compute: nova: values: conf: nova: ephemeral_storage_encryption: hash: ripemd160
Upgrade to Epoxy.
Existing VMs continue to work without requiring per-instance
metadata. New VMs also default to ripemd160, which is a less
secure hashing algorithm than the sha256 default introduced in
Epoxy.
Security
[58728] The managed: false field is added for auditd after cluster update
After update of a management cluster to 2.31.0, the managed: false field
is added to the auditd configuration in the Cluster object of MOSK
clusters that have auditd enabled. This behaviour is expected and does not
affect the auditd functionality. Therefore, no action is required before MOSK
cluster update to 26.1 or 26.1.x.
For release changes in the auditd configuration and actions required after the MOSK cluster update to 26.1 or 26.1.x, see Migration of the auditd settings from the Cluster object to the auditd module.
StackLight
[48581] OpenSearchClusterStatusCritical is firing during cluster update
During update of a management or MOSK cluster with StackLight enabled in HA
mode, the OpenSearchClusterStatusCritical alert may trigger when the next
OpenSearch node restarts before shards from the previous node finish assigning.
This can push some indices to red temporarily, making them unavailable for
reads and writes, possibly causing some new logs being lost.
The issue does not affect the cluster during the update, no workaround is needed, and you can safely ignore it.
[55317] The Dropped sample for series errors in the Prometheus logs
When the experimental feature memory-snapshot-on-shutdown, which is enabled
by default, is used together with remote_write, Prometheus may emit
multiple log messages, such as Dropped sample for series that was not
explicitly dropped via relabelling. For more details, see the upstream issue
description in the Prometheus GitHub project.
Workaround:
On the related management cluster, open the affected MOSK
Clusterobject for editing:kubectl edit cluster <affectedMOSKClusterName> -n <affectedMOSKClusterProjectName>
Remove the
memory-snapshot-on-shutdownfeature from theprometheusServer:enabledFeatureslist:spec: ... providerSpec: ... value: ... helmReleases: ... - name: stacklight values: ... prometheusServer: enabledFeatures: []