Newer documentation is now live.You are currently reading an older version.

Known issues

This section lists MOSK known issues with workarounds for the MOSK release 26.1.1.

Bare metal

[63019] Management cluster upgrade gets stuck with a machine in prepare phase

Fixed in MOSK 26.1.2

Due to an issue with the AnsibleExtra configuration, a management cluster upgrade may get stuck with one of the machines remaining in the prepare phase with the following error message in the baremetal-provider logs :

Error updating machine \"default/master-2\":
failed to build AnsibleExtra Spec for Machine 'default/master-2'
from BareMetalHost 'default/master-2' HardwareDetails matching
BareMetalHostProfile 'default/master-2-profile':
rootFS 'http://httpd-http/distribution/ubuntu/noble/24.04~20260615122024/rootfs' is not yet available ...

You may also see the following error in the logs of the resource-controller container of the Ironic pod:

files_linux.go:68] Failed to ensure resource by-path
'/volume/html/distribution/ubuntu/jammy/22.04~20260615122024/rootfs':
download, create directory error: mkdir /volume/html/distribution: permission denied

As a workaround, delete the Ironic pod in the kaas namespace. A new pod will be spawned and missing artifacts for Ubuntu distributions will be downloaded.

[62999] Master node replacement fails on MOSK clusters with BGP announcement enabled

On MOSK clusters with BGP announcement enabled (useBGPAnnouncement: true in the Cluster object), master nodes are deployed one per rack, and the Kubernetes API Virtual IP (VIP), for example, 10.0.30.100/32 is assigned to the loopback (lo) interface of each master node. For details on BGP announcement, see Configure BGP announcement for cluster API LB address.

During master node replacement, lcm-agent on the new node tries to reach the API VIP on the lo interface for the first time but cannot connect to it. This happens because cloud-init binds the API VIP to the lo interface before the local containerized API proxy starts listening on it, so the connection is refused. The lcm-agent cannot report back to the management cluster. As a result, the Machine object remains stuck in the PendingLCMAgent state.

As a workaround, temporarily remove the API VIP from the lo interface to force API traffic through the physical network to a healthy master node. Once the node finishes provisioning, the local VIP handling takes over automatically.

Workaround:

  1. From the management cluster, obtain the IP address of the new node from its IpamHost object:

    kubectl get ipamhost <NODE_NAME> -o jsonpath="{.status.serviceMap['ipam/SVC-k8s-lcm'][0].ipAddress}"
    
  2. SSH to the new node:

    ssh -i <path-to-ssh-key> mcc-user@<NEW_NODE_IP>
    
  3. Wait for cloud-init to complete on the new node:

    sudo cloud-init status --wait
    

    Before proceeding, ensure that the output reports status: done.

    Warning

    Do not proceed with networking modifications until cloud-init has finished executing, so it does not overwrite your changes.

  4. Verify that the API VIP is bound to the lo interface:

    ip addr show dev lo
    

    Example of system response on the affected node:

    1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1000
        inet 127.0.0.1/8 scope host lo
        inet 10.0.30.100/32 scope global lo
    
  5. Confirm that lcm-agent is failing with a connection refused error:

    sudo journalctl -u lcm-agent* -n 20 --no-pager
    
  6. Remove the /32 API VIP from the lo interface:

    sudo ip addr del 10.0.30.100/32 dev lo
    

    Substitute 10.0.30.100 with the actual API VIP of your cluster.

  7. Verify outbound network connectivity to the API. It should now route across the network to one of the healthy master nodes:

    curl -k -v https://10.0.30.100:443/api
    

    An HTTP response such as HTTP/2 401 or 403 confirms that the traffic is reaching an active Kubernetes API server. For example:

    * Connected to 10.0.30.100 (10.0.30.100) port 443
    < HTTP/2 401
    {
      "kind": "Status",
      "message": "Unauthorized",
      "code": 401
    }
    
  8. Reload systemd to clear the warnings about unit files that have changed on disk:

    sudo systemctl daemon-reload
    
  9. Restart the lcm-agent daemon to trigger an immediate check-in:

    sudo systemctl restart lcm-agent*.service
    
  10. Verify that the node recovers:

    1. On the new node, inspect the agent logs to confirm a successful tick execution:

      sudo journalctl -u lcm-agent* -f
      

      In the system response, find the tick finished messages without connection errors.

    2. From the management or seed node, verify the machine phase transition:

      kubectl get machine,lcmmachine -o wide
      

      The LCMPHASE for the new node should transition from PendingLCMAgent to Prepare, Deploy, and then to Ready.

Ceph

[54195] Ceph OSD experiencing slow operations in BlueStore during MOSK deployment

Note

Since MOSK 26.2, this issue is documented in Troubleshooting Guide: Ceph OSD experiencing slow operations in BlueStore.

The description and workaround below remain valid for this release.

During MOSK cluster deployment, the following false-positive example alert for Ceph may raise:

Failed to configure Ceph cluster: ceph cluster verification is failed:
[BLUESTORE_SLOW_OP_ALERT: 3 OSD(s) experiencing slow operations in BlueStore]

The issue occurs due to the following upstream Ceph issues:

To verify whether the cluster is affected:

  1. Enter the pelagia-ceph-tools pod:

  2. Verify the Ceph cluster status:

    • Verify Ceph health:

      ceph -s
      

      Example of a positive system response in the affected cluster:

      cluster:
        id:     6ae41eb3-262e-4da9-8847-25efed2fcaa2
        health: HEALTH_WARN
                2 OSD(s) experiencing slow operations in BlueStore
      
      services:
        mon: 3 daemons, quorum a,b,c (age 9h)
        mgr: a(active, since 9h), standbys: b
        osd: 4 osds: 4 up (since 9h), 4 in (since 9h)
        rgw: 2 daemons active (2 hosts, 1 zones)
      
      data:
        pools:   15 pools, 409 pgs
        objects: 1.67k objects, 4.6 GiB
        usage:   11 GiB used, 2.1 TiB / 2.1 TiB avail
        pgs:     409 active+clean
      
      io:
        client:   85 B/s rd, 500 KiB/s wr, 0 op/s rd, 27 op/s wr
      
    • Verify Ceph health details:

      ceph health detail
      

      Example of a positive system response in the affected cluster:

      HEALTH_WARN 2 OSD(s) experiencing slow operations in BlueStore
      [WRN] BLUESTORE_SLOW_OP_ALERT: 2 OSD(s) experiencing slow operations in BlueStore
           osd.2 observed slow operation indications in BlueStore
           osd.3 observed slow operation indications in BlueStore
      
  3. Exit the pelagia-ceph-tools pod.

Workaround:

Configure the bluestore_slow_ops_warn options as follows:

kubectl -n ceph-lcm-mirantis edit cephdeployment
spec:
  cephClusterSpec:
    rookConfig:
      osd|bluestore_slow_ops_warn_lifetime: "600"
      osd|bluestore_slow_ops_warn_threshold: "10"

Wait for up to five minutes for the change to apply and the alert to disappear during cluster deployment.

This configuration triggers the alert only if at least 10 BlueStore slow operations occur during last 10 minutes. If triggered, it indicates a potential hardware disk issue on the BlueStore host that must be verified and reconfigured accordingly.

[58609] Ceph rebalancing gets stuck during disabled node removal

When disabling or removing a Ceph node during operations such as a rolling reboot, Ceph may not finish rebalancing if only two of three OSD nodes remain active. The CephDeployment object can remain in Maintenance, causing the rebalance process to wait indefinitely for Ceph to become ready. The issue may only affect environments with a small number of Ceph OSD nodes, pool replica count set to one less than the number of storage nodes (replicas=storage_nodes_count-1), and failure domain host.

As a workaround, run the following command for the affected Ceph OSD node:

ceph osd reweight <osdId> 0

Cluster update

[8106] Frequent node disconnections with mcc-keepalived forcing new election

Note

Since MOSK 26.2, this issue is documented in Troubleshooting Guide: Frequent node disconnections with mcc-keepalived forcing new election.

The description and workaround below remain valid for this release.

After cluster update, some nodes may remain in an unstable Ready state with mcc-keepalived constantly reelecting the leader and failing to acquire the VIP address, which produces forcing new election messages in logs.

Workaround:

  1. Identify the leader node that owns the VIP:

    1. On any control plane node, run the following command:

      cat /etc/keepalived/keepalived.conf
      

      In the system response, capture the VIP used for the cluster.

    2. Using the VIP, identify the leader node:

      ip a| grep <VIP>
      

      If the VIP is not found, run the command on another control plane node until you find the leader.

  2. Connect to the non-leader control plane nodes and change the priority on these nodes in keepalived.conf:

    vi /etc/keepalived/keepalived.conf
    

    For example, change the priority on each node to 150 and 200 respectively:

    vrrp_instance VRRP1 {
        state MASTER
        garp_master_delay 15
        interface k8s-lcm
        virtual_router_id 154
        priority 100        # Change it on one node to 150 and on the other node to 200
        virtual_ipaddress {
            10.205.88.181
        }
    
  3. Restart the mcc-keepalived service on the control plane nodes where the priority was changed:

    systemctl restart mcc-keepalived
    
  4. In 10-15 minutes, verify the logs of the node identified in step 1:

    journalctl -u mcc-keepalived -f | grep election
    

    You should no longer see the forcing new election messages, and the flapping node status should be resolved.

[64721] MOSK cluster upgrade gets stuck due to runc-ee package downgrade

Clusters deployed with MOSK 26.1 (Cluster release 21.1.0) starting from 2026-08-04 may have issues upgrading to more recent versions, such as MOSK 26.1.x (Cluster releases 21.1.x) patch releases or MOSK 26.2 (Cluster release 21.2.0). The issue occurs during lcm-ansible execution at the step that installs the runc-ee package.

Symptom: the first control plane node gets stuck in the deploy phase with the following error in the lcm-ansible logs:

TASK [containerd : Install containerd.io packages]
task path: .../lcm-ansible-<version>/roles/containerd/tasks/Debian.yml:62
The following packages will be upgraded:
  containerd.io* → 1.7.31m1+fips-0ubuntu0.24.04.1
The following packages will be DOWNGRADED:
  runc-ee → 1.4.2m1-0ubuntu0.24.04.1
E: Packages were downgraded and -y was used without --allow-downgrades.

The issue occurs because the runc-ee package version was not pinned and was installed from a constantly updated repository in MOSK 26.1 (Cluster release 21.1.0). Starting from MOSK 26.1.x (Cluster releases 21.1.x) patch releases and MOSK 26.2 (Cluster release 21.2.0), the package version is pinned.

Workaround:

  1. SSH to any node of the affected cluster.

  2. Install the pinned version of the runc-ee package:

    apt-get --allow-downgrades install runc-ee=1.4.2m1-0ubuntu0.24.04.1 -y --allow-change-held-packages
    
  3. Hold the package to prevent future automatic updates:

    apt-mark hold runc-ee
    
  4. Repeat steps 1-3 for the remaining nodes of the affected cluster.

LCM

[42889] Graceful reboot gets stuck when Kubernetes and OpenStack control planes are drained simultaneously

When a GracefulRebootRequest targets both the Kubernetes and OpenStack control plane machines, either by listing machines of both types in spec.machines or by leaving the list empty to reboot all cluster nodes, the rolling reboot may get stuck. This happens because both node groups are drained in parallel, and the OpenStack workload manager running on the Kubernetes control plane becomes unavailable while the OpenStack control plane nodes are simultaneously being drained.

Workaround:

  1. Identify the machines that have not yet been rebooted.

  2. Delete the stuck GracefulRebootRequest:

    kubectl -n <projectName> delete gracefulrebootrequest <gracefulRebootRequestName>
    
  3. Recreate the reboot requests in two sequential steps as described in Perform a rolling reboot of a cluster using CLI: first for the Kubernetes control plane machines, then, once that request completes and is deleted, for the remaining machines that still require a reboot.

MOSK management console

[50168] Inability to use a new project right after creation

A newly created project does not display all available tabs in the MOSK management console and contains different access denied errors during first five minutes after creation.

To work around the issue, refresh the browser in five minutes after the project creation.

OpenSDN

[40032] tf-rabbitmq fails to start after rolling reboot

Occasionally, RabbitMQ instances in tf-rabbitmq pods fail to enable the tracking_records_in_ets during the initialization process.

To work around the issue, restart the affected pods manually.

[51101] tf-config pods fail to process API calls

Note

Since MOSK 26.2, this issue is documented in Troubleshooting Guide: The OpenSDN tf-config pods fail to process API calls.

The description and workaround below remain valid for this release.

The OpenSDN tf-config pods may fail to process API calls when the uWSGI listen queue is full. As a result, pods report Unhealthy and OpenSDN deployments can fail. In the pod logs, repeated messages appear such as:

*** uWSGI listen queue of socket "10.10.0.155:8082" (fd: 3) full !!!
(101/100) ***

Workaround:

Delete all tf-config pods one by one so they are recreated.

  1. List the tf-config pods:

    kubectl get pods -l tungstenfabric=config -n tf
    
  2. Delete one tf-config pod:

    kubectl delete pod <POD_NAME> -n tf
    

    Wait for the new pod to be created.

  3. Verify that the new pod has status Running and the restart count does not increase:

    kubectl get pods -l tungstenfabric=config -n tf
    

    Example of a positive system response:

    tf-config-jcfrr     4/4     Running     0     2m
    
  4. Repeat steps 2-3 for the remaining tf-config pods one by one.

OpenStack

[31186,34132] Pods get stuck during MariaDB operations

Note

Since MOSK 26.2, this issue is documented in Troubleshooting Guide: Pods get stuck during MariaDB operations.

The description and workaround below remain valid for this release.

During MariaDB operations on a management cluster, Pods may get stuck in continuous restarts with the following example error:

[ERROR] WSREP: Corrupt buffer header: \
addr: 0x7faec6f8e518, \
seqno: 3185219421952815104, \
size: 909455917, \
ctx: 0x557094f65038, \
flags: 11577. store: 49, \
type: 49

Workaround:

  1. Create a backup of the /var/lib/mysql directory on the mariadb-server Pod.

  2. Verify that other replicas are up and ready.

  3. Remove the galera.cache file for the affected mariadb-server Pod.

  4. Remove the affected mariadb-server Pod or wait until it is automatically restarted.

After Kubernetes restarts the Pod, the Pod clones the database in 1-2 minutes and restores the quorum.

[53401] Credential rotation reports success without performing action

Occasionally, the password rotation procedure for admin or service credentials may incorrectly report success without actually initiating the rotation process. This can result in unchanged credentials despite the procedure indicating completion.

To work around the issue, restart the rotation procedure and verify that the credentials have been successfully updated.

[54570] The rfs-openstack-redis pod gets stuck in the Completed state

Note

Since MOSK 26.2, this issue is documented in Troubleshooting Guide: The rfs-openstack-redis pod gets stuck in the Completed state.

The description and workaround below remain valid for this release.

After node reboot, the rfs-openstack-redis pod may get stuck in the Completed state blocking synchronization of the Redis cluster.

As a workaround, delete the rfs-openstack-redis pod that remains in the Completed state:

kubectl -n openstack-redis delete <pod-name>

[57473] OpenStack update fails due to neutron-ovs-agent-default start failure

During OpenStack update from Caracal to Epoxy, neutron-ovs-agent-default may fail to start with the The DaemonSet neutron-ovs-agent-default is not ready error in the Rockoon logs due to Kopf missing the OsDpl update events.

As a workaround, recreate the rockoon pod of the affected MOSK cluster:

kubectl -n osh-system rollout restart deployment rockoon

[63801] Encrypted ephemeral VM disk becomes undecryptable

Rebooting a virtual machine (VM) with encrypted ephemeral storage after upgrading from OpenStack Caracal to Epoxy makes its data undecryptable.

Symptoms:

After upgrade from OpenStack Caracal to Epoxy, a VM can no longer access existing data on its encrypted ephemeral disk after a hard reboot or a stop/start. This issue affects VMs that meet both of the following conditions:

  • A VM uses an encrypted ephemeral storage that was created on OpenStack Caracal or earlier

  • A VM was created on OpenStack Caracal, and OpenStack was later upgraded to Epoxy

Warning

The data is not lost and remains recoverable as long as no application inside the affected VM writes to the ephemeral disk, for example, by reformatting it or re-creating partitions.

Cause:

OpenStack Caracal container images in MOSK are based on Ubuntu 22.04. OpenStack Epoxy images are based on Ubuntu 24.04, which includes cryptsetup 2.7. This cryptsetup version changes the default hashing algorithm for disk encryption from ripemd160 to sha256. For details, see Cryptsetup 2.7.0 Release Notes.

OpenStack Nova does not explicitly set the hashing algorithm, as per upstream known issue #1639221, when it configures disk encryption. As a result, after the upgrade to Epoxy, Nova sets up encryption with different parameters than those used when the disk was originally created, which breaks decryption. MOSK does not automatically re-encrypt affected disks with the original hashing algorithm.

Workaround:

The fix for this issue is included in the nova image delivered with MOSK 26.2 for OpenStack Epoxy. The image includes fixes that let you control the hashing algorithm that Nova uses for disk encryption:

  • Nova now explicitly specifies the hashing algorithm when it creates disk encryption. On Epoxy, the default is sha256, matching the cryptsetup default in Epoxy-based MOSK images. This default applies to all VMs on a given compute node. You can override it in the configuration of any nova-compute service.

  • You can also override the hashing algorithm for an individual VM using the dmcrypt_hash instance metadata key.

To fix VMs that are already broken or at risk of breaking after the upgrade to Epoxy:

  1. Pin the MOSK 26.2 nova image for the nova-compute service using the <OPENSTACKDEPLOYMENT-NAME>-artifacts ConfigMap:

    apiVersion: v1
    kind: ConfigMap
    metadata:
      labels:
        openstack.lcm.mirantis.com/watch: "true"
      name: <OPENSTACKDEPLOYMENT-NAME>-artifacts
      namespace: openstack
    data:
      epoxy: |
        nova_compute: mirantis.azurecr.io/openstack/nova:epoxy-noble-20260829052712
    

    Caution

    Remove the pinned image from the ConfigMap only after you update your cluster to MOSK 26.2 or newer. Otherwise, restarting an existing virtual machine breaks its ephemeral disk the same way.

  2. Add the following metadata to each affected instance:

    openstack server set --property dmcrypt_hash=ripemd160 <INSTANCE-ID>
    
  3. If the instance is already broken (can not read its ephemeral disk), power-cycle the instance.

The disk becomes decryptable again.

To prevent existing VMs from breaking before you upgrade from Caracal to Epoxy:

  1. While still on Caracal, pin the MOSK 26.2 nova image for the nova-compute service using the <OPENSTACKDEPLOYMENT-NAME>-artifacts ConfigMap. The image is pinned per OpenStack release and takes effect when you upgrade to Epoxy:

    apiVersion: v1
    kind: ConfigMap
    metadata:
      labels:
        openstack.lcm.mirantis.com/watch: "true"
      name: <OPENSTACKDEPLOYMENT-NAME>-artifacts
      namespace: openstack
    data:
      epoxy: |
        nova_compute: mirantis.azurecr.io/openstack/nova:epoxy-noble-20260829052712
    

    Caution

    Remove the pinned image from the ConfigMap only after you update your cluster to MOSK 26.2 or newer. Otherwise, restarting an existing VM breaks its ephemeral disk the same way.

  2. While still on Caracal, set the following configuration option to ripemd160. The setting has no effect until you upgrade to Epoxy.

    kind: OpenStackDeployment
    spec:
      services:
        compute:
          nova:
            values:
              conf:
                nova:
                  ephemeral_storage_encryption:
                    hash: ripemd160
    
  3. Upgrade to Epoxy.

Existing VMs continue to work without requiring per-instance metadata. New VMs also default to ripemd160, which is a less secure hashing algorithm than the sha256 default introduced in Epoxy.

Security

[58728] The managed: false field is added for auditd after cluster update

After update of a management cluster to 2.31.0, the managed: false field is added to the auditd configuration in the Cluster object of MOSK clusters that have auditd enabled. This behaviour is expected and does not affect the auditd functionality. Therefore, no action is required before MOSK cluster update to 26.1 or 26.1.x.

For release changes in the auditd configuration and actions required after the MOSK cluster update to 26.1 or 26.1.x, see Migration of the auditd settings from the Cluster object to the auditd module.

StackLight

[48581] OpenSearchClusterStatusCritical is firing during cluster update

Fixed in MOSK 26.2

During update of a management or MOSK cluster with StackLight enabled in HA mode, the OpenSearchClusterStatusCritical alert may trigger when the next OpenSearch node restarts before shards from the previous node finish assigning. This can push some indices to red temporarily, making them unavailable for reads and writes, possibly causing some new logs being lost.

The issue does not affect the cluster during the update, no workaround is needed, and you can safely ignore it.

[55317] The Dropped sample for series errors in the Prometheus logs

When the experimental feature memory-snapshot-on-shutdown, which is enabled by default, is used together with remote_write, Prometheus may emit multiple log messages, such as Dropped sample for series that was not explicitly dropped via relabelling. For more details, see the upstream issue description in the Prometheus GitHub project.

Workaround:

  1. On the related management cluster, open the affected MOSK Cluster object for editing:

    kubectl edit cluster <affectedMOSKClusterName> -n <affectedMOSKClusterProjectName>
    
  2. Remove the memory-snapshot-on-shutdown feature from the prometheusServer:enabledFeatures list:

    spec:
      ...
      providerSpec:
        ...
        value:
          ...
          helmReleases:
            ...
            - name: stacklight
              values:
                ...
                prometheusServer:
                  enabledFeatures: []