26. Release Notes

26.1. Robin Cloud Native Platform v5.9.0

The Robin Cloud Native Platform (CNP) v5.9.0 release notes document has pre- and post-upgrade considerations, new features, improvements, fixed issues, and known issues.

Release Date: September 30, 2026

26.1.1. Infrastructure Versions

The following software applications are included in this CNP release:

Software Application

Version

Kubernetes

1.35.5

Docker

25.0.2 (RHEL 8.10 or Rocky Linux 8.10)

Podman

5.4.0 (RHEL 9.6)

Prometheus

3.13.1

Prometheus Adapter

0.10.0

Node Exporter

1.8.0

Calico

3.31.6

HAProxy

3.4.2

PostgreSQL

15.17

Grafana

11.6.3

CRI Tools

1.35.0

cert-manager

1.19.1

26.1.2. Supported Operating Systems

The following are the supported operating systems and kernel versions for Robin CNP v5.9.0:

OS Version

Kernel Version

Red Hat Enterprise Linux 8.10

4.18.0-553.el8_10.x86_64

Rocky Linux 8.10

4.18.0-553.el8_10.x86_64

Red Hat Enterprise Linux 9.6

5.14.0-570.17.1.el9_6.x86_64

Note

  • Robin CNP supports both RT and non-RT kernels on the above supported operating systems.

  • Python 3 is mandatory for Robin CNP v5.9.0 and later. When installing or upgrading to Robin CNP v5.9.0 or later on Rocky Linux 8 or RHEL 8, Robin CNP automatically checks for Python 3 and installs it if it is not available. Python 3 is the default Python version on RHEL 9 and must be available before installing Robin CNP v5.9.0; if it is missing, the installation will fail. Robin CNP does not remove any existing Python versions, and both Python 2 and Python 3 can coexist on the same OS.

26.1.3. Upgrade Paths

The following are the supported upgrade paths for Robin CNP v5.9.0:

  • Robin CNP v5.7.1-311 (Refresh) to Robin CNP v5.9.0-862

  • Robin CNP v5.7.2-330 to Robin CNP v5.9.0-862

  • Robin CNP v5.7.2-330 (Refresh) to Robin CNP v5.9.0-862

26.1.3.1. Pre-upgrade considerations

  • For a successful upgrade, you must run the possible_job_stuck.py script before and after the upgrade. Contact the Robin Support team for the upgrade procedure using the script.

  • When upgrading from supported Robin CNP versions to Robin CNP v5.9.0, if your cluster already has cert-manager installed, you must uninstall it before upgrading to Robin CNP v5.9.0.

  • Before upgrading to Robin CNP v5.9.0, if the robin-certs-check job or CronJob is running, you must stop it. To stop the robin-certs-check job, run the kubectl delete job robin-certs-check -n robinio command, and to stop the robin-certs-check CronJob, run the robin cert check --stop-cronjob command.

26.1.3.2. Post-upgrade considerations

  • After upgrading to Robin CNP v5.9.0, verify that the value of the k8s_resource_sync config parameter is set to 60000 using the robin schedule list | grep -i K8sResSync command. If it is not set, you must run the robin schedule update K8sResSync k8s_resource_sync 60000 command to update the value of the robin schedule K8sResSync config parameter.

  • After upgrading to Robin CNP v5.9.0, you must run the robin-server validate-role-bindings command. To run this command, you need to log in to the robin-master Pod. This command verifies the roles assigned to each user in the cluster and corrects them if necessary.

  • After upgrading to Robin CNP v5.9.0, the k8s_auto_registration config parameter is disabled by default. The config setting is deactivated to prevent all Kubernetes apps from automatically registering and consuming resources. The following are the points you must be aware of with this change:

    • You can register the Kubernetes apps using the robin app register command manually and use Robin CNP for snapshots, clones, and backup operations of the Kubernetes app.

    • As this config parameter is disabled, when you run the robin app nfs-list command, the mappings between Kubernetes apps and NFS server Pods are not listed in the command output.

    • If you need mapping between a Kubernetes app and an NFS server Pod when the k8s_auto_registration config parameter is disabled or the k8s app is not manually registered, get the PVC name from the Pod YAML file (kubectl get pod -n <name> -o YAML) and run the robin nfs export list | grep <pvc name> command.

    • The robin nfs export list command output displays the PVC name and namespace.

  • After upgrading to Robin CNP v5.9.0, you must start the robin-certs-check CronJob using the robin cert check -start-cronjob command.

26.1.3.3. Pre-upgrade steps

Upgrading from the supported Robin CNP v5.7.1 or Robin CNP v5.7.2 to Robin CNP v5.9.0

Before upgrading from Robin CNP v5.7.1 or Robin CNP v5.7.2 to Robin CNP v5.9.0, perform the following steps:

  1. Update the value of the suicide_threshold config parameter to 1800:

    # robin config update agent suicide_threshold 1800
    
  2. Set the toleration seconds for all NFS server Pods to 86400 seconds. After the upgrade, you must change the toleration seconds according to the post-upgrade steps.

    for pod in kubectl get pod -n robinio -l robin.io/instance=robin-nfs --output=jsonpath={.items..metadata.name}; do echo "Updating $pod tolerationseconds to 86400";     kubectl patch pod $pod -n robinio --type='json' -p='[{"op": "replace", "path": "/spec/tolerations/1/tolerationSeconds", "value": 86400}, {"op": "replace", "path": "/spec/tolerations/2/tolerationSeconds", "value": 86400}]'; done
    
  3. Verify the webhooks are enabled by running the robin config list | grep -I robin_k8s_extension command. It should be true. If it is disabled, then enable it:

    # robin config update manager robin_k8s_extension True
    

26.1.3.4. Post-upgrade steps

Upgrading from the supported Robin CNP v5.7.1 or Robin CNP v5.7.2 to Robin CNP v5.9.0

After upgrading from Robin CNP v5.7.1 or Robin CNP v5.7.2 to Robin CNP v5.9.0, perform the following steps:

  1. Verify the state of the hosts, Pods, services, apps, and nodes.

  2. Update the value of the suicide_threshold config parameter to 40:

    # robin config update agent suicide_threshold 40
    
  3. Set the check_helm_apps config parameter to False:

    # robin config update cluster check_helm_apps False
    
  4. Verify the robin_k8s_extension config parameter is set to True. If not, set it to True.

    # robin config update manager robin_k8s_extension True
    
  5. Set the toleration seconds for all NFS server Pods to 60 seconds when the node is in the notready state and set it to 0 seconds when the node is in the unreachable state.

    # for pod in `kubectl get pod -n robinio -l robin.io/instance=robin-nfs --output=jsonpath={.items..metadata.name}`; do     echo "Updating $pod tolerationseconds";     kubectl patch pod $pod -n robinio --type='json' -p='[{"op": "replace", "path": "/spec/tolerations/0/tolerationSeconds", "value": 60}, {"op": "replace", "path": "/spec/tolerations/1/tolerationSeconds", "value": 0}]'; done 2>/dev/null
    

26.1.4. New Features

26.1.4.1. NetworkManager support for RHEL 9.6

Starting with Robin CNP v5.9.0, Robin CNP supports NetworkManager through the netmgr-wrapper for Red Hat Enterprise Linux (RHEL) 9.6 nodes. NetworkManager replaces the legacy network scripts (network.service daemon) used in previous RHEL OS versions.

NetworkManager is responsible for managing, configuring, and detecting network devices and connections dynamically. For more information, see NetworkManager.

Important points

  • When you install Robin CNP on RHEL 9.6 nodes, it uses the NetworkManager daemon.

  • When you add a new RHEL 9.6 node to your Robin CNP cluster, it uses the NetworkManager daemon for the newly added node only.

  • When you upgrade your Robin CNP cluster to Robin CNP v5.9.0, it uses the network.service daemon. It does not change to the NetworkManager daemon.

  • The legacy network-scripts method is deprecated for fresh installations with the Robin CNP v5.9.0 release and will no longer be supported beginning with release 5.10.0 for fresh installations.

26.1.4.2. Built-in HAProxy Prometheus exporter

Starting with Robin CNP v5.9.0, the HAProxy binary in Robin CNP includes a built-in Prometheus exporter module. This module, compiled with the USE_PROMEX flag, provides HAProxy metrics in Prometheus format through an HTTP /metrics endpoint on all master nodes. This feature eliminates the need for manual builds or configurations.

Default behavior

The exporter has the following default characteristics:

  • Provides metrics on port 29470 on all master nodes.

  • The port is configurable using a port file or an environment variable.

To verify the exporter, run the following command:

curl -Ss localhost:29470/metrics | head -5

Upgrade behavior

When you upgrade from a Robin CNP version earlier than Robin CNP v5.9.0, the upgrade process automatically performs the following actions:

  • Adds the HAPROXY_PROMETHEUS_PORT=29470 entry to the robin-bootstrap-config ConfigMap.

  • Updates the keepalived.conf.template file.

  • Enables the /metrics endpoint on port 29470 across all master nodes.

For more information, see HAProxy Prometheus exporter.

26.1.4.3. Support for Patroni log rotation

Starting with Robin CNP v5.9.0, you can manage PostgreSQL log rotation for Patroni directly through Robin configuration commands. This feature introduces the log_rotation_age attribute to control rotation frequency.

The following are the supported frequencies:

  • Minutes (min)

  • Hours (h)

  • Days (d)

  • The log_filename attribute to define strftime naming patterns for archived files. The default rotation age is 30d, and you can disable rotation by setting the value to 0.

26.1.4.4. Support for rotating Robin master key using HashiCorp Vault

Robin CNP v5.9.0 supports the rotation of the robin-master key using HashiCorp Vault as an external Key Management Service (KMS). The robin-master key, which is generated during installation and stored in the Vault KV store, serves as a Key Encryption Key (KEK) to encrypt and decrypt node and volume-level encryption keys. This rotation feature enhances security by replacing the existing master key with a new one and automatically re-encrypting all dependent key references (keyrefs) in the database, including host, encrypted volume, and backup keyrefs.

To rotate the robin-master key, run the following command:

# robin kms rotate-key

For more information, see Rotate Robin master key using Vault.

26.1.4.5. Modify the replica count for volumes

Robin CNP v5.9.0 enables you to modify the replica count of existing volumes, including clones and imported non-RWX volumes, by using the robin volume add-replica and robin volume remove-replica CLI commands.

This feature allows you to increase or decrease replica count one at a time. You must increase from one to two, two to three, and decrease from three to two, two to one.

Robin CNP automatically manages replica placement based on node health, available space, and defined fault domains. For more information, see Modify replica count for volumes.

26.1.4.6. Kubernetes control-plane certificate management

Robin CNP v5.9.0 provides a new feature to manage the lifecycle of Kubernetes control-plane certificates. This feature allows you to monitor certificate health, configure automated renewal policies, and perform manual renewals for certificates located in /etc/kubernetes/pki/ on control-plane nodes.

  • Use the robin k8s-cert list command to view the status of all control-plane certificates across the cluster, including expiration dates, remaining validity, and issuer details.

  • Configure the robin-cert-monitor to automatically renew certificates by setting a renewal_offset (default is 14 days) and a check_interval (default is 24 hours). You can enable or disable this schedule using the robin k8s-cert update command.

  • Trigger immediate renewals using robin k8s-cert renew --force. For situations requiring an immediate refresh regardless of the current expiry date, use the --force-renew option.

  • Simulate renewal processes without affecting services or restarting control-plane components by using the --dry-run flag.

Note

Manual and automated renewals trigger a restart of control-plane components, such as the API server and Etcd, which may cause a brief service disruption of 10 to 30 seconds per master node.

For more information, see Manage Kubernetes Control-Plane Certificates.

26.1.4.7. Client-side RPC request timeout support

Robin CNP v5.9.0 supports client-side Remote Procedure Call (RPC) request timeouts. This feature prevents application I/O hangs and maintains system responsiveness during server overloads by enforcing classified service-level agreements (SLAs) for Control, IO, and Blob I/O.

The following table provides the configuration parameters and the default values:

Parameter

Purpose

Default value

rio_io_rpc_timeout_ms

Timeout for RIO I/O RPC requests. The default value is 10 seconds.

10 seconds

rdvm_io_rpc_timeout_ms

Timeout for RDVM I/O RPC requests. The default value is 10 seconds

10 seconds

rdvm_rio_control_rpc_timeout_ms

Timeout for RDVM-RIO control RPC operations. The default value is 30 seconds.

30 seconds

rdvm_blob_io_rpc_timeout_ms

Timeout for RDVM-blob control RPC operations. The default value is 30 seconds.

30 seconds

For more information, see Configure and monitor client-side RPC timeouts.

26.1.4.8. Update KVM emulator pinning after installation

Robin CNP v5.9.0 enables you to update KVM emulator CPU pinning on running clusters. Previously, you could only define these settings during the initial installation. This feature lets you modify resource isolation settings as workloads scale or hardware requirements change, without reinstalling Robin CNP.

How it works

A new utility script, update-emulatorpin-cpuset.sh, is available on all nodes. This script automates updates to internal Robin configuration files and Kubernetes ConfigMaps.

To use this feature, you must perform the following:

  • Update the kubelet configuration and run the script on each node.

  • Reboot the node to apply the changes to all running Pods.

  • Update the application manifest and recreate existing applications to enable pinning for those workloads.

Limitations

  • NUMA node scope: Multiple NUMA nodes are not supported.

  • The emulator pin configuration only affects VMs hosted on the same NUMA node as the reserved CPU cores.

For more information, see Update node configuration for emulator pinning.

26.1.4.9. Asynchronous Disaster Recovery (Tech Preview)

Starting from Robin CNP v5.9.0, Robin CNP provides the snapshot-based Asynchronous Disaster Recovery (DR) feature.

The feature enables you to replicate your Kubernetes-based stateful applications along with its constructs (PVC, StatefulSet, config maps, secrets, services, etc.) onto a remote secondary peer cluster (site), and you can manually failover to it in the event of a disaster or maintenance activities. You can enable encryption when transmitting data over the wire to a peer cluster.

The Robin Asynchronous Disaster Recovery feature allows you to bring your applications online faster by failing over to the secondary cluster (site) in the event of a disaster with a minimum application downtime and failback later. For more information, see Asynchronous Disaster Recovery

26.1.4.10. Data locality tracking for volumes

Starting with Robin CNP v5.9.0, Robin CNP provides visibility into the data locality ratio for each volume mount. The data locality ratio indicates the percentage of a volume’s data that is physically stored on the node where the volume is currently mounted.

When the data locality percentage is high, I/O operations are served locally, avoiding network hops. This reduces latency and improves storage performance.

The data locality ratio is reported as a percentage between 0% and 100%.

  • 100% - All leader data for a volume resides on the mount node (fully local).

  • 0% - None of the leader data for a volume resides on the mount node (fully remote).

Viewing data locality ratio

You can view the data locality ratio for a volume using the following ways:

  • The Mount column in the robin volume list command.

  • The Data Locality column in the Mounts table in the robin volume info command.

  • The mount_data_locality label in the robin_vol_mount_node_ids metric.

For more information, see Data Locality Tracking for Volumes.

26.1.4.11. Support for filesystem mount option for PVCs

Starting with Robin CNP v5.9.0, you can specify filesystem mount options for Persistent Volume Claims (PVCs) by adding the following parameter in the StorageClass configuration:

robin.io/mount_options

This feature provides control over how volumes are mounted on nodes, which is particularly useful for managing block discard behavior.

The following values are supported:

  • discard: Enables the discard/TRIM command. This is the default value.

  • nodiscard: Disables the discard/TRIM command.

To use this feature, add the robin.io/mount_options parameter to the StorageClass definition as shown in the following example:

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: robin-nodiscard
  labels:
    app.kubernetes.io/instance: robin
    app.kubernetes.io/managed-by: robin.io
parameters:
  faultdomain: host
  fstype: ext4
  protection: quorum-replication
  replication: "3"
  robin.io/mount_options: "nodiscard"
provisioner: robin
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer

Note

If any option other than the supported mount option is specified, the PVC remains in the Pending state and provisioning fails with an error message listing the supported values.

26.1.4.12. Enable device ownership for non-root containers

Starting with Robin CNP v5.9.0, you can enable device ownership for non-root containers. This feature allows non-root containers to access host character devices such as /dev/vhost-net, /dev/vfio/<group>, and /dev/net/tun without manual chown commands on each node.

By default, non-root containers cannot access host character devices because Linux enforces device node ownership and permission bits at the host level. If the container’s UID or GID defined in the Pod’s securityContext does not match with the device’s ownership or group, the container receives a permission denied error even when the device is mounted and the required Linux capabilities are granted.

By default, this feature is disabled. When this feature is enabled, the container runtime automatically changes the ownership of the mounted device node to match the UID and GID of the non-root user defined in the Pod’s securityContext during startup.

The method Robin CNP uses to manage device ownership depends on the container runtime installed on each node.

CRI-O runtime for RHEL 9

When this feature is enabled, Robin CNP programs RHEL 9 nodes running CRI-O with the enable_device_ownership_from_security_context = true flag in the CRI-O configuration.

With this flag enabled, CRI-O automatically changes the ownership of the mounted device node to match the runAsUser and runAsGroup values defined in the Pod’s securityContext at container startup. No manual node configuration is required.

udev rules for RHEL 8 and Rocky Linux 8

Because the CRI-O (enable_device_ownership_from_securit_context) mechanism does not work with nodes using Docker as CRI, Robin CNP installs udev rules in the /etc/udev/rules.d/99-robin-device-ownership.rules location on each node. This rule sets device permissions to world-accessible (0666) whenever the device node is created or the system boots.

On bare-metal nodes, the vhost_net kernel module is also configured for automatic loading at boot to ensure the ``udev``rule triggers correctly on every startup.

To enable this feature at the cluster level, you must pass the following option as part of the GoRobin install, upgrade, or add node command:

--enable-device-ownership

26.1.4.13. VLAN Quality of Service (QoS) for Robin IP-Pool

Starting with Robin CNP v5.9.0, a new --qos option is added as part of the robin ip-pool add command. This option sets the 802.1p priority bits within an Ethernet VLAN tag to prioritize network traffic. The valid values are integers from 0 to 7, where higher values indicate a higher priority for network traffic. When creating an IP-Pool with the --qos option, you must specify the --vlan option. This is called VLAN QoS.

Note

The VLAN QoS can only be configured when creating a new IP pool. Adding VLAN QoS to an existing IP pool is not supported.

This option is only valid for OVS and SR-IOV driver-backed IP pools. When an IP address is assigned from an IP-Pool configured with these options, Robin CNP automatically tags outbound traffic from the Pod with the specified value in the VLAN header. This feature is used in Radio Access Network (RAN) environments to support latency-sensitive workloads, such as 5G Central Unit (CU)/Distributed Unit (DU), that require guaranteed traffic priority.

26.1.4.14. Support for tenant user groups

Starting with Robin v5.9.0, you can use Tenant User Groups (TUGs) in CNP. The feature enables you to manage users within a cluster tenant collectively. You can map a group of users to a single tenant to manage their roles, capabilities, and permissions as a single unit.

The TUG feature supports local users, users managed by an external Identity Provider (IdP) such as Keycloak or LDAP, or a mixture of both. When you update a TUG’s permissions, Robin CNP automatically applies the changes to all group members.

A TUG can be linked to a user group from an external IdP. In such a configuration, all members of the IdP user group will automatically be added to the TUG. If a user gets added to the IdP group, they will also be added to the TUG. Conversely, users who are removed from the IdP group will be removed from the TUG.

TUGs that are mapped to external IdP groups can only contain members from that group. For unmapped TUGs, however, there are no constraints on group membership or authentication type. A TUG can hold a mix of local, LDAP-imported, Keycloak-imported, or any other user type simultaneously.

The new robin tenant-user-group command provides comprehensive control, enabling you to add, update, and remove groups, as well as manage specific user capabilities.

For more information, see Tenant user groups.

26.1.4.15. Support for OpenID Connect (OIDC) Identity Providers

Robin CNP now supports external OpenID Connect (OIDC) identity providers. This feature enables single sign-on (SSO) across your applications and web properties by integrating with providers like Keycloak, Google, and other OAuth 2.0-compatible services.

Support for Multiple Identity Provider Types

You can now register and manage various identity provider (IdP) types to handle authentication and user management:

  • Keycloak (keycloak_provider): Integrates with Keycloak servers for advanced user and group management, including Tenant User Group (TUG) synchronization.

  • Google (google_provider): A pre-configured provider for Google OIDC services.

  • Generic OAuth (generic_oauth_provider): Supports standard OAuth 2.0/OIDC flows, including Authorization Code Flow and Device Flow.

  • LDAP (ldap_provider): Integrates with traditional LDAP directory services.

For more information, see OpenID Connect identity providers.

26.1.4.16. Group-tagged RoleBindings

Robin CNP v5.9.0 provides the group-tagged RoleBindings feature to optimize how the CNP manages access to tenant namespaces.

By binding permissions to a group Subject instead of individual users, this feature significantly reduces the total number of RoleBinding objects, which lowers etcd pressure and improves cluster scalability. For more information, see Group-tagged RoleBindings.

26.1.4.17. Users login session management

Robin CNP v5.9.0 supports comprehensive login session management, allowing you to maintain secure and independent access across multiple devices and interfaces.

When you log in, the Robin CNP assigns a unique session ID (UUID) to your environment and embeds it within a JSON Web Token (JWT) to verify your identity and session status.

This feature provides full visibility into active sessions, including those for ephemeral users who authenticate through external identity providers. You can now view detailed session information—such as the associated username, tenant, and unique device ID—for both CLI and UI-based access.

To enhance security, this update introduces role-based session control that allows you to revoke access immediately, even before a session naturally expires.

While super admins can manage all sessions across the cluster and tenant admins can manage sessions within their specific tenant, individual users can view and terminate their own active sessions. For more information, see Manage users login sessions.

26.1.4.18. Metrics for Stormgr, RIO, and RDVM process lifecycle

Starting with Robin CNP v5.9.0, new metrics are added to track the process lifecycle of the following Robin storage daemons:

  • Store Manager (Stormgr) on the master node

  • Robin I/O (RIO) on each worker node

  • Robin Distributed Volume Manager (RDVM) on each worker node

The following are the newly added metrics:

Metric

Description

robin_stormgr_start_time_epoch_seconds

Unix timestamp (in seconds) when Stormgr last started.

robin_stormgr_ready_time_epoch_seconds

Unix timestamp (in seconds) when Stormgr reached the ready state after its last start.

robin_rio_start_time_epoch_seconds

Unix timestamp (in seconds) when the RIO process last started.

robin_rio_ready_time_epoch_seconds

Unix timestamp (in seconds) when the RIO completed initialization after its last start.

robin_rdvm_start_time_epoch_seconds

Unix timestamp (in seconds) when the RDVM process last started.

robin_rdvm_ready_time_epoch_seconds

Unix timestamp (in seconds) when the RDVM completed initialization after its last start.

Note

All metrics carry an instance label set to the node’s hostname.

Example of Prometheus queries:

  • Detect iomgr restarts: changes(robin_iomgr_start_time_epoch_seconds[5m]) > 0

  • Detect stormgr restarts: changes(robin_stormgr_start_time_seconds[5m]) > 0

For more information, see Storage Manager Metrics.

26.1.4.19. Tag management using Robin CNP UI

Robin CNP v5.9.0 supports tag management from the Robin CNP UI, providing parity with existing command-line interface (CLI) capabilities. You can now use the UI to organize and categorize resources, including nodes and disks, more efficiently.

The following are the key features:

  • Apply tags to various resources such as nodes and disks to improve organization and filtering.

  • Add single or multiple values to a tag key.

  • To manage tags, log in to the UI and navigate to Settings > Tags.

26.1.5. Improvements

26.1.5.1. Enhanced GoRobin utility tool for better reliability and resumability

The GoRobin utility tool is an advanced orchestration and automation tool for deploying, managing, and upgrading Robin CNP clusters in on-premises environments. Starting with Robin CNP v5.9.0, this utility tool is enhanced with a robust, state-driven architecture designed for enterprise-grade reliability and resumability.

The following are the architectural and operational improvements in the GoRobin utility tool:

  • State machine-based architecture: Uses a state machine with transitions instead of linear, procedural execution.

  • Persistent state: Stores data in an SQLite database to track the state of operations instead of using in-memory storage that is lost on failure.

  • Resume capability: Allows you to resume operations from the last successful state instead of restarting from scratch.

  • Parallel task execution: Executes tasks in parallel using 40 worker threads. This significantly reduces the time required for large-scale deployments and upgrades compared to the sequential execution in the previous release.

  • Resilient long-running operations: Provides auto-recovery for long-running operations to reduce the risk of failure.

26.1.5.2. Google Secure LDAP integration

Robin CNP v5.9.0 supports integration with Google Secure LDAP (Google Workspace). Use the new GOOGLE_CLOUD LDAP server type to configure this integration. This update also adds default username search attributes to improve user discovery in Google Workspace environments.

For more information, see Add an LDAP server to a Robin cluster.

26.1.5.3. External PKI integration for Robin certificates

Robin CNP v5.9.0 supports integration with external Public Key Infrastructure (PKI) servers, such as EJBCA and Keyfactor. This improvement allows you to automate the issuance and renewal of internal and external certificates using an external, trusted Certificate Authority (CA) instead of relying on default self-signed certificates.

This feature is required for enterprises with strict CA policies that mandate all certificates be managed by an external, trusted CA. For more information, see Manage external PKI server.

26.1.5.4. Unix socket-based VNC for KVM virtual machines

Robin CNP v5.9.0 provides a secure alternative to traditional TCP-based VNC console access for KVM virtual machines (VMs). This enhancement resolves a security vulnerability that exposed VNC ports (5900+) on host network interfaces, potentially allowing unauthorized console access through host-level port forwarding.

To use this feature, set vnc_mode: socket in the application bundle manifest. Existing applications continue to use the default tcp mode to ensure backward compatibility. You can only configure this setting during deployment. To apply this change to existing VMs, you must redeploy them.

Note

This setting is supported for both Linux and Windows KVM engines.

For more information about configuration and administrative access, see Securing virtual machine console access.

26.1.5.5. Built-in Calico Felix Prometheus Metrics

Starting with Robin CNP v5.9.0, Felix Prometheus metrics are enabled by default on every calico-node DaemonSet Pod. In earlier versions, you had to manually enable metrics collection by running the kubectl set env command.

The setting is preserved during upgrades. For verification steps and available metrics, see Calico Prometheus metrics.

26.1.5.6. Export stalled I/O as metrics

The Sherlock diagnostic tool with the Robin CNP v5.9.0 release exposes stalled I/O data as a metric. The new robin_vol_iostall_pending_ios metric tracks the number of pending I/O operations on a volume. You can use this metric for monitoring and alerting through the /metrics endpoint.

Metric labels

Label

Description

name

The name of the volume, such as the PVC name.

volid

The volume ID.

node

The node on which the pending IOs are detected.

devname

The device name on which the pending IOs are detected.

Example

The following example shows the robin_vol_iostall_pending_ios metric:

robin_vol_iostall_pending_ios{name="pvc-7ec92c7e-58d7-49bc-bab4-e3c02d3a4912",volid="23",node="vnode-103-75",devname="sdd"} 1

For more information, see Volume Metrics.

26.1.5.7. Enhanced robin log collect command

The robin log collect command output displays the details for the following CLI commands with the Robin CNP v5.9.0 release:

  • rstorstat snapcache summary

  • stormgr snapshot list

  • stormgr snapshot stats

  • robin config list

26.1.5.8. Execution Timestamps for CNP CLI outputs

Starting with Robin CNP v5.9.0, when you run any CLI commands using the --urlinfo option, the command output displays the timestamp under the Execution timestamp field.

The execution timestamp is generated by the server and uses the YYYY-MM-DD HH:MM:SS,mmm format (for example, 2026-07-03 02:44:07,588 PDT).

Example

# robin host list --urlinfo
Id           | Hostname       | Version   | Status | RPool   | Avail. Zone | LastOpr | Roles | Cores        | GPUs  | Mem         | HDD(#/Alloc/Total) | SSD(#/Alloc/Total) | Pod Usage | Joined Time
-------------+----------------+-----------+--------+---------+-------------+---------+-------+--------------+-------+-------------+--------------------+--------------------+-----------+----------------------
1782893726:1 | hypervvm-73-62 | 5.9.0-613 | Ready  | default | N/A         | ONLINE  | S,C   | 6.45/5.55/12 | 0/0/0 | 19G/12G/31G | 2/-/200G           | -/-/-              | 74/36/110 | 01 Jul 2026 04:20:17
1782893726:2 | hypervvm-73-64 | 5.9.0-613 | Ready  | default | N/A         | ONLINE  | S,C   | 7.5/4.5/12   | 0/0/0 | 22G/8G/31G  | 2/-/200G           | -/-/-              | 83/27/110 | 01 Jul 2026 04:20:34
1782893726:3 | hypervvm-73-63 | 5.9.0-613 | Ready  | default | N/A         | ONLINE  | C,S   | 7.45/4.55/12 | 0/0/0 | 20G/10G/31G | 2/-/200G           | -/-/-              | 85/25/110 | 01 Jul 2026 05:49:55
* Note: all values indicated above in the format XX/XX/XX represent the Free/Allocated/Total values of the respective resource unless otherwise specified. In addition allocated values for compute resource such as cpu, memory and pod usage includes reserved values for the corresponding resource.
    CURL cmd :
        curl -g -i -k -X GET -H "X-Robin-Master: fd74:ca9b:3a09:868c:172:18:0:ae8d" -H "Authorization: eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzZXNzaW9uX2lkIjoiZTBmOWYxYjQtNzZhZi0xMWYxLTk3ODEtZjU5ZGU2YTBlZWI0IiwidXNlcl9pZCI6NSwidGVuYW50X2lkIjoxLCJleHAiOjE3ODMxNTcyNjR9.dMEG-_l4MFKmAOK_mDLqmeuGUpUTggnUg7_MCexJV-k" https://[fd74:ca9b:3a09:868c:172:18:0:ae8d]:29442/api/v5/robin_server/hosts
    HTTP code :
        200
    Execution timestamp :
        2026-07-03 02:44:07,588 PDT

Robin Pods such as robin-master, robin-worker, and iomgr also display the system time and timezone by default before the container name.

Example

[2026-05-27 11:56:13 PDT][robinmaster@master ~]#

26.1.5.9. Metrics for CSI sidecar containers

Starting with Robin CNP v5.9.0, the following metrics for CSI sidecar containers are exposed in the Prometheus format through an HTTP endpoint:

  • _total (Counters) - Total number of CSI operations (for example, CreateVolume, DeleteVolume).

  • _errors_total (Counters) - Total number of errors encountered during CSI operations.

  • _duration_seconds (Histograms) - Duration and latency of CSI calls in seconds.

Each metric includes labels that provide additional details, such as the CSI call type (driver_name, grpc_status_code, method_name).

Example

csi_sidecar_operations_seconds_count{driver_name="robin",grpc_status_code="OK",method_name="/csi.v1.Controller/ControllerGetCapabilities"} 1
csi_sidecar_operations_seconds_sum{driver_name="robin",grpc_status_code="OK",method_name="/csi.v1.Controller/ControllerGetCapabilities"} 0.000508898
csi_sidecar_operations_seconds_bucket{driver_name="robin",grpc_status_code="OK",method_name="/csi.v1.Identity/Probe",le="0.1"} 1

For more information, see CSI Sidecar Container Metrics.

26.1.5.10. FRR-K8s mode as default mode for MetalLB BGP backend

Starting with Robin CNP v5.9.0, the MetalLB version is updated from v0.14.9 to v0.16.1. In this version, the FRR mode is deprecated, and the FRR-K8s mode is now the default BGP backend for MetalLB. It provides the same functionality as the FRR mode with additional flexibility.

For more information, see Load Balancer Support using MetalLB.

26.1.5.11. Unicast mode support for Keepalived

Robin CNP v5.9.0 supports unicast mode for Keepalived Virtual Router Redundancy Protocol (VRRP) communication in Kubernetes High Availability (HA) setups. This improvement replaces the previous multicast-only requirement, allowing you to deploy Robin CNP in networking environments where multicast is restricted or unavailable.

Previously, Keepalived used multicast (224.0.0.18) for communication between master nodes, which required multicast routing to be enabled on all network switches and routers.

By utilizing unicast mode, Keepalived now sends VRRP advertisements directly to specific peer IP addresses. This change enables broader network compatibility and allows the use of duplicate Virtual Router IDs (VRIDs) within the same network segment.

Note

If you want to utilize multicast mode for your cluster, before upgrading to Robin CNP v5.9.0, you must create a Keepalived configuration with a .custom suffix or copy the existing /etc/keepalived/keepalived.conf to /etc/keepalived/keepalived.conf.custom. During the upgrade, Robin CNP checks for the .custom file; if present, it will use the configuration mentioned in this .custom file without modification.

26.1.6. Fixed Issues

Reference ID

Description

RSD-12477

When you perform a deferred clone hydration operation in Robin CNP, the IOMgr service could crash with an assertion failure in snapchot.c. This failure occurs due to a race condition between snapshot eviction and snapshot reloading operations. This issue is fixed

RSD-12642

When you scale down the cert-manager components (cert-manager, cert-manager-webhook, and cert-manager-cainjector deployments), the TLS secret is deleted. When the cert-manager components are scaled up, the secret will be automatically recreated. However, the regenerated certificate used an unexpected Subject Common Name (CN) instead of the original node’s Fully Qualified Domain Name (FQDN). This issue is fixed.

RSD-12500

When storage-compute affinity is configured on Robin CNP, it fails to allocate volume in certain scenarios when the first node (host) does not have resources without checking on other nodes. This issue is fixed.

RSD-12235, RSD-12315

In Robin CNP, when you try to relocate an application instance, the operation might fail, leaving the Pod stuck in a TopologyAffinityError state with this message: relocation. This issue occurs because Robin CNP incorrectly treats existing volume disk layouts as reserved resources during relocation. This issue is fixed.

RSD-12071

In a scenario, when two or more application creation jobs start at the same time with the same static IP, the validation check for the second job would return without raising an exception.

And this results in multiple bundle applications being allocated the same static IP, duplicate IP entries in the master logs for different job IDs, and resulting network conflicts and application reachability issues. This issue is fixed.

PP-42238

The Helm binary version on the host differs from the version in the robin-master pod or the robin-client, causing inconsistent Helm versions for different users (for example, root and cloud). This issue is fixed.

RSD-11546

When you enable VM metrics using the robin metrics start --enable-vm-metrics command, metrics for KVM applications do not appear in the Performance tab of the application details page. This issue is fixed.

RSD-11287

The issue of the kubelet service restarting automatically during a Robin CNP upgrade, which resulted in service interruptions, is fixed.

RSD-12381

The IOMgr crashes with a segmentation fault at seg_info_print in resync.c. The crash occurs during heavy IO workloads, frequent snapshots, or volume resync operations. The crash is caused by a race condition where a resync session is freed while the Remote Procedure Call (RCP) handler is still processing the segment list. This issue is fixed.

RSD-11395

The issue of Robin CNP not monitoring the validity of external identity certificates as part of its certificate management feature is fixed.

RSD-11453

When a Helm app is deployed on Robin CNP, it allocates more resources to a tenant than its limit. This issue is fixed.

RSD-10993

IOMgr crashes with an assertion in tcmu.c when rio_tcmur_siphon_rw_cmds() or rio_tcmur_handle_one_cmd() functions encounter a command length not aligned to the block size. This issue is fixed.

RSD-11165, RSD-10928

When upgrading from Robin CNP v5.5.0 to Robin CNP v5.5.1, the robin-master-bootstrap process might fail due to a non-idempotent insert into the user_credentials table. This issue occurs when a user has a blank current_context_ref in the users table while a corresponding entry already exists in user_credentials, causing a UniqueViolation that prevents the upgrade. This issue is fixed.

PP-40849

When you uninstall a Robin CNP v5.7.0 cluster using the GoRobin command, the uninstallation process fails during the robin-script.sh cleanup phase. This issue occurs because the GoRobin command did not forward the required credentials of the cluster to the underlying cleanup command. This issue is fixed.

PP-40645

In Robin CNP v5.5.1, the RDVM and RIO storage metrics were missing from the /metrics endpoint. This issue occurred because the monitor_iomgr_stats config parameter was disabled by default. This issue is fixed.

PP-39842

The robin host list command doesn’t account for long-running sidecar containers in the initContainers section for bundle applications. This causes Robin CNP to incorrectly place pods, which can lead to oversubscription or performance issues. This issue is fixed.

RSD-9478

When deploying applications with RWX PVCs, application Pods fail to mount volumes and get stuck in the ContainerCreating state because RPC requests get stuck in IO operation on the volumes, leading to degraded volumes and faulted storage drives. This issue is fixed.

PP-38331

In Robin CNP v5.5.0, the robin dev jobmgr info command does not display any output due to an API issue. This issue is fixed. This command now correctly displays the required output.

PP-38076

In Robin CNP v5.5.0, the issue of restoring a Robin bundle application from a backup fails if the associated storage repository is removed, and then re-registered with the same name is fixed.

RSD-11839

The issue where the robin-worker-bootstrap process caused upgrade timeouts while fetching the Consul encryption key and management token is fixed. As part of the fix, Robin CNP v5.9.0 now uses the service IP address or the fully qualified service name (robin-master.robinio.svc.cluster.local) to ensure faster and more reliable API connections.

RSD-12147

When a KVM instance is stopped or relocated, stale storage pool XML definitions remain on the worker node. If an operation triggers a libvirtd pool scan, all libvirtd threads attempt I/O against unmounted storage paths. This attempt blocks the kernel and causes subsequent virsh commands to hang until you restart libvirtd. This issue is fixed.

RSD-11007

In rare scenarios, after you perform an application upgrade on Robin CNP, the application’s Pod might enter into a flapping state where it repeatedly restarts. However, after a few redeployments, the Pod becomes stable on its own without any intervention, and this behavior is intermittent.

PP-41192

When creating a KVM app using the Robin bundle, the CPU core count must be specified in even numbers. If odd numbers are specified, the KVM will not be deployed; however, the respective Pod might be in the Running state. This issue is fixed.

PP-37965

In Robin CNP v5.5.0, when you scale up a Robin Bundle app, it is not considering the existing CPU cores and memory already in use by a vnode. As a result, Robin CNP is not able to find a suitable host, even though there are additional resources available. This issue is fixed.

PP-42237

When you try to deploy a KVM application with multiple vnodes that consume all available vfio-pci resources on a host, the application fails to restart. If you attempt to restart the application, it might result in displaying the following error, leaving the application in a FAULTED or NOTREADY state:

Failed to allocate resources for App

This issue is fixed.

26.1.7. Known Issues

Reference ID

Description

PP-45523

When you add a new volume (PVC) to a Kubernetes application that is already part of a protection group, the new volume will not be replicated to the secondary cluster.

PP-45445

When you replicate Helm-based or Kubernetes-based applications (flexapps) for Disaster Recovery (DR), Robin CNP does not propagate application-level configuration changes to the secondary cluster. While volume data is successfully replicated, changes to the application manifest—such as image tags, Helm chart revisions, or resource limits—are not captured.

Workaround If you update the application configuration on the primary cluster, you must manually apply the same configuration changes (such as running helm upgrade) on the secondary cluster to ensure both environments remain in sync

PP-45514

When you unpair a replication configuration and then re-pair a peer cluster using the same name, the initial synchronization of a protection group might fail with the following error message:

Error: Not authorized to perform this operation: – Authorization error: User no longer exists

This issue occurs if the storage repository from the previous pairing was not successfully removed during the unpairing process.

Workaround If you observe this issue, contact the Robin Customer Support team to resolve this issue.

PP-45503

In Robin CNP v5.9.0, when you run the robin user add-context command, it fails with a database NotNullViolation error and displays the following message:

Failure occurred during user update operation: (psycopg2.errors.NotNullViolation) null value in column “tenant_ref” of relation “user_contexts” violates not-null constraint

Workaround

If you observe this issue, contact the Robin Customer Support team to resolve this issue.

PP-45466

Symptom

When you restart a Pod that uses a ReadWriteMany (RWX) volume, the Pod might get stuck in the ContainerCreating state. This issue occurs when the volume mount fails with a return code 1 or 32, often accompanied by an ASSIGNED_ERR status in the Robin NFS export list.

Verify for the following symptoms:

  • Running kubectl describe pod <pod-name> shows a FailedMount warning with an error message similar to:

    Command ‘/bin/mount /dev/sdw /var/lib /robin/nfs/robin-nfs-shared-7/ganesha/pvc-64660d41-b541-4805-9583-75fe4b64a0e2 -o rw,discard’ failed with return code 32: mount: /var/lib/robin/nfs/robin-nfs-shared-7/ganesha/pvc-64660d41-b541-4805-9583-75fe4b 64a0e2: wrong fs type, bad option, bad superblock on /dev/sdw, missing codepage or helper program, or other error.

    dmesg(1) may have more information after failed mount system call.

  • The command robin nfs export-list shows the volume state as ASSIGNED_ERR.

  • System logs (dmesg) on the node hosting the NFS server pod show XFS corruption warnings, such as:

    XFS (sdw): Corruption warning: Metadata has LSN (11:103800) ahead of current LSN (11:101368). Please unmount and run xfs_repair (>= v4.3) to resolve.

    XFS (sdw): log mount/recovery failed: error -22

    XFS (sdw): log mount failed``

Workaround

If you encounter this issue, perform the following steps to repair the file system:

  1. Identify the PVC name associated with the failing Pod.

    # kubectl describe pod <pod-name>
    
  2. Find the NFS server Pod and the node where the export is assigned by running.

    # robin nfs export-list | grep <pvc_name>
    
  3. Log in to the node identified in the previous step.

  4. Check the dmesg output to identify the specific device (for example, /dev/sdw) associated with the XFS corruption error.

  5. Repair the file system by running the following command on the node.

    # xfs_repair -L <device_path>
    

    Note: Replace <device_path> with the device identified in step 4 (for example, /dev/sdw).

PP-45497

Symptom

In rare scenarios, an IOMgr Pod restart might cause drives to enter SUSPECTED_OFFLINE and ACCESS_FAILED states. This causes the host status to change to Partial, and the cluster phase to Degraded. The drives do not automatically recover, even after the IOMgr service returns to a healthy state.

Workaround 1. Identify the affected drives:

# robin disk list --host <hostname>

Look for drives with Status: ACCESS_FAILED.

  1. Unfault each affected drive:

# robin drive unfault <wwn>

Before running this on a production cluster, use robin volume list to confirm volumes on the affected drive have a healthy replica elsewhere (Protection should not be NONE without another healthy copy). If the volume has a healthy replica, proceed with the unfault. Otherwise, contact Robin Support — this includes cases where the volume has no healthy replica, or where a drive fails to unfault or re-enters ACCESS_FAILED shortly after.

PP-45487

Symptom

Volume allocation operation fails with faultdomain_customlabel parameter due to internal issues, and Robin CNP incorrectly reports a block size mismatch (for example, disk-4096 deviceset-512), even when the physical disk and volume block sizes match. This issue prevents the creation of Persistent Volume Claims (PVCs) and StatefulSets that rely on custom-label-based fault domains.

Workaround

If you observe this issue, contact the Robin Customer Support team to resolve this issue.

PP-39034

Symptom

When you delete ReadWriteMany (RWX) volumes, the associated shared nfs-server pods are not automatically deleted. By design, these pods stay active to allow for reuse by other volumes.

Workaround

Manually delete shared NFS server pods that are no longer in use.

  1. List all active exports and their associated NFS server pods:

    # robin nfs export-list
    
  2. In the output, check for active exports in the NFS Server Pod column. If a specific pod name does not appear in this list, it has no active exports and is safe to delete.

  3. Delete nfs-server pods that do not have an active export from the robinio namespace:

    # kubectl delete pod -n robinio <nfs_server_pod_name>
    

PP-41235

Symptom

When creating a clone from a volume snapshot, it might fail, and the robin job info command displays the following error:

Failed to create snapshot: 541 Vol <pvc_name>

The kubectl get volumesnapshot command displays that the volume snapshot is successful (READYTOUSE: true), but the robin volume-snapshot list displays no volume snapshot.

Workaround

  1. If a clone PVC was created from the snapshot, then delete the clone PVC:

    # kubectl delete pvc <clone-pvc>
    
  2. Retrieve the VolumeSnapshotContent object from VolumeSnapshot:

    # kubectl get volumesnapshot <snapshot-name>
    -o jsonpath='{.status.boundVolumeSnapshotContentName}'
    
  3. Remove the finalizer from the VolumeSnapshot object:

    # kubectl patch volumesnapshot <snapshot-name>
    -p '{"metadata":{"finalizers":[]}}' --type=merge
    
  4. Delete the VolumeSnapshot object:

    # kubectl delete volumesnapshot <snapshot-name>
    
  5. Remove the finalizer from the VolumeSnapshotContent object:

    # kubectl patch volumesnapshotcontent <snapcontent-name>
    -p '{"metadata":{"finalizers":[]}}' --type=merge
    
  6. Delete the VolumeSnapshotContent object:

    # kubectl delete volumesnapshotcontent <snapcontent-name>
    

PP-38078

Symptom

After a network partition, the robin-agent and iomgr-server may not restart automatically, and stale devices may not be cleaned up. This issue occurs because the consulwatch thread responsible for monitoring Consul and triggering restarts may fail to detect the network partition. As a result, stale devices may not be cleaned up, potentially leading to resource contention and other issues.

Workaround

Manually restart the robin-agent and iomgr-server using supervisorctl:

# supervisorctl restart robin-agent iomgr-server

PP-38471

Symptom

When StatefulSet Pods restart, the Pods might get stuck in the ContainerCreating state with the error: CSINode <node_name> does not contain driver robin due to stale NFS mount points and failure of the csi-nodeplugin-robin Pod due to CrashLoopBackOff state.

Workaround

If you notice this issue, restart the csi-nodeplugin Pod:

# kubectl delete pod <csi-nodeplugin> -n robinio

PP-42457

Symptom

When you upgrade your cluster from Robin CNP v5.7.2 with the best-effort-qos parameter enabled to Robin CNP v5.9.0, the robin-patroni and robin-auth-server pods might still request non-zero CPU.

In a best-effort configuration, all infrastructure Pods should have a CPU request of 0.

This issue occurs because the fix for best-effort Quality of Service (QoS) does not automatically apply during cluster upgrades. If you notice this issue, you need to apply the workaround.

Workaround

To resolve this issue and set the CPU requests to 0, manually update the Patroni operator and PostgreSQL configurations:

  1. Update the Patroni operator configuration:

    Run the following command to edit the operator configuration:

    # kubectl edit operatorconfigurations robin-patroni-postgres-operator
    -n robinio
    

    In the editor, set the following parameters to "0":

    • connection_pooler_default_cpu_limit

    • connection_pooler_default_cpu_request

    • postgres_pod_resources.default_cpu_limit

    • postgres_pod_resources.default_cpu_request

    • min_cpu_limit

  2. Update the PostgreSQL resource:

    Run the following command to edit the PostgreSQL resource:

    # kubectl edit postgresql -n robinio robin-patroni
    
  3. In the resources section, set the CPU limits and requests to "0":

    resources:
      limits:
        cpu: "0"
        memory: 8G
      requests:
        cpu: "0"
        memory: 4G
    
  4. Save the changes to start a rollout of new Patroni Pods and update the Patroni statefulset.

  5. Restart the operator pod:

    Restart will delete the existing operator pod to force a restart:

    # kubectl delete pod -n robinio <operator_pod_name>
    

    Note

    Replace <operator_pod_name> with the actual ID of your robin-patroni-postgres-operator Pod.

RSD-12729

Symptom

In the event of a power failure on a single-node cluster, Autopilot might fail to redeploy vnodes. If the SR-IOV NICs are not fully ready when the recovery job starts, the deployment fails with a PLAN_FAILED status. Autopilot does not automatically retry the deployment after this specific failure, leaving the affected pods in an Error state.

Workaround

To recover vnodes in this state, follow these steps:

  1. Identify the vnodes that are in the PLAN_FAILED state.

  2. Run a host probe for the affected node using the following command:

    # robin host probe <node-name> --rediscover --wait
    

    Replace <hostname> with the fully qualified hostname of the node to be probed.

PP-41172

Symptom

After upgrading from a supported Robin CNP version to Robin CNP v5.9.0, NFS mounts on client nodes can become unresponsive, leading to critical issues such as the following:

  • kubelet instability (frequent restarts)

  • patronictl command failures (connection refused)

  • kubectl exec operations failing for pods on affected nodes.

This problem is primarily observed when the NFS client (the node where the PVC is mounted) experiences prolonged unresponsiveness from the NFS server.

Workaround

If an NFS mount is hung, you can recover the system by forcing new NFS sessions:

  1. Identify hung NFS Mounts:

    • Attempt to access the NFS mount path.

      Example

      /var/lib/kubelet/pods/<pod_uid>/volumes/kubernetes.io~csi/<pvc_name>/mount.

    • If the command hangs (for example, ls /path/to/mount with no output and requiring Ctrl+C to exit), the mount is hung.

      Example

      $ ls /var/lib/pods/0cab5468-b43f-4afd-bad3/volumes/
      kubernetes.io~csi/pvc-7f31e2fc-b5b7-4991-ab97/mount
      
  2. Confirm that the NFS export for the hung PVC is in a READY state using robin nfs export-list.

    Example

    $ robin nfs export-list|grep pvc-7f31e2fc-b5b7-4991-ab97
    |READY|19|pvc-7f31e2fc-b5b7-4991-ab97|robin-nfs-shared-23|
    ["sm-compute02"]|192.02.204.31:/pvc-7f31e2fc-b5b7-4991-ab97|
    loaddbehg-fio|sachin|
    
  3. From the robin nfs export-list output, note the NFS Server Pod name serving the hung export. For example, in the above output, the NFS server Pod is robin-nfs-shared-23.

  4. Delete the identified NFS server pod. This action forces new NFS sessions and typically resolves the hung mount issue:

    # kubectl delete pod -n robinio <nfs_shared_server_pod_name>
    

    Example

    # kubectl delete pod robin-nfs-shared-23 -n robinio
    

PP-41195

Symptom

After you perform a force unmount operation for an RWX (ReadWriteMany) volume on a host where it was previously mounted, or during certain failover scenarios involving RWX volumes, the associated Robin NFS server Pod might transition into an ASSIGNED_ERR state.

When the Pod is in this state, the NFS server Pod is unable to export the volume, rendering the volume inaccessible via NFS.

Workaround

Contact the Robin Customer Support team to resolve this issue.

PP-41159

Symptom

After upgrading to Robin CNP v5.7.2, some Pods might get stuck in the ContainerCreating state because the VolumeUnmount job holds a lock on the volume, and the VolumeUnmount job shows the following error:

Target /var/lib/robin/nfs/robin-nfs-shared-107/ganesha/pvc-b84ed376-2a58-484f-8031-4530c1899b2c is busy, please retry later. (In some cases useful info about processes that use the device is found by lsof(8) or fuser(1))

Workaround

Apply the following workaround steps:

  1. Verify no pending I/Os at the RIO layer:

    # rio snapshot iolist
    
  2. Verify no in-flight I/Os on the relevant block device:

    # /sys/block/<dev>/inflight
    
  3. If there are no pending and in-flight I/Os, do a lazy unmount with the path shown in the VolumeUnmount job error message:

    # umount -l <mountpath>
    

PP-35015

Symptom

After renewing the expired Robin license successfully, Robin CNP incorrectly displays the License Violation error when you try to add a new user to the cluster. If you notice this issue, apply the following workaround.

Workaround

You need to restart the robin-server-bg service.

# rbash master
# supervisorctl restart robin-server-bg

PP-39901

Symptom

After rebooting a worker node that is hosting Pods with Robin RWX volumes, one or more application Pods using these volumes might get stuck in the ContainerCreating state indefinitely.

Workaround

If you notice the above issue, contact the Robin CS team.

PP-39645

Symptom

Robin CNP v5.9.0 may rarely fail to honor soft Pod anti-affinity, resulting in uneven Pod distribution on labeled nodes.

When you deploy an application with the recommended preferred DuringSchedulingIgnoredDuringExecution soft Pod Anti-Affinity, pods may not be uniformly distributed across the available, labeled nodes as expected. Kubernetes routes nodes to Robin CNP for pod scheduling. In some situations, a request to the Robin CNP from Kubernetes may not have the required node to honor soft affinity.

Workaround

Bounce the Pod that has not honored soft affinity.

PP-34226

Symptom

When a PersistentVolumeClaim (PVC) is created, the CSI provisioner initiates a VolumeCreate job. If this job fails, the CSI provisioner calls a new VolumeCreate job again for the same PVC. However, if the PVC is deleted during this process, the CSI provisioner will continue to call the VolumeCreate job because it does not verify the existence of the PVC before calling the VolumeCreate job.

Workaround

Bounce the CSI provisioner Pod.

# kubectl delete pod -n robinio <csi-provisioner-robin>

PP-34414

Symptom

In rare scenarios, the IOMGR service might fail to open devices in the exclusive mode when it starts as other processes are using these disks. You might observe the following issue:

Some app Pods get stuck in the ContainerCreating state after restarting.

Steps to identify the issue:

Check the following type of faulted error in the EVENT_DISK_FAULTED event type in the robin event list command:

disk /dev/disk/by-id/scsi-SATA_Micron_M500_MTFD_1401096049D5 on node default:poch06 is faulted

# robin event list --type EVENT_DISK_FAULTED

If you see the disk is faulted error, check the IOMGR logs for dev_open() and Failed to exclusively open error messages on the node where disks are present.

# cat iomgr.log.0 | grep scsi-SATA_Micron_M500_MTFD_1401096049D5
| grep "dev_open"

If you see the Device or resource busy error message in the log file, use fuser command to confirm whether the device is in use:

# fuser /dev/disk/by-id/scsi-SATA_Micron_M500_MTFD_1401096049D5

Workaround

If the device is not in use, restart the IOMGR service on the respective node:

# supervisorctl restart iomgr

PP-34492

Symptom

When you run the robin host list command and if you notice a host is in the NotReady and PROBE_PENDING states, follow these workaround steps to diagnose and recover the host:

Workaround

Run the following command to check which host is in the NotReady and PROBE_PENDING states:

# robin host list

Run the following command to check the current (Curr) and desired (Desired) states of the host in the Agent Process (AP) report:

# robin ap report | grep <hostname>

Run the following command to probe the host and recover it:

# robin host probe <hostname> --wait

This command forces a probe of the host and updates its state in the cluster.

Run the following command to verify the host’s state:

# robin host list

The host should now transition to the Ready state.

PP-35478

Symptom

In rare scenarios, the kube-scheduler may not function as expected when many Pods are deployed in a cluster due to issues with the kube-scheduler lease.

Workaround

Complete the following workaround steps to resolve issues with the kube-scheduler lease:

Run the following command to identify the node where the kube-scheduler Pod is running with the lease:

# kubectl get lease -n kube-system

Log in to the node identified in the previous step.

Check if the kube-scheduler Pod is running using the following command:

# docker ps | grep kube-scheduler

As the kube-scheduler is a static Pod, move its configuration file to temporarily stop the Pod:

# mv /etc/kubernetes/manifests/kube-scheduler.yaml /root

Run the following command to confirm that the kube-scheduler Pod is deleted. This may take a few minutes.

# docker ps | grep kube-scheduler

Verify that the kube-scheduler lease is transferred to a different Pod:

# kubectl get lease -n kube-system

Copy the static Pod configuration file back to its original location to redeploy the kube-scheduler Pod:

# mv /root/kube-scheduler.yaml /etc/kubernetes/manifests/

Confirm that the kube-scheduler container is running:

# docker ps | grep kube-scheduler

PP-36865

Symptom

After rebooting a node, the node might not come back online after a long time, and the host BMC console displays the following message for RWX PVCs mounted on that node:

Remounting nfs rwx pic timed out, issugin SIGKILL

Workaround

Power cycle the host system.

PP-37330

Symptom

During or after upgrading to Robin CNP v5.9.0, the NFSAgentAddExport job might fail with an error message similar to the following:

/bin/mount /dev/sdn /var/lib/robin/nfs/robin-nfs-shared-35/ganesha/pvc-822e76f0-9bb8-4629-8aae-8318fb2d3b41 -o discard failed with return code 32: mount: /var/lib/robin/nfs/robin-nfs-shared-35/ganesha/pvc-822e76f0-9bb8-4629-8aae-8318fb2d3b41: wrong fs type, bad option, bad superblock on /dev/sdn, missing codepage or helper program, or other error.

Workaround

If you notice this issue, contact the Robin Customer Support team for assistance.

PP-37416

Symptom

In rare scenarios, when upgrading from Robin CNP v5.7.2 to Robin CNP v5.9.0, the upgrade might fail with the following error during the Kubernetes upgrade process on other master nodes:

Failed to execute kubeadm upgrade command for K8S upgrade. Please make sure you have the correct version of kubeadm rpm binary installed

Steps to identify the issue:

  1. Check the /var/log/robin-install.log file to know why the upgrade failed.

    Example

    etcd container: {etcd_container_id} and exited status: {is_exited}

    Killing progress PID 4168272

    Failed to execute kubeadm upgrade command for K8S upgrade. Please make sure you have the correct version of kubeadm rpm binary installed

    Install logs can be found at /var/log/robin-install.log

    Caught EXIT signal. exit_code: 1

    Note

    You can get the above error logs for any static manifests of api-server, etcd, scheduler, and controller-manager.

  2. If you notice the above error, run the following command to inspect the Docker containers for the failed component. The containers will likely be in the Exited state.

    # docker ps -a | grep schedule
    

Workaround

If you notice the above error, restart the kubelet:

# systemctl restart kubelet

PP-38044

Symptom

When attempting to detach a repository from a hydrated Helm application, the operation might fail with the following error:

Can’t detach repo as the application is in IMPORTED state, hydrate it in order to detach the repo from it.

This issue occurs even if the application has already been hydrated. The system incorrectly marks the application in the IMPORTED state, preventing the repository from being detached.

Workaround

To detach the repository, manually rehydrate the application and then retry the detach operation:

  1. Run the following command to rehydrate the application.

    # robin app hydrate --wait
    
  2. Once the hydration is complete, detach the repository.

    # robin app detach-repo --wait –y
    

PP-38251

Symptom

When evacuating a disk from an offline node in the large cluster, the robin drive evacuate command fails with the following error message:

Json deserialize error: invalid value: integer -10, expected u64 at line 1 column 2440.

Workaround

If you notice the above issue, contact the Robin CS team.

PP-38087

Symptom

In certain cases, the snapshot size allocated to a volume could be less than what is requested. This occurs when the volume is allocated from multiple disks.

PP-38924

Symptom

After you delete multiple Helm applications, one of the Pods might get stuck in the Error state, and one or more ReadWriteMany (RWX) volumes might get stuck in the Terminating state.

Workaround

On the node where the Pod stuck in the Error state, restart Docker and Kubelet.

PP-34451

Symptom

In rare scenarios, the RWX Pod might be stuck in the ContainerCannotRun state and display the following error in the Pod’s event:

mount.nfs: mount system call failed

Perform the following steps to confirm the issue:

  1. Run the robin volume info command and check for the following details:

    1. Check the status of the volume. It should be in the ONLINE status.

    2. Check whether the respective volume mount path exists.

    3. Check the physical and logical sizes of the volume. If the physical size of the volume is greater than the logical size, then the volume is full.

  2. Run the following command to check whether any of the disks for the volume are running out of space:

    # robin disk info
    
  3. Run the lsblk and blkid commands to check whether the device mount path works fine on the nodes where the volume is mounted.

  4. Run the ls command to check if accessing the respective filesystem mount path gives any input and output errors.

If you notice any input and output errors in step 4, apply the following workaround:

Workaround

  1. Find all the Pods that are using the respective PVC:

    # kubectl get pods --all-namespaces -o=jsonpath='{range .items[]}
    {.metadata.namespace} /{.metadata.name}{"\t"}{.spec.volumes[].
    persistentVolumeClaim.claimName}{"\n"}{end}' | grep <pvc_nmae>
    
  2. Bounce all the Pods identified in step 1:

    # kubectl delete pod <pod> -n <namespace>
    

PP-40819

Symptom

From the Robin CNP UI, when you try to deploy an application by cloning from a snapshot, the operation might fail with the following similar error message indicating an invalid negative CPU value: Invalid value: “-200m”: must be greater than or equal to 0.

You might observe this issue specifically when the application has sidecar containers configured with CPU requests/limits. This is a CNP UI issue. You can use the CNP CLI to perform the same operation successfully.

Workaround

Use the following Robin CLI command to clone the snapshot and create an app:

# robin app create from-snapshot <new_app_name>
<snapshot_id> --rpool default --wait

PP-41022

Symptom

The robin host list command might incorrectly display negative values for CPU resources cores (specifically “Free” or “Allocated” CPU) on certain nodes. This occurs even when there are no user applications consuming significant CPU, suggesting a miscalculation or misreporting of available resources. The issue impacts the ability to accurately assess node capacity and schedule new workloads.

Workaround

If you notice this issue, apply the workaround.

Restart kubelet on the affected node:

# systemctl restart kubelet

PP-40514

Symptom

When an NFS shared server Pod becomes unavailable (for example, during a node reboot) and NFS failover is disabled, the kubelet on affected nodes might enter a defunct (zombie) state or restart continuously. This issue typically occurs in large clusters when ReadWriteMany (RWX) volumes are in use.

Workaround

Reboot the affected node.

# kubectl cordon <node-name>

PP-40993

Symptom

During large cluster upgrades, the upgrade might fail during Robin pre‑upgrade actions if Robin Auto Pilot creates active jobs. This occurs when multiple Robin Auto Pilot watchers are configured for a single pod, resulting in lingering jobs (for example, VnodeDeploy) that block the upgrade process.

Workaround

Restart the robin-master-bg service on the master node to clear active Auto Pilot jobs, then retry the upgrade.

PP-41599

Symptom

When creating a clone from a volume snapshot, it might fail, and the robin job info command displays the following error:

Failed to create snapshot: 541 Vol <pvc_name>

The kubectl get volumesnapshot command displays that the volume snapshot is successful (READYTOUSE: true), but the robin volume-snapshot list displays no volume snapshot.

Workaround

  1. If a clone PVC was created from the snapshot, then delete the clone PVC.

    # kubectl delete pvc <clone-pvc>
    
  2. Retrieve the VolumeSnapshotContent object from VolumeSnapshot.

    # kubectl get volumesnapshot <snapshot-name>
    -o jsonpath='{.status.boundVolumeSnapshotContentName}'
    
  3. Remove the finalizer from the VolumeSnapshot object.

    # kubectl patch volumesnapshot <snapshot-name>
    -p '{"metadata":{"finalizers":[]}}' --type=merge
    
  4. Delete the VolumeSnapshot object.

    # kubectl delete volumesnapshot <snapshot-name>
    
  5. Remove the finalizer from the VolumeSnapshotContent object.

    # kubectl patch volumesnapshotcontent <snapcontent-name>
    -p '{"metadata":{"finalizers":[]}}' --type=merge
    
  1. Delete the VolumeSnapshotContent object.

    # kubectl delete volumesnapshotcontent <snapcontent-name>
    

PP-45047

Symptom

In Robin CNP version v5.9.0, the application details page in the UI for a cloned application might appear blank. While the Info, Resources, Storage, and Manage tabs are visible, the data fields within these tabs (such as Pods, Memory, Volumes, Namespace,etc.) do not display any information

Workaround

To view the details of a cloned application, use the Robin CLI.

For example, run the following command:

# robin app info <cloned-app-name> --namespace <namespace>

PP-45076

Symptom

When a node is in a NotReady state due to a Container Runtime Interface (CRI) crash, the robin-file-server Pod might become stuck in the Init:0/1 state on a new node. This occurs because a stale VolumeAttachment remains on the failed node, preventing the ReadWriteOnce (RWO) PersistentVolumeClaim (PVC) from attaching to the new node.

Workaround

To resolve this issue, you must manually remove the stuck Pod and the stale volume attachment:

  1. Force-delete the pod that is stuck in the Terminating state on the failed node.

    # kubectl delete pod <pod-name> -n robinio --force --grace-period=0
    
  2. Delete the stale VolumeAttachment object associated with the failed node.

    # kubectl delete volumeattachment <attachment-name>
    

PP-44438

Symptom

The setup-gpu-operator option in the Config JSON is not supported in Robin CNP v5.9.0.

PP-45491

Symptom

If the robin-master key rotation is triggered without running the robin kms sync-backup-keyrefs command after the previous robin-master key rotation, the Robin master Pods enter an infinite restart loop.

This issue occurs because the rotate_kms_key flag remains stuck in the true state in the robin-bootstrap-config ConfigMap. The Robin server will not start as the Robin master bootstrap gets stuck with the following error:

Exception: rotate_kms_key: a previous rotation’s backup repo keyref sync is still pending (kms_backup_repo_sync_needed=true); refusing to start a new rotation until ‘robin kms sync-backup-keyrefs’ completes

Workaround

If you notice this issue, apply the following workaround:

  1. Set the rotate_kms_key flag to false in the robin-bootstrap-config ConfigMap:

    # kubectl edit cm -n robinio robin-bootstrap-config
    
  2. Find the current active Robin master Pod:

    # kubectl get lease -n robinio | grep master
    
  3. Delete the current active Robin master Pod:

    # kubectl delete pod -n robinio <current_active_master_pod>
    

26.1.8. Technical Support

Contact Technical support for any assistance.