Troubleshooting Kubernetes PVC Stuck in Pending on CentOS Stream / Rocky Linux
Resolve Kubernetes Persistent Volume Claims (PVCs) stuck in pending status on CentOS Stream and Rocky Linux by diagnosing StorageClass, provisioner, and underlying storage issues.
Resolve Kubernetes Persistent Volume Claims (PVCs) stuck in pending status on CentOS Stream and Rocky Linux by diagnosing StorageClass, provisioner, and underlying storage issues.
Introduction
A Kubernetes Persistent Volume Claim (PVC) stuck in a Pending status is a common and critical issue that prevents applications from starting or functioning correctly. When a PVC remains Pending, it means Kubernetes cannot find or provision an appropriate Persistent Volume (PV) to satisfy the claim's requirements. This guide provides a comprehensive, expert-level approach to diagnose and resolve such issues specifically tailored for environments running on CentOS Stream or Rocky Linux, where underlying host configurations and security policies might subtly influence storage provisioning.
Symptom & Error Signature
When a PVC is stuck in Pending, you will observe that your application pods requesting this PVC fail to start or remain in a ContainerCreating or Pending state themselves, often due to a FailedAttachVolume or similar error.
The primary symptom is seen when inspecting the PVC:
kubectl get pvc -n your-namespace
Expected output showing the stuck PVC:
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE
my-app-data-pvc Pending standard-rwo 2m
Further details, including specific events, can be retrieved from the PVC's describe output:
kubectl describe pvc my-app-data-pvc -n your-namespace
Typical error signatures in the Events section:
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning ProvisioningFailed 30s persistentvolume-controller failed to provision volume with StorageClass "standard-rwo": storageclass.storage.k8s.io "standard-rwo" not found
Warning ProvisioningFailed 30s persistentvolume-controller failed to provision volume with StorageClass "standard-rwo": no volume plugin matched name "kubernetes.io/nfs"
Warning ProvisioningFailed 30s persistentvolume-controller Failed to get StorageClass "standard-rwo" for volume: storageclass.storage.k8s.io "standard-rwo" not found
Warning VolumeBindingFailed 30s persistentvolume-controller no persistent volumes available for this claim and no storage class is set
Warning VolumeBindingFailed 30s persistentvolume-controller waiting for first consumer to bind volume; waiting for pod "my-app-pod-xxxx" to be scheduled
Root Cause Analysis
A PVC gets stuck in Pending when the Kubernetes control plane cannot successfully bind it to an available Persistent Volume (PV) or dynamically provision a new one. The underlying reasons typically fall into these categories:
Missing or Misconfigured StorageClass:
- The
StorageClassreferenced by the PVC does not exist. - The
StorageClassexists but references a non-existent or incorrectly configured provisioner. - No default
StorageClassis defined, and the PVC doesn't explicitly specify one.
- The
Storage Provisioner Issues:
- The Container Storage Interface (CSI) driver or in-tree provisioner specified by the
StorageClassis not installed or not running in the cluster. - The CSI driver's components (e.g., controller, node-driver registrar, actual driver pods) are not healthy or lack necessary permissions.
- For
hostPathor local-path provisioners, the target directories on the nodes might not exist or have incorrect permissions. - For network storage (NFS, iSCSI, Ceph), the underlying storage system is unreachable, misconfigured, or the client utilities are not installed/configured on the worker nodes (e.g.,
nfs-utilsfor NFS).
- The Container Storage Interface (CSI) driver or in-tree provisioner specified by the
Insufficient or Mismatched Persistent Volumes (for static provisioning):
- If dynamic provisioning is not used or fails, the PVC requires a pre-created PV.
- No available PV matches the PVC's requested
capacity,accessModes, orStorageClass. - A PV exists, but its
volumeMode(Filesystem vs. Block) does not match the PVC's implicit or explicit request.
Access Mode Incompatibility:
- The PVC requests
ReadWriteMany(RWX), but theStorageClassor available PVs only supportReadWriteOnce(RWO) orReadOnlyMany(ROX). Many storage systems (likehostPathor basic block devices) do not natively support RWX.
- The PVC requests
Volume Binding Mode (Immediate vs. WaitForFirstConsumer):
- If
volumeBindingMode: WaitForFirstConsumeris set on the StorageClass, the PVC will remainPendinguntil a Pod attempts to use it. This is normal behavior, but if the Pod itself is stuck (e.g., due to resource constraints or scheduling issues), the PVC will also appear stuck. - If
volumeBindingMode: Immediateis set, the PV should be bound as soon as the PVC is created, regardless of Pod scheduling. If it'sPendingwith this mode, it indicates a provisioning failure.
- If
Node-Specific Configuration Issues (CentOS/Rocky Linux focus):
- SELinux: The default enforcing SELinux policy on CentOS/Rocky can block mounting or accessing volumes, especially for
hostPathor network-mounted volumes, if appropriate SELinux contexts are not applied. - FirewallD:
firewalldmight be blocking necessary ports for network storage protocols (NFS, iSCSI, Ceph) between worker nodes and the storage server, or even between Kubernetes components. - Missing Utilities: Essential client utilities for network storage (e.g.,
nfs-utils,iscsi-initiator-utils,ceph-common) might be missing on worker nodes.
- SELinux: The default enforcing SELinux policy on CentOS/Rocky can block mounting or accessing volumes, especially for
Step-by-Step Resolution
Follow these steps to diagnose and resolve PVC Pending issues. Each step builds upon the previous, focusing on the most common causes first.
1. Inspect the PVC and its Events
Always start by examining the PVC itself and, critically, its events. This provides the most direct clues.
kubectl describe pvc <pvc-name> -n <namespace>
Look for Warning or Failed events. These often explicitly state the problem, such as "StorageClass not found" or "no volume plugin matched."
2. Verify StorageClass Existence and Configuration
The absence or misconfiguration of the StorageClass is a leading cause.
a. Check if the StorageClass exists:
If your PVC specifies a StorageClass (e.g., storageClassName: standard-rwo in the PVC YAML), ensure it actually exists.
kubectl get sc
If it's missing, you need to create it. Example for a standard-rwo StorageClass using hostPath (for testing/single-node clusters, not production):
# storageclass-hostpath.yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: standard-rwo
provisioner: k8s.io/minikube-hostpath # Or a real CSI driver like hostpath.storage.k8s.io
volumeBindingMode: Immediate # Or WaitForFirstConsumer
reclaimPolicy: Delete
---
The
k8s.io/minikube-hostpathprovisioner is a common placeholder for simplehostPathin minikube. For real clusters, you'd use a specific CSI driver (e.g.,local-path.storage.k8s.ioiflocal-path-provisioneris installed, or a vendor-specific CSI driver). If you intend to usehostPathon a multi-node cluster, understand its limitations as PVs are bound to specific nodes.
Apply the StorageClass:
kubectl apply -f storageclass-hostpath.yaml
b. Inspect the StorageClass details:
If the StorageClass exists, examine its provisioner and parameters.
kubectl describe sc <storageclass-name>
Ensure the Provisioner field matches a CSI driver or in-tree provisioner installed in your cluster.
3. Verify Storage Provisioner/CSI Driver Status
If the StorageClass is present and correctly defined, the issue often lies with the provisioner itself.
a. Identify the Provisioner:
From kubectl describe sc <storageclass-name>, note the Provisioner field. Common provisioners include:
kubernetes.io/no-provisioner(for static PVs)kubernetes.io/aws-ebs,kubernetes.io/gce-pd,kubernetes.io/azure-disk(in-tree cloud providers – mostly deprecated for CSI)pd.csi.storage.gke.io,ebs.csi.aws.com,disk.csi.azure.com(CSI cloud providers)hostpath.storage.k8s.io(e.g., used bylocal-path-provisioner)nfs.csi.k8s.io,cephfs.csi.ceph.com,iscsi.csi.k8s.io(external CSI drivers)
b. Check Provisioner Pods/Deployments:
Search for the corresponding CSI driver deployments, statefulsets, or daemonsets in your cluster, usually in the kube-system or a dedicated namespace (e.g., local-path-storage).
kubectl get pods -A | grep -i <provisioner-keyword>
# Example for local-path-provisioner:
kubectl get pods -A | grep -i local-path
Check their status and logs:
kubectl logs <provisioner-pod-name> -n <provisioner-namespace>
kubectl describe pod <provisioner-pod-name> -n <provisioner-namespace>
Look for errors indicating failure to initialize, communicate with the API server, or interact with the underlying storage.
4. Address Static PV Binding Issues (if not using dynamic provisioning)
If you're using static provisioning (i.e., you manually create PVs, or the StorageClass has provisioner: kubernetes.io/no-provisioner), ensure there's a suitable PV available.
kubectl get pv
Look for a PV with STATUS: Available. Ensure it meets the PVC's requirements:
CAPACITY: Must be equal to or greater than the PVC's request.ACCESS MODES: Must include all modes requested by the PVC.STORAGECLASS: Must match the PVC'sstorageClassName, or both must be empty for binding.VOLUME MODE: Must match (Filesystem is default).
If no suitable PV exists, create one manually:
# pv-static-hostpath.yaml
apiVersion: v1
kind: PersistentVolume
metadata:
name: my-static-pv
spec:
capacity:
storage: 1Gi
volumeMode: Filesystem
accessModes:
- ReadWriteOnce
persistentVolumeReclaimPolicy: Retain # Or Delete, Recycle
storageClassName: standard-rwo # Must match the PVC's StorageClass
hostPath:
path: "/mnt/data/my-static-volume" # Path on the worker node
---
Using
hostPathfor static PVs in a multi-node cluster requires careful management to ensure the PV is scheduled on a node where the path actually exists and is not used by other PVs. It's generally not recommended for production multi-node environments.
Apply the PV:
kubectl apply -f pv-static-hostpath.yaml
5. Check Node-Specific Configuration (CentOS Stream / Rocky Linux)
For hostPath volumes, or CSI drivers relying on node-level access (e.g., NFS, iSCSI, local-path), the host OS configuration is crucial.
a. HostPath Permissions and Existence:
If using a hostPath volume (either directly in a PV or via a local-path-provisioner), ensure the directory exists on the relevant worker nodes and has correct permissions.
Connect to the worker node(s) where your pod is scheduled (or potentially could be scheduled):
ssh <worker-node-ip>
Check directory existence and permissions:
sudo mkdir -p /mnt/data/my-static-volume # Example path from PV
sudo chmod 777 /mnt/data/my-static-volume # Or more restrictive permissions
sudo chown 1000:1000 /mnt/data/my-static-volume # Example UID:GID if known
The ownership (UID/GID) of the host directory often needs to match the user/group ID of the process inside the container attempting to write to the volume. Use
securityContextin your Pod definition to specifyrunAsUserandfsGroupif needed.
b. SELinux Policies: SELinux can block mount operations or file access even with correct permissions. Check audit logs for AVC denials:
sudo ausearch -c 'containerd' -m AVC,USER_AVC -ts recent
sudo ausearch -c 'kubelet' -m AVC,USER_AVC -ts recent
If you see SELinux denials related to paths used by volumes, you may need to add a policy or relax the context for the specific path. Temporarily set SELinux to permissive mode for testing (DO NOT do this in production unless fully understood):
sudo setenforce 0
If setting to permissive resolves the issue, you need to create a proper SELinux policy or apply the correct context. For hostPath or network mounts, container_file_t or container_var_lib_t are common contexts.
sudo semanage fcontext -a -t container_file_t "/mnt/data/my-static-volume(/.*)?"
sudo restorecon -Rv /mnt/data/my-static-volume
c. FirewallD Rules:
For network storage, firewalld on CentOS/Rocky nodes might be blocking access to the storage server.
Check the active zones and rules:
sudo firewall-cmd --get-active-zones
sudo firewall-cmd --list-all-zones
If using NFS (port 2049), iSCSI (port 3260), or Ceph (various ports), ensure outbound connections are allowed to the storage server. You might need to add service or port rules, e.g., for NFS:
sudo firewall-cmd --zone=public --add-service=nfs --permanent
sudo firewall-cmd --reload
d. Install Storage Client Utilities: For network storage, ensure the necessary client utilities are installed on all worker nodes.
- NFS:
sudo dnf install -y nfs-utils sudo systemctl enable --now nfs-client.target - iSCSI:
sudo dnf install -y iscsi-initiator-utils sudo systemctl enable --now iscsid - Ceph: Depends on the specific CSI driver, but often
ceph-commonor similar tools are required.
6. Review Volume Binding Mode and Pod Scheduling
If your StorageClass uses volumeBindingMode: WaitForFirstConsumer, the PVC will wait until a Pod that uses it is scheduled. If the Pod is stuck, the PVC will also be stuck.
a. Check Pod Status:
kubectl get pod -n <namespace> -o wide
kubectl describe pod <pod-name> -n <namespace>
Look for reasons why the Pod might not be scheduling, such as:
No nodes are available that match all of the following predicates...(resource constraints, node selectors, taints/tolerations)Pendingdue to an image pull error.
Resolve the Pod scheduling issue, and the PVC should bind.
7. Final Checks and Clean Up
- Recreate PVC/Pod: Sometimes, after fixing the underlying issue, simply deleting and recreating the PVC (and the dependent Pod) can help trigger a fresh binding attempt.
kubectl delete pvc <pvc-name> -n <namespace> kubectl delete pod <pod-name> -n <namespace> # If directly referencing the PVC # Or delete deployment/statefulset if managing the pod kubectl apply -f <original-pvc-manifest.yaml> kubectl apply -f <original-pod-or-deployment-manifest.yaml> - Kubernetes Control Plane Health: While rare, ensure the Kubernetes API server and controller manager are healthy.
This command is deprecated in newer Kubernetes versions, but checkingkubectl get componentstatuseskube-apiserver,kube-controller-manager, andkube-schedulerpods inkube-systemnamespace is still valid.
By meticulously following these steps, you should be able to identify and rectify the root cause of a Kubernetes PVC stuck in Pending on your CentOS Stream or Rocky Linux environment.
Our Production Verification Guarantee
Encountering a bug not covered here or running a non-standard kernel configuration? Our solutions are continually refined against real production incidents. Submit an environment trace for our editorial team to replicate.