Showing posts with label Kuberneties (K8s). Show all posts
Showing posts with label Kuberneties (K8s). Show all posts

Thursday, December 4, 2025

Kubernetes: Complete Feature Summary for Executive Decision‑Makers

Kubernetes: Complete Feature Summary for Executive Decision‑Makers


Kubernetes is a full‑scale platform that modernizes how applications are deployed, scaled, secured, and operated. It delivers value across eight major capability areas, each directly tied to business outcomes.

1. Reliability & High Availability

Self‑healing containers
Automatic failover
Rolling updates & instant rollbacks
Health checks (liveness/readiness probes)
Multi‑node clustering
ReplicaSets for redundancy
Business impact: Keeps applications online, reduces outages, and improves customer experience.

2. Scalability & Performance

Horizontal Pod Autoscaling (HPA)
Vertical Pod Autoscaling (VPA)
Cluster Autoscaler
Built‑in load balancing
Resource quotas & limits
Business impact: Handles traffic spikes automatically and optimizes resource usage.

3. Security & Compliance

Role‑Based Access Control (RBAC)
Network Policies
Secrets encryption
Pod Security Standards
Image scanning & signing
Namespace isolation
Audit logging
Business impact: Strengthens security posture and supports compliance requirements.

4. Automation & DevOps Enablement

CI/CD integration
GitOps workflows
Automated deployments & rollbacks
Declarative configuration
Infrastructure as Code (IaC)
Business impact: Accelerates delivery, reduces manual errors, and standardizes operations.

5. Environment Standardization

Namespaces for dev/test/prod
Consistent container images
ConfigMaps & Secrets for environment configs
Multi‑OS container support (CentOS, Ubuntu, Debian, etc.)
Business impact: Eliminates “works on my machine” issues and improves developer productivity.

6. Cost Optimization

Efficient bin‑packing
Autoscaling to reduce idle resources
Spot instance support
Multi‑cloud flexibility
High container density
Business impact: Lowers infrastructure costs and prevents over‑provisioning.

7. Multi‑Cloud & Hybrid Cloud Flexibility

Runs on AWS, Azure, GCP, on‑prem, or hybrid
No vendor lock‑in
Disaster recovery across regions
Edge computing support
Business impact: Future‑proofs the organization and enables global deployments.

8. Observability & Monitoring

Metrics (Prometheus, Metrics Server)
Logging (ELK, Loki)
Tracing (Jaeger, OpenTelemetry)
Dashboards (Grafana, Lens)
Business impact: Improves visibility, speeds up troubleshooting, and supports data‑driven decisions.

Wednesday, November 26, 2025

Kubernetes Knowledge Transfer Pack

Kubernetes Knowledge Transfer Pack


1. Namespaces
Definition: Logical partitions in a cluster, used to separate environments or teams.
Commands:
kubectl get namespaces
kubectl create namespace dev-team
kubectl delete namespace dev-team

2. Pods
Definition: Smallest deployable unit in Kubernetes, wraps one or more containers.

Commands:
kubectl get pods
kubectl get pods --all-namespaces
kubectl describe pod <pod-name>
kubectl delete pod <pod-name>

3. Containers
Definition: Actual running processes inside pods (Docker/containerd images).

Commands:
kubectl logs <pod-name> -c <container-name>
kubectl exec -it <pod-name> -c <container-name> -- /bin/sh

4. Deployments
Definition: Controller that manages pods, scaling, and rolling updates.

Commands:
kubectl create deployment nginx-deploy --image=nginx
kubectl scale deployment nginx-deploy --replicas=5
kubectl get deployments
kubectl delete deployment nginx-deploy

5. Services
Definition: Provides stable networking to pods.
Types: ClusterIP, NodePort, LoadBalancer.

Commands:
kubectl expose deployment nginx-deploy --port=80 --target-port=80 --type=ClusterIP
kubectl get svc
kubectl delete svc nginx-deploy

6. ConfigMaps
Definition: Store non‑confidential configuration data.

Commands:
kubectl create configmap app-config --from-literal=ENV=prod
kubectl get configmaps
kubectl describe configmap app-config

7. Secrets
Definition: Store sensitive data (passwords, tokens).

Commands:
kubectl create secret generic db-secret --from-literal=DB_PASSWORD=banking123
kubectl get secrets
kubectl describe secret db-secret

8. Volumes & Storage
Definition: Persistent storage for pods.

Commands:
kubectl get pv
kubectl get pvc --all-namespaces

9. StatefulSets
Definition: Manage stateful apps (databases, Kafka).

Commands:
kubectl apply -f redis-statefulset.yaml
kubectl get statefulsets

10. DaemonSets
Definition: Ensures one pod runs on every node (logging, monitoring).

Commands:
kubectl get daemonsets -n kube-system

11. Jobs & CronJobs
Job: Runs pods until completion.
CronJob: Runs jobs on a schedule.

Commands:
kubectl create job pi --image=perl -- perl -Mbignum=bpi -wle 'print bpi(2000)'
kubectl get jobs
kubectl create cronjob hello --image=busybox --schedule="*/1 * * * *" -- echo "Hello World"
kubectl get cronjobs

12. Ingress
Definition: Manages external HTTP/HTTPS access to services.

Commands:
kubectl apply -f ingress.yaml
kubectl get ingress

🏗 Kubernetes Architecture

Control Plane Components

API Server → Entry point for all requests.

etcd → Cluster state database.

Controller Manager → Ensures desired state.

Scheduler → Assigns pods to nodes.

Node Components

Kubelet → Agent ensuring containers run.

Kube-proxy → Networking rules.

Container Runtime → Runs containers (Docker, containerd).

Add‑ons
CoreDNS → DNS service discovery.

CNI Plugin (Flannel/Calico) → Pod networking.

Metrics Server → Resource monitoring.

📊 Monitoring & Health Commands

kubectl get nodes -o wide
kubectl get pods --all-namespaces -w
kubectl get events --all-namespaces --sort-by=.metadata.creationTimestamp
kubectl top nodes
kubectl top pods
systemctl status kubelet
systemctl status containerd

Start and Stop Kubernetes Services

Start and Stop Kubernetes Services


Proper Stop/Start Cycle

1. Stop services:

sudo systemctl stop kubelet
sudo systemctl stop containerd

2. Verify stopped:

systemctl status kubelet
systemctl status containerd

How to Confirm They’re Really Dow

# Check kubelet process
ps -ef | grep kubelet

# Check container runtime process
ps -ef | grep containerd

# List running containers (if using containerd)
sudo crictl ps

# If using Docker runtime
sudo docker ps

3. Start services again:

sudo systemctl start containerd
sudo systemctl start kubelet

4. Verify Recovery

After a minute or two, check again:

kubectl get nodes
kubectl get pods -n kube-system
kubectl get componentstatuses


How to Build Multi‑Environment Pods with CentOS, Ubuntu, Debian, and More

How to Build Multi‑Environment Pods with CentOS, Ubuntu, Debian, and More


Step 1: Create Namespaces

kubectl create namespace mqmdev
kubectl create namespace mqmtest
kubectl create namespace mqmprod

Step 2: Create Pods (example with 7 OS containers)

Save YAML files (mqmdev.yaml, mqmtest.yaml, mqmprod.yaml) with multiple containers inside each pod.

Example for mqmtest:

apiVersion: v1
kind: Pod
metadata:
name: mqmtest-pod
namespace: mqmtest
spec:
containers:
- name: centos-container
image: centos:7
command: ["/bin/bash", "-c", "sleep infinity"]
- name: redhat-container
image: registry.access.redhat.com/ubi8/ubi
command: ["/bin/bash", "-c", "sleep infinity"]
- name: ubuntu-container
image: ubuntu:22.04
command: ["/bin/bash", "-c", "sleep infinity"]
- name: debian-container
image: debian:stable
command: ["/bin/bash", "-c", "sleep infinity"]
- name: fedora-container
image: fedora:latest
command: ["/bin/bash", "-c", "sleep infinity"]
- name: oraclelinux-container
image: oraclelinux:8
command: ["/bin/bash", "-c", "sleep infinity"]
- name: alpine-container
image: alpine:latest
command: ["/bin/sh", "-c", "sleep infinity"]

Apply:

kubectl apply -f mqmdev.yaml
kubectl apply -f mqmtest.yaml
kubectl apply -f mqmprod.yaml

Step 3: Check Pod Status

kubectl get pods -n mqmdev
kubectl get pods -n mqmtest
kubectl get pods -n mqmprod

Step 4: List Container Names Inside a Pod

kubectl get pod mqmtest-pod -n mqmtest -o jsonpath="{.spec.containers[*].name}"

Output example:

centos-container redhat-container ubuntu-container debian-container fedora-container oraclelinux-container alpine-container

Step 5: Connect to a Particular Container

Use kubectl exec with -c <container-name>:

# CentOS
kubectl exec -it mqmtest-pod -n mqmtest -c centos-container -- /bin/bash

# Red Hat UBI
kubectl exec -it mqmtest-pod -n mqmtest -c redhat-container -- /bin/bash

# Ubuntu
kubectl exec -it mqmtest-pod -n mqmtest -c ubuntu-container -- /bin/bash

# Debian
kubectl exec -it mqmtest-pod -n mqmtest -c debian-container -- /bin/bash

# Fedora
kubectl exec -it mqmtest-pod -n mqmtest -c fedora-container -- /bin/bash

# Oracle Linux
kubectl exec -it mqmtest-pod -n mqmtest -c oraclelinux-container -- /bin/bash

# Alpine (use sh instead of bash)
kubectl exec -it mqmtest-pod -n mqmtest -c alpine-container -- /bin/sh

Step 6: Verify OS Inside Container

Once inside, run:
cat /etc/os-release

This confirms which OS environment you’re connected to.

✅ Summary

Create namespaces → mqmdev, mqmtest, mqmprod.
Apply pod YAMLs with 7 containers each.
Check pod status → kubectl get pods -n <namespace>.
List container names → kubectl get pod <pod> -n <namespace> -o jsonpath=....
Connect to container → kubectl exec -it ... -c <container-name> -- /bin/bash.
Verify OS → cat /etc/os-release.

Nodes in Kubernetes: The Unsung Heroes of Container Orchestration

Nodes in Kubernetes: The Unsung Heroes of Container Orchestration


What is a Node in Kubernetes?
A Node is a worker machine in Kubernetes. It can be a physical server or a virtual machine in the cloud. Nodes are where your pods (and therefore your containers) actually run.

Think of it like this:

Cluster = a team of machines.
Node = one machine in that team.
Pod = a unit of work scheduled onto a node.
Container = the actual application process inside the pod.


🔹 Node Components
Each node runs several critical services:

Kubelet → Agent that talks to the control plane and ensures pods are running.
Container Runtime → Runs containers (Docker, containerd, CRI‑O).
Kube‑proxy → Manages networking rules so pods can communicate with each other and with services.

🔹 Types of Nodes

Control Plane Node → Runs cluster management components (API server, etcd, scheduler, controller manager).
Worker Node → Runs user workloads (pods and containers).

🔹 Example: Checking Nodes

When you run:

[root@centosmqm ~]# kubectl get nodes
NAME        STATUS   ROLES           AGE   VERSION
centosmqm   Ready    control-plane   28h   v1.30.14

Explanation of Output:
NAME → centosmqm → the hostname of your node.
STATUS → Ready → the node is healthy and can accept pods.
ROLES → control-plane → this node is acting as the master/control plane, not a worker.
AGE → 28h → the node has been part of the cluster for 28 hours.
VERSION → v1.30.14 → the Kubernetes version running on this node.

👉 In this example, your cluster currently has one node (centosmqm), and it is the control plane node. If you added worker nodes, they would also appear in this list with roles like <none> or worker.

🔹 Commands to Work with Nodes

# List all nodes
kubectl get nodes

# Detailed info about a node
kubectl describe node centosmqm

# Show nodes with more details (IP, OS, version)
kubectl get nodes -o wide

✅ Summary

A Node is the machine (VM or physical) that runs pods.
Nodes can be control plane (managing the cluster) or worker nodes (running workloads).
Your example shows a single control-plane node named centosmqm, which is healthy and running Kubernetes v1.30.14.

Monday, May 13, 2024

Mastering Etcd Backup and Restore in Kubernetes: A Step-by-Step Guide

Mastering Etcd Backup and Restore in Kubernetes: A Step-by-Step Guide

Kubernetes relies heavily on etcd as its primary storage backend to keep all its configuration data, state, and metadata. Here’s why backups and restores of etcd are crucial for maintaining the health and resilience of Kubernetes environments:

1. Critical Data Preservation

  • Single Point of Truth: etcd holds the entire cluster state including pods, services, and controller information. Losing etcd data means losing the entire cluster state.
  • Configuration Data: All resource configurations such as deployments, services, and network configurations are stored in etcd. Backing up etcd ensures that these configurations are not permanently lost in case of a disaster.

2. Disaster Recovery

  • Cluster Integrity: In the event of a physical disaster, software bug, or data corruption, a recent backup can be the fastest way to restore cluster operations without reconstructing configurations from scratch.
  • Data Corruption Recovery: If the etcd database becomes corrupted, the only way to restore its operation without losing all the cluster data is to restore it from a backup.

3. High Availability and Durability

  • Avoid Downtime: Regular backups help in minimizing downtime. In a high-availability setup, if one of the etcd nodes fails, etcd can still serve data from its other nodes. However, in catastrophic scenarios where multiple nodes are affected, backups are critical.
  • Redundancy: Regular backups contribute to the redundancy strategies, essential for business continuity and compliance with data protection regulations.

4. Versioning and Rollbacks

  • Cluster State Reversions: Backing up etcd allows administrators to rollback the cluster state to a previous point in time. This is vital for recovery from bad updates or configurations that may lead to system instability or downtime.
  • Audit and Compliance: For some organizations, maintaining a history of changes in etcd can be important for audit trails and compliance with regulatory requirements.

5. Operational Flexibility

  • Migration: Backups can facilitate the migration of Kubernetes clusters between different environments or cloud providers by ensuring that all etcd data can be consistently transferred.
  • Testing and Development: Backups of etcd data can be used to clone environments for testing and development without affecting the production environment.

6. Prevent Data Loss

  • Human Errors: Mistakes such as accidental deletions or misconfigurations can be mitigated by restoring data from backups.

Pre and Post Steps for etcd Backup

Pre-Backup Steps:

  1. Verify Cluster Health: Ensure that your etcd cluster is healthy before taking a backup. You can check the health of etcd with the following command:

ETCDCTL_API=3 etcdctl endpoint health --endpoints=https://192.168.163.134:2379 \ --cacert=/etc/kubernetes/pki/etcd/ca.crt \ --cert=/etc/kubernetes/pki/etcd/server.crt \ --key=/etc/kubernetes/pki/etcd/server.key

  1. Plan the Backup Window: Although etcd v3 supports live backups (without downtime), it's still prudent to plan backups during periods of low activity to minimize performance impact.

  2. Check Storage Space: Ensure that there is sufficient disk space where you plan to store the backup.

Post-Backup Steps:

  1. Verify the Backup: Check that the backup file is not corrupted and is of a reasonable size compared to previous backups.

ls -lh /path/to/your/backup/backup.db

  1. Secure the Backup: Move the backup to a secure, off-site storage location if possible. Ensure that the backup is encrypted if it’s stored off-cluster.

  2. Document the Backup: Record details about the backup such as the date, size, etcd version, and checksum for integrity.

Pre and Post Steps for etcd Restore

Pre-Restore Steps:

  1. Determine the Need for Restore: Understand why a restore is necessary (e.g., data corruption, loss, etc.) and determine if a restore is the best course of action.
  2. Notify Stakeholders: Inform all relevant stakeholders about the planned downtime and impact, as etcd restoration involves downtime.
  3. Prepare the Environment: Ensure the server where etcd will be restored has all necessary software installed and is configured similarly to the original etcd environment.
  4. Backup Current State: Before performing a restore, backup the current state of etcd even if it is believed to be corrupted. This step is crucial for having a fallback option.

Post-Restore Steps:

  1. Verify Cluster Functionality: After restoration, check all Kubernetes services and ensure that the cluster returns to its expected operational state.
  2. Test Workloads: Verify that critical workloads are functioning correctly and that data integrity is maintained post-restore.
  3. Monitor Performance: Observe the cluster's performance after the restore. Look for any unexpected behavior or errors in the logs.
  4. Document the Process: Record the details of the restore process and any issues encountered or lessons learned.

Important Points to Consider for Backup and Restore

  • Backup Type: etcd backups are online and do not require stopping the service. They can be done live while the cluster is running.
  • Restore Type: Restoring etcd is an offline process. The etcd service must be stopped, and the data directory must be replaced with the one from the backup.
  • Data Integrity: Use checksums to validate the integrity of the backup files before and after transport or storage.
  • Security: Use secure connections (TLS) for interactions with etcd during backup and restore. Ensure backup files are encrypted and securely stored.
  • Regular Testing: Regularly test your backup and restore procedures to ensure they work as expected. This is crucial for disaster recovery planning.

By following these detailed steps and considerations, you can effectively manage the backup and restore processes for etcd in your Kubernetes environment, ensuring that your cluster can be reliably recovered in case of an emergency.


Below is a concise summary of identifying etcd configuration, including data directory location, and steps for backup and restore on a single-node Kubernetes setup:

Identifying etcd Configuration and Data Directory


  1. Static Pod: Check /etc/kubernetes/manifests for etcd.yaml and look for the --data-dir argument within the command section.
cat /etc/kubernetes/manifests/etcd.yaml

2. System Service: Review the systemd service file for etcd.

systemctl cat etcd.service

Look for --data-dir in ExecStart in the output command.

3. Common Locations: Check typical directories like /var/lib/etcd or /var/lib/etcd/default.

4. Process Arguments: Use ps aux | grep etcd to view the running process and its arguments.

5. Kubernetes API Server: Inspect kube-apiserver.yaml for --etcd-servers to identify the etcd endpoint.

Detailed Backup Command Explanation:

ETCDCTL_API=3 etcdctl snapshot save /root/backup_db_mqm/backup1.db \ --endpoints=https://192.168.163.134:2379 \ --cacert=/etc/kubernetes/pki/etcd/ca.crt \ --cert=/etc/kubernetes/pki/etcd/server.crt \ --key=/etc/kubernetes/pki/etcd/server.key

Command Components:

  • ETCDCTL_API=3: This environment variable specifies that the etcdctl tool should use the v3 API, which is necessary for most modern etcd features, including snapshots.
  • etcdctl: The command line utility used for interacting with etcd.
  • snapshot save: The snapshot command manages snapshot files of etcd data. save is used to create a new snapshot.
  • /root/backup_db_mqm/backup1.db: The file path where the snapshot will be saved. This should be a secure location with adequate storage space.
  • --endpoints=https://192.168.163.134:2379: Specifies the etcd member to connect to. In this case, it's the local etcd instance running on the Kubernetes master node.
  • --cacert, --cert, --key: These options provide the paths to the certificate authority file, the certificate, and the key for secure communication with etcd over TLS.

Detailed Restore Command Explanation:

ETCDCTL_API=3 etcdctl snapshot restore /root/backup_db_mqm/backup1.db \ --data-dir /var/lib/etcd-from-backup \ --name master-node \ --initial-cluster master-node=https://192.168.163.134:2380 \ --initial-cluster-token etcd-cluster-1 \ --initial-advertise-peer-urls https://192.168.163.134:2380

Command Components:

  • snapshot restore: This command restores an etcd member’s state from a saved snapshot.
  • /root/backup_db_mqm/backup1.db: The path to the snapshot file that will be used for restoration.
  • --data-dir /var/lib/etcd-from-backup: Specifies the directory to store the etcd state after the restore. This should be different from the current data directory to avoid overwriting live data during testing.
  • --name master-node: The name for the etcd member, which should match the name used when the etcd cluster was first initialized.
  • --initial-cluster: Configuration for the etcd cluster. This should list all member names and their peer URLs. For a restore, typically, you'll use the same initial cluster configuration as before unless you're changing the topology.
  • --initial-cluster-token: A new cluster token to ensure that the restored etcd nodes join a new cluster instance, preventing conflicts with any existing etcd clusters.
  • --initial-advertise-peer-urls: The URL that this member will use to communicate with other etcd nodes in the cluster. This must be reachable from the other etcd nodes.

Points to Consider:

  • Backup Regularity: It's important to perform backups regularly, based on how frequently your data changes. The frequency of backups will influence your recovery point objective (RPO).
  • Secure Backup Storage: Store backups in a secure, offsite location to protect against data loss scenarios such as data center failures.
  • Validate Backups: Regularly validate the integrity of backups by performing test restores to ensure that your backup files are not corrupted and can be successfully restored.
  • Document Procedures: Maintain detailed documentation of your backup and restore procedures, including command syntax and operational considerations, to ensure that team members can perform recoveries during critical incidents.

By understanding each part of these commands and considering the operational best practices around them, you can effectively manage and protect your Kubernetes cluster's state data stored in etcd.