Blog
1844+ articles — Page 74 of 77
Platform Engineering
Monitoring Slurm GPU Clusters with Prometheus
Set up Prometheus and Grafana monitoring for Slurm clusters with NVIDIA DCGM, job metrics, and queue utilization dashboards.
4 min read Platform Engineering
Slurm Multi-Node Distributed AI Training
How to run distributed PyTorch and DeepSpeed training across multiple GPU nodes using Slurm with NCCL, InfiniBand, and fault tolerance.
5 min read Platform Engineering
Slurm with Pyxis and Enroot for GPU Containers
Run containerized AI workloads on Slurm using NVIDIA Pyxis and Enroot. Faster than Docker, native GPU support, no daemon.
4 min read DevOps
SOC 2 Compliance for Cloud-Native Applications Guide
A practical engineering guide to SOC 2 compliance for Kubernetes-based applications. Automate evidence collection, implement controls, and pass audits.
2 min read DevOps
Sovereign Cloud: Building EU-Compliant Infrastructure...
EU data sovereignty requirements are reshaping cloud architecture. Practical patterns for multi-region deployments, data residency, and regulatory compliance.
5 min read DevOps
Supply Chain Security: SBOMs, Sigstore, and SLSA in Practice
Software supply chain attacks are surging. Here's how to implement SBOMs, container signing with Sigstore, and SLSA compliance in your CI/CD pipeline.
3 min read DevOps
Supply Chain Security in 2026: SLSA, Sigstore, and...
Secure your software supply chain with SLSA levels, Sigstore signing, and SBOM generation. Practical implementation for container-based workflows.
2 min read Platform Engineering
Sustainable Computing
AI workloads are exploding energy consumption. Practical strategies for carbon-aware scheduling, right-sizing GPU instances, and carbon measurement.
3 min read Automation
Testing Ansible Automation with Molecule and GitHub Actions
Comprehensive testing strategy for Ansible content. Unit tests with Molecule, integration tests in containers, and CI/CD with GitHub Actions.
2 min read AI
Troubleshooting OpenClaw Docker Deployments: Common...
A field-tested troubleshooting guide for OpenClaw Docker deployments covering networking, volumes, permissions, health checks, and container lifecycle.
5 min read Platform Engineering
WebAssembly on Kubernetes: SpinKube and Wasm
Deploy WebAssembly workloads alongside containers on Kubernetes using SpinKube. Faster cold starts, smaller footprint, and polyglot serverless functions.
3 min read Platform Engineering
WebAssembly Beyond the Browser
Wasm components boot in microseconds and consume minimal memory. How WebAssembly is reshaping serverless, edge computing, and plugin architectures.
3 min read AI
What is OpenClaw? The Open-Source AI Agent Gateway You...
Discover OpenClaw, the open-source AI agent gateway that connects LLMs to messaging platforms like Discord, Telegram, and Slack.
3 min read Platform Engineering
Zero Trust Architecture for Kubernetes Workloads
Implement zero trust security for Kubernetes. mTLS with service mesh, workload identity, network policies, and runtime security enforcement.
2 min read AI
Containerized AI Workloads with Podman on RHEL
Deploy and manage AI models using Podman containers on Red Hat Enterprise Linux—including GPU passthrough, rootless containers, and production.
6 min read AI
Multi-Tenant GPU Orchestration on OpenShift AI
A preview of my KubeCon Europe 2026 talk — what I'll cover about multi-tenant GPU orchestration on OpenShift AI with NVIDIA KAI (G/H200), and why this topic.
5 min read Conferences
Red Hat Summit 2026: GPU Platform Engineering Talk
Luca Berton presents 'GPUs take flight: Safety-first multi-tenant Platform Engineering with NVIDIA and Red Hat OpenShift AI' — a lightning talk at Red Hat.
5 min read AI
GPU Hardware Selection Guide for RHEL AI
Compare NVIDIA A100, H100, and AMD MI300X GPUs for RHEL AI workloads. Performance benchmarks, cost analysis, and deployment recommendations for each tier.
6 min read Conferences
KubeCon Europe 2026 Side Events Guide
Your ultimate guide to KubeCon + CloudNativeCon Europe 2026 side events, parties, and meetups in Amsterdam — from Cloud Native Rejekts to KuBBBecon.
45 min read Open Source
FOSDEM 2026: Open Source, RISC-V & Community Energy in...
Luca Berton's experience at FOSDEM 2026 in Brussels — connecting with DeepComputing, RISC-V International Foundation, and 8,000+ open source minds.
5 min read Platform Engineering
Manage Schema Evolution in Real‑Time Data
Master schema evolution strategies for real-time data systems in this hands-on Coursera course. Handle breaking changes without pipeline downtime.
2 min read Platform Engineering
Orchestrate & Recover Real-Time Data Pipelines
Master data pipeline orchestration and recovery in this hands-on Coursera course. Build fault-tolerant streaming architectures with Apache tools.
2 min read Platform Engineering
Stream & Unify Data Schemas with CDC
Master Change Data Capture techniques in this hands-on Coursera course. Stream, transform, and unify data schemas across heterogeneous systems.
2 min read AI
Ansible Automation for RHEL AI Deployments
Automate your entire RHEL AI lifecycle with Ansible playbooks—from GPU provisioning to model deployment, monitoring setup, and CI/CD pipeline integration.
6 min read