About

Career progression and operating style.

Background, work history, and engineering preferences in one place.

Shubham Sharma
Author & Reliability SpecialistVerified

Shubham Sharma

Senior Platform & Reliability Engineer with 6+ years specializing in Kubernetes cluster internals, CloudNativePG, Longhorn CSI storage optimization, and enterprise OpenTelemetry/eBPF runtime telemetry.

Independent Research & Editorial ScopeAll architectural case studies, performance analyses, and reliability postmortems published on this site are independent technical research, synthetic archetypes, or based on personal platform experimentation (nucops.com). They do not represent, disclose, or reflect the internal architecture, proprietary systems, production incidents, or security policies of any employer or commercial client.
Current lens

Platform engineering with frontend restraint.

RoleSenior Platform & Reliability Engineer
Core SpecializationKubernetes Internals, Storage CSI, Telemetry & Observability
Current PlatformTalos Linux + Proxmox VE Multi-Node Baremetal Cluster
Resume notes

Context that explains the current operating taste.

Industry ExperienceBanking & High-Throughput

Authored telemetry and runtime systems running across top Indian banking infrastructures where stability and ultra-low latency are non-negotiable.

Homelab & SRE150+ Pod Bare-Metal Cluster

Self-host and operate a multi-node production Kubernetes cluster built on Talos Linux, automated via GitOps, and observed with end-to-end OpenTelemetry.

EducationB.Tech Computer Science

B.Tech in Computer Science from HBTU Kanpur and diploma in Computer Engineering from AMU Aligarh.

Timeline

Experience across observability, product engineering, and platform systems.

A concise view of roles and shipped outcomes.

Sep 2019 – Present

VuNet Systems

Senior Software Engineer (Platform & Observability)

  • Architected low-overhead CLR bytecode instrumentation and native .NET profilers capturing runtime database queries, method arguments, and payloads across tier-1 Indian financial platforms.
  • Engineered an enterprise Browser Real User Monitoring (BRUM) SDK with rrweb for DOM session recording, Core Web Vitals extraction, and automated PII masking.
  • Designed zero-code auto-injection pipelines with OpenTelemetry Operator, Beyla eBPF, and Traefik sidecars across Kubernetes clusters.
  • Honored with VuNet Innovation Award (Q2 2021) and Teamwork Award (Q1 2022).
2020 – Present

Production Homelab & SRE Platform

Platform Architect & Operations Lead

  • Operate a multi-node Talos Linux and Proxmox VE Kubernetes platform hosting 150+ pods managed via GitOps (Flux CD + Kustomize).
  • Architected zero-downtime migration of 10+ CloudNativePG PostgreSQL clusters to single-replica Longhorn with Barman S3 WAL archiving, slashing commit latency by 45%.
  • Authored Kyverno admission webhooks enforcing pod security standards, strict resource governance, and automated durability defaults.
  • Built end-to-end OpenTelemetry pipelines, ClickHouse distributed storage, and 50+ custom Grafana dashboards.
Jan 2019 – Sep 2019

CubeRoot Edu Labs

Co-Founder & Lead Engineer

  • Developed full-stack web applications and analytics platforms using React, Node.js, and relational database schemas.
  • Built reliable hosting infrastructure and CI/CD pipelines for early-stage interactive education portals.
Preferred stack

Go for durable tooling, TypeScript for refined surfaces.

Kubernetes & SRE

Talos LinuxCloudNativePGLonghorn CSIFlux CDKyvernoCilium CNITraefik IngressKube-VIP

Observability & Telemetry

OpenTelemetryPrometheus & ThanosGrafana DashboardsClickHouserrweb Session ReplayCore Web VitalseBPF (Beyla)Alertmanager

Programming & Runtimes

TypeScriptReact 19C# / .NET CoreGoC++ (CLR Profilers)PythonBash / ShellPostgreSQL
Operating principles

The same platform rules still guide the work.

01

Decouple Storage from Compute

Don't pay a double-replication tax. Let application engines handle transactional consensus and replication, and let block storage stay lightweight and performant.

02

Deterministic Failure Testing Over Hope

Test every failure mode directly—split-brain replicas, OOMKilled cgroups, and network partitions—before software reaches high-throughput production.

03

Declarative Governance Over Manual Tweaks

Enforce security standards, resource bounds, and automated mutations through admission policies (Kyverno) rather than fragile shell commands.

04

Observability Built as a Platform Default

Telemetry, metrics, trace propagation, and session replay must be integrated at the runtime layer, never treated as an afterthought.

Continue exploring

Continue into systems or start a direct conversation.