🚀 Building Reliable Infrastructure

Hi, I'm Gagan Gandhi

A Site Reliability Engineer who transforms chaos into stability. With 9 months of hands-on experience, I've helped startups scale their infrastructure while cutting costs by 60%. I design systems that don't just work—they work flawlessly, even under pressure. Let's build something reliable together.

7K→2K
Alerts Reduced
60%
Cost Savings
10K
Concurrent Requests

About Me

I'm obsessed with building systems that are reliable, scalable, and cost-efficient. The kind of systems that let engineers sleep peacefully and allow companies to scale without fear.

My journey started with a simple question: "How do we make systems that never fail?" Now, as an SRE at FarEye, I'm answering it every single day. I've reduced alert noise by 71% (from 7,000 to 2,000), optimized cloud costs by 60%, and handled load testing for 10,000 concurrent users—all because I believe that reliability isn't a feature, it's a foundation.

Whether it's setting up Kubernetes clusters, designing CDC pipelines, implementing observability stacks, or automating away manual toil—I approach every challenge with one goal: make operations invisible and operations teams heroes. I love taking complex infrastructure problems and turning them into elegant, automated solutions.

When I'm not designing systems, you'll find me exploring new DevOps tools, contributing to open-source, or helping teams understand that SRE isn't just about firefighting—it's about building fire prevention into the foundation.

Professional Experience

Real-world impact: reducing operational chaos, scaling systems, and building tools that make engineers' lives better.

Site Reliability Engineer (SRE)
FarEye • Noida, India
March 2026 – Present
Cloud Engineer Intern
Mist Avinya • Remote
July 2025 – March 2026

Featured Projects

Production-grade systems that prove reliability engineering isn't just about theory—it's about building systems that actually work.

🏦 Banking Platform

Production-ready microservices banking system with real-time data synchronization, comprehensive monitoring, and tested for 10K concurrent users.

Kubernetes Docker AWS EKS Kafka CDC
  • Architecture: Deployed 5+ microservices on EKS with auto-healing and auto-scaling
  • Data Consistency: Built Debezium CDC pipeline streaming PostgreSQL changes to Kafka → Elasticsearch for live updates
  • Security: Multi-tier VPC setup with Jump Hosts, Security Groups, and IAM policies
  • Load Tested: Apache Bench with 10,000 concurrent requests—no drops, no timeouts
  • Observability: Prometheus + Grafana + custom metrics at service, pod, and container level

📈 Zerodha Trading Platform

Stock trading platform containerized and deployed on AWS with full CI/CD pipeline, monitoring, and infrastructure automation.

Docker AWS GitHub Actions RDS
  • Containerization: Multi-container setup with proper networking, volume management, and environment isolation
  • Infrastructure: EC2 instances, RDS database, S3 storage, and Load Balancing
  • CI/CD: Automated build, test, and deployment using GitHub Actions—from commit to production in minutes
  • Monitoring: Health checks, log aggregation, and alerting for production readiness
  • Scaling: Auto-scaling groups and load balancer configurations for traffic spikes
View Source Code →

📋 Project Info Prints

Automation tool that eliminates manual documentation effort and streamlines project onboarding at scale.

Python Automation DevOps
  • Problem Solved: Reduces hours of manual documentation to minutes—because humans shouldn't copy-paste
  • Automation: Scans projects, generates comprehensive documentation automatically
  • Integration: Works seamlessly with CI/CD pipelines for automated reporting on every commit
  • Scalability: Handles projects of any size—from solo apps to enterprise systems
  • Impact: Improves onboarding time and knowledge sharing across teams
View Source Code →

Technical Arsenal

The tools and technologies I use to build, scale, and maintain systems that work flawlessly at scale.

Cloud & Infrastructure

AWS (EKS, EC2, ECR, S3) Azure (AKS, ACR, VMs) VPC & Networking CloudWatch Route 53 Load Balancing

Containers & Orchestration

Docker Kubernetes (K8s) Helm Container Registry Pod Management YAML Configuration

Monitoring & Observability

Prometheus Grafana CloudWatch Squadcast Node Exporter cAdvisor

CI/CD & Automation

Git & GitHub GitHub Actions GitLab CI Pipeline Design Infrastructure as Code Automation Scripts

Data & Messaging

Apache Kafka Redis Debezium (CDC) PostgreSQL MongoDB Elasticsearch

Programming & Scripting

Python Bash/Shell Java JavaScript YAML Linux

Let's Connect

Ready to build reliable systems? Let's talk about your infrastructure challenges, DevOps practices, or how to make your systems bulletproof.

📧
Email
gagangandhi4080@gmail.com
📱
Phone
+91 99116 79155
📍
Location
Delhi, India