APURV
  • Home
  • Journey
  • Projects
  • Blogs
  • Interview
  • Exams
Resume
APURV

Building scalable, secure, and production-ready cloud infrastructure. Automation first.

NAVIGATION

HomeExperienceProjectsCertificationsSkills

TECH STACK

AWSGCPK8sCI/CDLinuxDocker

CONNECT

LinkedInGitHubEmailResume

© 2026 Apurv Gujjar. All rights reserved.
Apurv Gujjar
Apurv GujjarDevOps & Cloud Engineer
|Interview Documentation
Portfolio
Handbooks
🎯Linux🐙Git & GitHub🤖GitHub Actions🌐Networking☁AWS🛠Terraform🐳Docker☸Kubernetes📊Monitoring🛡DevSecOps💰Cost Optimization🚨Incident Scenarios👤HR & Behavioral☁GCP🐍Python
Interview DocumentationMonitoring

Monitoring

Observability, Prometheus Metrics & Grafana Guide for Reliability Engineers

Q1

How do you explain the difference between Monitoring and Observability? Are they the same thing?

💬Answer
  • Monitoring: Focuses on known failure modes. It tells you when a system is broken by gathering predefined metrics (e.g., CPU > 90%, disk space full).
  • Observability: Focuses on unknown failure modes. It provides the context and raw telemetry (metrics, logs, and distributed traces) to help you understand why a system is broken, especially in complex, distributed microservice architectures.
Q2

What are the Four Golden Signals of Monitoring, and why are they critical?

Q3

How does the RED methodology help monitor request-driven microservices?

Q4

What is the USE methodology, and in what scenarios is it preferred over RED?

Q5

How do you define and contrast SLIs, SLOs, and SLAs? Who is the target audience for each?

Q6

What is an Error Budget, and how does it help balance velocity with system reliability?

Q7

What is Alert Fatigue, and what architectural strategies do you use to prevent it?

Q8

# Prevention:

Q9

What is the role of Alertmanager in a Prometheus monitoring stack?

Q10

Can you walk me through the architecture of Prometheus? How does it collect and store metrics?

Q11

How does Prometheus Service Discovery automatically discover scaling resources?

Q12

What are Prometheus Exporters, and when are they necessary?

Q13

What is Prometheus Federation, and in what environments would you implement it?

Q14

What is Prometheus Remote Write, and how does it facilitate long-term metric retention?

Q15

What is OpenTelemetry (OTel), and how does it standardize telemetry collection?

Q16

What is Jaeger, and how does it help troubleshoot microservices transactions?

Q17

What is Distributed Tracing, and how do spans connect to form a trace?

Q18

What is a Correlation ID, and how does it help reconstruct logs across multiple microservices?

Q19

What is Log Aggregation, and what are the main components of a modern log pipeline?

Q20

Why is Continuous Monitoring crucial in a modern DevOps lifecycle?

Q21

# Why it is Important:

Q22

If you are tasked with monitoring the health of a Kubernetes cluster, what layers and metrics would you target?

Q23

# 1. Cluster/Node Layer (Infrastructure)

Q24

# 2. Kubernetes Control Plane (Control Layer)

Q25

# 3. Pod/Application Layer (Workload Layer)

Q26

What is Prometheus, and can you describe its active scraping model?

Q27

# How it Works:

Q28

How would you design and implement a scalable logging architecture for a distributed system?

Q29

What key metrics would you monitor to measure the health and efficiency of a CI/CD build pipeline?

Apurv Gujjar - DevOps & Cloud Engineer
Created by

Apurv Gujjar

DevOps & Cloud Engineer

Specialized in:DevOpsAWSGCPKubernetesTerraformDocker
View Portfolio