Ensiklopedia VibeKoding: Principles of Monitoring, Logging, and Alerting.Ensiklopedia VibeKoding: Principles of Monitoring, Logging, and Alerting.
> ๐ก Learning Guide: This chapter requires no programming background. Through interactive demos, you'll gain a comprehensive understanding of operations โ from monitoring and alerting to troubleshooting, from capacity planning to automated operations, mastering all the skills needed to run production systems.> ๐ก Learning Guide: This chapter requires no programming background. Through interactive demos, you'll gain a comprehensive understanding of operations โ from monitoring and alerting to troubleshooting, from capacity planning to automated operations, mastering all the skills needed to run production systems.
Many beginners think: "Once the code is deployed, the job is done."Many beginners think: "Once the code is deployed, the job is done."
That couldn't be more wrong!That couldn't be more wrong!
Deployment is merely the starting point of operations work. It's like buying a new car โ the real work of maintenance, repairs, and refueling is what follows.Deployment is merely the starting point of operations work. It's like buying a new car โ the real work of maintenance, repairs, and refueling is what follows.
Operations has three goals:Operations has three goals:
------
Monitoring is the "eyes" of operations. A system without monitoring is like driving blind โ you won't even know when something goes wrong.Monitoring is the "eyes" of operations. A system without monitoring is like driving blind โ you won't even know when something goes wrong.
Infrastructure Monitoring: Tracking server hardware resourcesInfrastructure Monitoring: Tracking server hardware resources
Application Monitoring: Tracking software runtime stateApplication Monitoring: Tracking software runtime state
Business Monitoring: Tracking business healthBusiness Monitoring: Tracking business health
| Tool | Purpose | Characteristics |
|---|---|---|
| Prometheus | Metric collection & storage | Time-series database, ideal for monitoring data |
| Grafana | Visualization dashboards | Powerful charts and dashboards |
| Zabbix | Comprehensive monitoring | Veteran tool with full-featured capabilities |
| Datadog | SaaS monitoring platform | One-stop solution, paid |
Key Point: Monitoring must be layered, covering everything from infrastructure to business to avoid blind spots.Key Point: Monitoring must be layered, covering everything from infrastructure to business to avoid blind spots.
------
Once monitoring detects an issue, operations staff need to be notified promptly โ that's alerting.Once monitoring detects an issue, operations staff need to be notified promptly โ that's alerting.
Proper alert classification helps prevent "alert fatigue":Proper alert classification helps prevent "alert fatigue":
| Level | Response Time | Typical Scenario | Notification Channels |
|---|---|---|---|
| P0 | Immediate (within 5 min) | Core service down, payment failures | Phone + SMS + IM |
| P1 | Within 30 minutes | Partial feature outage, severe performance degradation | SMS + IM + Email |
| P2 | Same day | High resource usage, occasional errors | IM + Email |
| P3 | Within the week | Non-critical issues, optimization suggestions |
Pain Point: A single small issue can trigger hundreds or thousands of alerts, numbing on-call staff.Pain Point: A single small issue can trigger hundreds or thousands of alerts, numbing on-call staff.
Solutions:Solutions:
Key Point: Alerts should be "few but meaningful" โ every alert must be worth acting on.Key Point: Alerts should be "few but meaningful" โ every alert must be worth acting on.
------
Logs are the "black box" for troubleshooting.Logs are the "black box" for troubleshooting.
javascript console.debug('Verbose debug info') // Used during development console.info('General information') // Normal flow logging console.warn('Warning') // Potential issues console.error('Error') // Errors that need attention
Traditional logging (not ideal):Traditional logging (not ideal):
CODE 2024-01-15 10:23:45 ERROR User john failed to login, attempts=3, ip=192.168.1.100
Structured logging (recommended):Structured logging (recommended):
json { "timestamp": "2024-01-15T10:23:45Z", "level": "ERROR", "message": "User login failed", "user": "john", "attempts": 3, "ip": "192.168.1.100", "service": "auth-service" }
ELK = Elasticsearch + Logstash + KibanaELK = Elasticsearch + Logstash + Kibana
Best Practices:Best Practices:
------
In a microservices architecture, a single request may pass through dozens of services โ how do you trace its complete path?In a microservices architecture, a single request may pass through dozens of services โ how do you trace its complete path?
Trace ID and Span IDTrace ID and Span ID
OpenTelemetry (OTel) is the industry standard for distributed tracing, providing a unified API and SDK.OpenTelemetry (OTel) is the industry standard for distributed tracing, providing a unified API and SDK.
javascript // Example: Recording a Span with OpenTelemetry import { trace } from '@opentelemetry/api' const tracer = trace.getTracer('my-service') async function processOrder(orderId) { // Create a Span const span = tracer.startSpan('processOrder') try { // Set attributes span.setAttribute('order.id', orderId) // Business logic... await validateOrder(orderId) await saveToDatabase(orderId) span.setStatus({ code: SpanStatusCode.OK }) } catch (error) { span.recordException(error) span.setStatus({ code: SpanStatusCode.ERROR, message: error.message }) } finally { span.end() // End the Span } }
Key Point: Distributed tracing quickly identifies performance bottlenecks and failure points โ an essential tool for microservices.Key Point: Distributed tracing quickly identifies performance bottlenecks and failure points โ an essential tool for microservices.
------
Production incidents are inevitable. The key is fast response and fast recovery.Production incidents are inevitable. The key is fast response and fast recovery.
| Tool | Purpose | Typical Scenario |
|---|---|---|
| tcpdump | Packet capture analysis | Network issues, packet loss |
| strace | System call tracing | Process hanging, file permission issues |
| Arthas | Java diagnostics | CPU spikes, memory leaks, deadlocks |
| top/htop | System resource monitoring | High CPU/memory usage |
| netstat | Network connection inspection | Port conflicts, abnormal connection counts |
| lsof | Open file inspection | File locks, disk full |
Arthas Example (Alibaba's open-source Java diagnostic tool):Arthas Example (Alibaba's open-source Java diagnostic tool):
bash # View top 5 threads by CPU usage $ top -H -p 12345 # Trace the execution time of a method $ trace com.example.OrderService createOrder # View a class's static fields $ getstatic com.example.Config MAX_CONNECTIONS # Hot-reload code (no restart needed) $ mc /tmp/Test.java $ redefine /tmp/Test.class
A post-mortem is not a blame session!A post-mortem is not a blame session!
The purpose of a post-mortem is:The purpose of a post-mortem is:
The 5 Whys Analysis:The 5 Whys Analysis:
Ask "why" at least 5 times to find the root cause:Ask "why" at least 5 times to find the root cause:
Key Point: Build a blameless culture โ focus on process improvement, not individual accountability.Key Point: Build a blameless culture โ focus on process improvement, not individual accountability.
------
Top-down optimization approach:Top-down optimization approach:
CODE User Experience โ Frontend Optimization (reduce requests, CDN, lazy loading) โ Network Optimization (HTTP/2, compression, persistent connections) โ Backend Optimization (caching, async, batching) โ Database Optimization (indexes, query tuning, sharding) โ System Optimization (kernel parameters, JVM tuning)
Index Optimization:Index Optimization:
sql -- Slow query (no index) SELECT * FROM orders WHERE user_id = 12345; -- 100x faster after creating an index CREATE INDEX idx_user_id ON orders(user_id);
Query Optimization:Query Optimization:
sql -- โ Avoid SELECT * SELECT * FROM users WHERE id = 123; -- โ Only query needed fields SELECT id, name, email FROM users WHERE id = 123; -- โ Avoid overly large IN clauses SELECT * FROM orders WHERE user_id IN (1, 2, 3, ..., 10000); -- โ Use JOIN or batch queries SELECT * FROM orders o JOIN user_ids u ON o.user_id = u.id;
Multi-level Cache Architecture:Multi-level Cache Architecture:
CODE Browser Cache (CDN) โ Local Cache (In-memory/Guava) โ Distributed Cache (Redis/Memcached) โ Database (MySQL/PostgreSQL)
Cache Update Strategies:Cache Update Strategies:
| Strategy | Pros | Cons | Use Case |
|---|---|---|---|
| Cache-Aside | Simple, reliable | Slow on first query | Read-heavy, write-light |
| Write-Through | Good data consistency | Slow writes | Balanced read/write |
| Write-Behind | Extremely fast writes | Potential data loss | Write-heavy, tolerates brief inconsistency |
Key Point: Caching is not a silver bullet โ consider consistency, avalanche, and penetration issues (refer to the "System Cache Design" chapter).Key Point: Caching is not a silver bullet โ consider consistency, avalanche, and penetration issues (refer to the "System Cache Design" chapter).
------
Tool Selection:Tool Selection:
| Tool | Characteristics | Use Case |
|---|---|---|
| JMeter | Feature-rich, visual | HTTP API stress testing |
| wrk/ab | Lightweight, command-line | Quick benchmarking |
| Locust | Python scripting, distributed | Complex scenario testing |
| K6 | Modern, JS scripting | CI/CD integration |
wrk Example:wrk Example:
bash # Install wrk $ brew install wrk # macOS $ apt install wrk # Ubuntu # Stress test an HTTP endpoint (10 threads, 30 seconds) $ wrk -t10 -c100 -d30s http://example.com/api/users # Output: # Running 30s test @ http://example.com/api/users # 10 threads and 100 connections # Thread Stats Avg Stdev Max +/- Stdev # Latency 45.32ms 12.45ms 120.50ms 87.56% # Req/Sec 2.12k 123.45 3.45k 89.01% # 632450 requests in 30.00s, 1.23GB read # Requests/sec: 21081.67
Auto-scaling in the cloud-native era:Auto-scaling in the cloud-native era:
yaml # Kubernetes HPA (Horizontal Pod Autoscaler) apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: my-app-hpa spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: my-app minReplicas: 2 maxReplicas: 10 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 70
When CPU usage exceeds 70%, pods automatically scale up (up to 10)When CPU usage exceeds 70%, pods automatically scale up (up to 10)
Key Point: Combine business forecasting (e.g., Black Friday sales) with proactive scaling to avoid last-minute scrambling.Key Point: Combine business forecasting (e.g., Black Friday sales) with proactive scaling to avoid last-minute scrambling.
------
Principle of Least Privilege:Principle of Least Privilege:
Jump Server (Bastion Host):Jump Server (Bastion Host):
All operations tasks go through the bastion host, which records complete operation logs.All operations tasks go through the bastion host, which records complete operation logs.
The 3-2-1 Backup Rule:The 3-2-1 Backup Rule:
Backup Strategies:Backup Strategies:
| Type | Frequency | Retention | RTO | RPO |
|---|---|---|---|---|
| Full Backup | Weekly | 1 month | 4 hours | 24 hours |
| Incremental Backup | Daily | 1 week | 2 hours | 1 hour |
| Real-time Backup | Per second | 7 days | Minutes | Seconds |
RTO (Recovery Time Objective): The maximum acceptable downtime durationRTO (Recovery Time Objective): The maximum acceptable downtime duration
RPO (Recovery Point Objective): The maximum acceptable data lossRPO (Recovery Point Objective): The maximum acceptable data loss
Regular Scanning:Regular Scanning:
bash # npm audit example $ npm audit found 3 vulnerabilities (1 moderate, 2 high) Package Severity Vulnerable versions lodash high <4.17.21 express moderate 4.0.0 - 4.18.2 # Auto-fix $ npm audit fix
------
yaml # .gitlab-ci.yml example stages: - test - build - deploy test: stage: test script: - npm install - npm test tags: - docker build: stage: build script: - docker build -t myapp:$CI_COMMIT_SHA . - docker push registry.example.com/myapp:$CI_COMMIT_SHA only: - main deploy: stage: deploy script: - kubectl set image deployment/myapp myapp=registry.example.com/myapp:$CI_COMMIT_SHA environment: name: production when: manual # Manually triggered deployment
Terraform Example (managing cloud resources):Terraform Example (managing cloud resources):
hcl # main.tf resource "aws_instance" "web" { ami = "ami-0c55b159cbfafe1f0" instance_type = "t2.micro" tags = { Name = "WebServer" Env = "production" } } resource "aws_security_group" "web" { name = "web-sg" ingress { from_port = 80 to_port = 80 protocol = "tcp" cidr_blocks = ["0.0.0.0/0"] } }
Advantages:Advantages:
GitOps = Git + IaC + AutomationGitOps = Git + IaC + Automation
Core principle: The Git repository is the single source of truth for infrastructureCore principle: The Git repository is the single source of truth for infrastructure
Workflow:Workflow:
CODE 1. Modify config files (push to Git) โ 2. Git repository changes trigger CI/CD โ 3. Automatically run terraform apply / kubectl apply โ 4. Infrastructure updates automatically โ 5. Monitor and reconcile actual state vs. desired state
Tools: ArgoCD, Flux (Kubernetes deployment)Tools: ArgoCD, Flux (Kubernetes deployment)
------
Operations is a vast domain, but the core can be distilled into the following:Operations is a vast domain, but the core can be distilled into the following:
| Level | Characteristics | Practices |
|---|---|---|
| Beginner | Reactive, manual operations | Fix issues only when they arise, manual deploys |
| Intermediate | Automated, standardized | CI/CD, monitoring & alerting, documentation |
| Advanced | Proactive, self-healing | Capacity planning, chaos drills, auto-scaling |
| Expert | Intelligent, unattended | AIOps, chaos engineering, serverless |
CODE 09:00 - Review overnight alerts, confirm system status 10:00 - Handle user-reported issues 11:00 - Attend engineering weekly, assess operational risk of new proposals 14:00 - Optimize slow queries, improve performance 15:00 - Code review 16:00 - Write deployment docs, update monitoring rules 17:00 - Chaos engineering drills 18:00 - On-call handoff
Beginner Stage (1โ3 months):Beginner Stage (1โ3 months):
Intermediate Stage (3โ6 months):Intermediate Stage (3โ6 months):
Advanced Stage (6โ12 months):Advanced Stage (6โ12 months):
Expert Stage (1+ year):Expert Stage (1+ year):
------
| Term | Full Name | Explanation |
|---|---|---|
| Monitoring | - | Real-time observation of system health. |
| Alerting | - | Notifying relevant personnel when anomalies occur. |
| Logging | - | Recording events during system operation. |
| Tracing | - | Tracking the full path of a request across a distributed system. |
| QPS | Queries Per Second | Queries per second, a measure of system throughput. |
| Latency | - | The time from request initiation to response. |
| RTO | Recovery Time Objective | Maximum acceptable downtime duration. |
| RPO | Recovery Point Objective | Maximum acceptable data loss. |
| Post-mortem | - | Incident review to analyze root causes and improvement actions. |
| CI/CD | Continuous Integration/Delivery | Automated testing and deployment. |
| IaC | Infrastructure as Code | Managing servers, networks, and other resources via code. |
| GitOps | - | Git-driven operations โ Git is the single source of truth. |
| ELK | Elasticsearch + Logstash + Kibana | The log collection, storage, and visualization trifecta. |
| SLA | Service Level Agreement | Committed service availability (e.g., 99.9%). |
| Blameless | - | A no-blame culture where post-mortems focus on process over individuals. |
------