As a NOC Analyst at Lightning AI, I work on infrastructure monitoring, telemetry, and incident response across next-generation compute clusters. My focus is turning system signals into actionable information for reliable operations.
Monitoring & telemetry
I participate in developing and scaling telemetry infrastructure, improving real-time data ingestion, log collection, and system visibility. I create and configure Grafana, Datadog, and Prometheus dashboards to detect performance anomalies and system constraints.
Diagnosis & incident response
- Troubleshoot Linux systems and hardware health through command-line tools, log analysis, and diagnostic utilities.
- Investigate network-layer vulnerabilities, routing anomalies, and interface errors using TCP/IP and DNS knowledge, alongside tcpdump, traceroute, and netstat.
- Triage incidents and escalate to SRE, hardware, and network engineering teams with precise logs and actionable findings.
Making the next response better
I document incidents and analyze alert trends to refine monitoring coverage and operational runbooks, helping teams respond with better context.