Search Jobs

Search by job, company or skills

Data Center Engineer

Data Center Engineer

Lintasarta
1-2 Years
  • Posted 7 hours ago
  • Be among the first 10 applicants

Job Description

About the Role :

The L1 Engineer forms the foundational layer of the tiered support model, operating across two complementary sub-functions within the same 24×7 onsite coverage model: Surveillance (NOC-based continuous monitoring) and Data Center (onsite physical response). Together, these two sub-functions ensure that every infrastructure event, whether detected remotely through dashboards or observed directly on the data center floor, is captured, classified, and escalated with precision and speed. L1 is the first point of contact for all incidents and the initiator of the escalation path to L2 and L3.

Responsibilities:

  • Maintains continuous 24×7 monitoring of all GPU infrastructure dashboards, covering compute health, GPU utilization, fabric and network status, power and cooling parameters, and environmental conditions, using platforms including DCGM, NetQ, UFM, Grafana, and ServiceNow.
  • Classifies all alarms by severity (P1–P4), validates against false-positive filters, and dispatches through the appropriate escalation workflow with full SLA tracking.
  • Creates accurate, complete incident tickets in the ITSM platform and dispatches to the appropriate tier (L1 DC for physical check, L2 for technical diagnosis, or L3 for complex escalation).
  • Provides P1 status updates every 30 minutes until resolution; monitors SLA countdown for all active tickets and proactively escalates tickets at risk of breach.
  • Conducts daily synthetic health checks: canary jobs, NCCL bandwidth tests, fabric monitoring, and log pipeline health verification.
  • Conducts regular physical walkthroughs and rack inspections, verifying LED indicators, cabling integrity, power supply status, and the physical condition of GPU servers, NVLink switches, and CDUs.
  • Responds to ticket dispatch from L1 Surveillance for direct on-floor physical checks, and performs basic hardware verification and initial corrective actions (e.g., reboot/power-cycle via BMC) before escalating to L2.
  • Executes Emergency Response Procedures (EPO or Loop Isolation protocols) within 15 minutes of alert confirmation for P1 environmental incidents including cooling failures, power irregularities, or liquid coolant leaks.
  • Monitors liquid-cooling system parameters (CDU and secondary loop) and coordinates with the DC facilities team on environmental anomalies.
  • Supports preventive maintenance activities: rack deep cleaning (quarterly), node health sweep (monthly), and cold-spare rotation.
  • Executes a structured shift handover at every transition, ensuring full situational awareness of active incidents, open tickets, pending dispatches, and infrastructure anomalies is formally transferred to the incoming shift with zero information loss.

Qualifications & Skills :

  • Min. 1–2 years in Data Center / NOC / IT Operations.
  • Basic Linux & BMC/IPMI access.
  • Basic TCP/IP networking.
  • Proficient with monitoring tools (Grafana / Zabbix / DCIM / NMS) & ITSM ticketing systems.
  • Solid understanding of SLA concepts, incident prioritization, and escalation workflows.
  • Language: English & Bahasa.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

TCP/IP networking

DCIM

ITSM ticketing systems

About Company

Similar Jobs

2-4 yrs
Indonesia
Skills:
containerization VMwareBackupWindows ServerRestorevirtualizationServer HardwareMonitoring ToolsOperating SystemsGrafanaStorageZabbixLinuxNetwork DevicesApmDisaster RecoveryArchiving
3-5 yrs
Indonesia
Skills:
MqttShellSNMPPythonBuilding AutomationBACnetSCADADCIMBmsPlcModbusDDCEPMS