OpenAI
Network Operations Engineer, AI Networking
About this role
OpenAI seeks a Network Operations Engineer to operate and optimize large-scale Ethernet fabrics supporting GPU clusters across global data centers. You'll handle production incidents, infrastructure maintenance, deployments, and automation while partnering with multiple teams to ensure high availability of AI training and inference networks.
What you'll do
- Monitor, troubleshoot, and resolve network incidents while meeting SLOs and minimizing MTTD/MTTR
- Operate large-scale Ethernet fabrics supporting GPU compute, storage, and management networks across 1P and 3P data centers
- Execute production changes, maintenance windows, and capacity expansions with minimal impact
- Manage hardware lifecycle including switch/optics replacements, RMA coordination, and software upgrades
- Support AI cluster deployments, data center expansions, and infrastructure migrations
- Build monitoring, telemetry, dashboards, and automation to improve observability and reduce operational toil
What they're looking for
- Large-scale data center or cloud network operations (5+ years)
- Cisco NX-OS, Arista EOS, NVIDIA Spectrum/Cumulus Linux, or Juniper JunOS
- Layer 2/3 networking fundamentals (BGP, OSPF, ECMP, MLAG, LACP, VRFs, VLANs)
- Physical infrastructure troubleshooting (fiber optics, transceivers, high-speed Ethernet)
- Python scripting and infrastructure automation
- RCA and incident response methodologies
- AI/HPC networking environments (preferred)
- RoCE v2, RDMA, PFC, ECN, QoS, and lossless Ethernet (preferred)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe your experience operating production Ethernet fabrics at scale—what was the largest deployment you've worked with?
- Walk us through a complex network incident you debugged: what was the root cause and how did you resolve it?