Oracle

Principal Systems Engineer

Nashville, Tennessee · Posted today

Opens eeho.fa.us2.oraclecloud.com

Get a version of your resume written for this job.

Salary
Not listed
Job type
Not specified
Work mode
Not specified
Source
Oracle (employer's hiring system)

Skills mentioned

Linux, Cybersecurity, Incident Response, Python, Bash, Grafana, Distributed Systems, Account Management

About the role

We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Systems Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational tooling for GPU fleets across multiple regions, ensuring high availability. You will collaborate closely with engineering and operations teams to build robust automation, observability, and reliability solutions while continuously improving GPU operations.

  • Build and maintain automation and operational tooling for OCI GPU infrastructure across multiple geographic regions.

  • Drive collaboration with software engineers, hardware teams, and operations partners to maintain a highly available GPU fleet.

  • Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using tools such as Grafana.

  • Serve as the senior escalation point for complex GPU host and repair issues

  • Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.

  • Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.

  • Participate in on-call rotations and provide support for critical infrastructure issues.

  • Document operational procedures, automation workflows, troubleshooting guides, and runbooks.

  • Build and improve AI agents, ensuring safe rollout, execution and monitoring.

  • Mentor and guide junior engineers in operational best practices, provide senior technical support, and drive constant improvement.

Required Qualifications:  

  • 8+ years of software operations or infrastructure automation experience with strong proficiency in Python and Bash.

  • Expert Linux administration experience, particularly Ubuntu and Oracle Linux, in large-scale production environments.

  • Strong understanding of distributed systems, including peer-to-peer, node-to-node, and service-to-service communication patterns.

  • Strong data-center and host-lifecycle experience, including provisioning, validation, repair workflows, hardware replacement, and fleet recovery.

  • Strong problem-solving and troubleshooting skills.

  • Excellent communication and teamwork skills.

  • Experience with observability tooling, including metrics, logging, dashboards, and alerting.

  • Experience with AI agents and tooling

  • Experience leading on-call operations and incident response.

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

Preferred Skills:

  • Hands-on experience with GPU infrastructure, including NVIDIA and AMD based systems, and familiarity with GPU shapes or instance families.

  • Experience operating or automating GPU, compute, or other large-scale cloud infrastructure.

Minimum Job Qualifications
Education and/or Experience:
11 years of experience in computer network administration, database management, systems architecture, systems administration, or related field

OR

Bachelor's Degree in information technology, computer science, engineering or related field AND 7 years of experience in computer network administration, database management, systems architecture, systems administration, or related field

OR

Master’s Degree in information technology, computer science, engineering or related field AND 5 years of experience in computer network administration, database management, systems architecture, systems administration, or related field.

Job Skills:
Same skills as prior level plus;
Cybersecurity Trends Demonstrated ability in or knowledge of cybersecurity trends, including staying current with industry threats, best practices, and emerging technologies.
Incident Management and Response Demonstrated ability in or knowledge of incident management and response, including timely handling and escalation of incidents to minimize business impact.
Training and Development Demonstrated ability to design and deliver effective training programs to build team and individual capabilities.
Technical Account Management Demonstrated ability to manage technical relationships and ensure successful outcomes with key accounts.
Cloud Architecture Demonstrated ability in or knowledge of cloud architecture, including designing scalable, reliable, and performant cloud services.

Preferred Job Qualifications
Education and/or Experience:
12 years of experience in computer network administration, database management, systems architecture, systems administration, or related field

OR

Bachelor's Degree in information technology, computer science, engineering or related field AND 8 years of experience in computer network administration, database management, systems architecture, systems administration, or related field

OR

Master’s Degree in information technology, computer science, engineering or related field AND 6 years of experience in computer network administration, database management, systems architecture, systems administration, or related field.

Job Skills:
Same skills as prior level

Job ID or-eeho-fa-us2-oraclecloud-com-cx-45001-345571 · Original posting ↗