LTD Global
Site Reliability Engineer
Berkeley, California · Posted 23 days ago
Opens ltdglobal.applytojob.com
Get a version of your resume written for this job.
- Salary
- Not listed
- Job type
- Contract
- Work mode
- Not specified
- Source
- Jazzhr (employer's hiring system)
Skills mentioned
Python, Java, C++, Perl, Kubernetes, Linux, Prometheus, SRE
About the role
📍 Hybrid — Berkeley, CA
📅 1 Year Contract Assignment with possibility of extension based on performance and organizational needs.
đź’° $80/hr
Ever wondered what powers breakthrough research in energy, physics, materials science, and chemistry? You're looking at it. This national HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption.
If you love solving real problems on live infrastructure, thrive on ownership, and want your work to directly enable world-class science, this is your seat.
What You'll Own
📅 1 Year Contract Assignment with possibility of extension based on performance and organizational needs.
đź’° $80/hr
Ever wondered what powers breakthrough research in energy, physics, materials science, and chemistry? You're looking at it. This national HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption.
If you love solving real problems on live infrastructure, thrive on ownership, and want your work to directly enable world-class science, this is your seat.
What You'll Own
- Monitor and triage alerts across compute, storage, network, and facility systems in real time
- Build automation that prevents issues before they become outages
- Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action)
- Walk the data center floor to keep power, cooling, and environmental systems humming
- Coordinate maintenance activities across teams and keep incidents accurately tracked
- Dig into complex, ambiguous problems and drive them to resolution
- Comfort working Owl shift (12am–8am), 5 days/week, hybrid onsite in Berkeley, CA
- Solid Linux/command-line (SSH) chops
- Programming/scripting experience:Â Python, C, C++, Perl, or Java
- A self-starter mindset, eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
- Network security fundamentals (ACLs, firewalls)
- Strong cross-team communication and collaboration skills
- Experience building or deploying Agentic AI / autonomous automation for technical workflows
- ServiceNow implementation experience
- ITSM best-practice know-how
Job ID jz-ltdglobal-20260901234925_jdqllw4h5wk3dzc0 · Original posting ↗
Similar jobs
Site Reliability Engineer, Kubernetes Platform (Top Secret Clearance)NewSpaceXHawthorne, California
Senior Cloud Site Reliability Engineer ManagerNewScotiabankBogota, District of Columbia
Staff Site Reliability EngineerNewAnduril IndustriesCosta Mesa, California
Senior Site Reliability Engineer, SecOpsNewCambridge Mobile TelematicsCambridge, Massachusetts
Senior Site Reliability Engineer - VolcanoNewKongRemote, United States · Remote · US$150,000–170,000 / year