Sherwin-Williams
Lead Site Reliability Engineer
Cleveland, Ohio · Posted today
Opens ejhp.fa.us6.oraclecloud.com
Get a version of your resume written for this job.
- Salary
- Not listed
- Job type
- Not specified
- Work mode
- Not specified
- Source
- Oracle (employer's hiring system)
Skills mentioned
Azure, DevOps, Kubernetes, Distributed Systems, Incident Response, AWS, CI/CD
About the role
The Lead Site Reliability Engineer role is responsible for optimizing the organization's IT products, services, systems, and digital products. The incumbent works to develop customized solutions to automate the administration, monitoring, and operation of business-critical applications and services, supervises tests for resiliency, redundancy, and failover to ensure uptime, and troubleshoots potential issues to ensure that IT products, services, systems, and digital products are running efficiently and effectively. In addition, the role is responsible for leading the design and implementation of scalable and reliable application and service solutions that can run across multiple environments and technologies. The incumbent fosters a culture of collaboration between cross-functional departments, including development teams, infrastructure teams, and support organizations, to enhance and improve system operability and provide training and knowledge transfer to team members. The role is also responsible for implementing leading practices and emerging technologies that will drive additional efficiencies across the IT organization. The role is also responsible for overseeing project planning, cost analysis, and vendor comparisons when assessing potential solutions and implementing technology and operational improvements.
WHAT THE ROLE WILL DO:
Optimize IT products, services, systems, and digital products by proactively analyzing application, service, and operational health metrics to identify potential inefficiencies and thereby ensuring improvements in performance, availability, and reliability.
Develop customized solutions to automate the deployment, administration, monitoring, and operation of applications and services and train other team members on the use of automation tools.
Develop a formal process for continuously reviewing and monitoring system SLIs, SLOs, SLAs, and OKRs and build optimization plans to address areas of improvement.
Foster a culture of collaboration between cross-functional teams to ensure improvement in IT products, services, systems, and digital products.
Optimize the use of applications, services, integrations, observability tools, infrastructure components, and load-balancing technologies by tracking performance, identifying potential issues, and ensuring optimal operation.
Lead efforts to continuously improve application and service reliability, performance, resiliency, observability, and operational readiness and take steps to mitigate potential issues.
Evaluate application and service requirements, lead cross-functional implementation teams, and conduct post-implementation reviews to share lessons learned from the project.
Create detailed implementation plans for the integration of new technologies, products, and services into the existing environment that will improve application and service resilience, performance, and reduce costs.
Provide leadership and knowledge-sharing to mentor engineers on reliability engineering, observability, automation, incident management, and operational excellence practices and assess their effectiveness.
This position is not hybrid/remote and will be located at our Cleveland Headquarters office.
This position is not eligible for sponsorship for work authorization now or in the future, including conversion to H1-B visa.
Job duties include contact with other employees and access confidential and proprietary information and/or other items of value, and such access may be supervised or unsupervised. The Company therefore has determined that a review of criminal history is necessary to protect the business and its operations and reputation and is necessary to protect the safety of the Company’s staff, employees, and business relationships.
Must be eighteen years or older
Required Qualifications:
Bachelor’s degree in Computer Science or Information Systems, or in lieu of a degree, at least 9 years of experience in the field of site reliability engineering
6+ years of experience in Site Reliability Engineering, Software Engineering, DevOps Engineering, Platform Engineering, Systems Engineering, or a related technical discipline.
Experience supporting and operating production applications, services, or enterprise technology platforms.
Experience with monitoring, logging, observability, and incident response practices.
Experience utilizing automation, scripting, or tooling to improve reliability, operational efficiency, and supportability.
Experience collaborating with cross-functional teams to troubleshoot, resolve, and prevent production issues.
Strong analytical and problem-solving skills
- Must be at least (18) eighteen years of age
- Must be legally authorized to work in the country of employment without now or in the future requiring sponsorship for employment visa status (e.g., OPT, CPT, H1B, EB-1, etc.)
Preferred Qualifications:
Experience supporting customer-facing or business-critical digital products and services at a global level.
Experience establishing and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
Experience leading major incident response, root cause analysis, and post-incident reviews.
Experience with cloud-native architecture (Microsoft Azure, AWS, Kubernetes, etc.)
Experience implementing observability, alerting, and operational excellence practices.
Experience with CI/CD pipelines and Infrastructure as Code (IaC).
Experience driving reliability, resiliency, and performance improvements in distributed systems.
Experience with leadership and executive-level communication.
Microsoft Certified: Azure Solutions Architect Expert
Relevant Site Reliability Engineering, Azure, Cloud, DevOps, or Platform Engineering certifications
Technical Skills:
Monitoring and Logging
Automation and Scripting
Reliability Engineering
Incident Management
Application and Service Operations
Performance and Capacity Management
Microsoft Azure
Kubernetes
Application Performance Monitoring
Observability and Telemetry
Infrastructure as Code
Distributed Systems
API and Integration Technologies
DevOps Practices
Job ID or-ejhp-fa-us6-oraclecloud-com-cx-2-2625805 · Original posting ↗
Visa sponsorship history
- 26 H-1B petitions approved for THE SHERWIN WILLIAMS COMPANY in fiscal year 2023 (USCIS).
From public government data. It shows this employer has sponsored workers before, not that this job offers sponsorship: check the job ad or ask the employer. More visa-friendly jobs
Similar jobs
- Site Reliability Engineer, Provider OperationsNewOpenRouterRemote, United States · Remote
- Senior Site Reliability EngineerNew2KAustin, Texas
Senior Site Reliability Engineer (SRE)NewTradewebUnited States
Site Reliability EngineerNewAnduril IndustriesWaltham, Massachusetts
Fielded Site Reliability EngineerNewAnduril IndustriesWaltham, Massachusetts