Staff Engineer, Site Reliability
Listed on 2026-01-06
-
IT/Tech
Systems Engineer, Cloud Computing
Linked In is the world’s largest professional network, built to create economic opportunity for every member of the global workforce. Our products help people make powerful connections, discover exciting opportunities, build necessary skills, and gain valuable insights every day. We’re also committed to providing transformational opportunities for our own employees by investing in their growth. We aspire to create a culture that’s built on trust, care, inclusion, and fun – where everyone can succeed.
Join us to transform the way the world works.
Job DescriptionSite Health Platform sits at the core of Linked In’s Reliability Infrastructure organization, with a primary focus on the end-to-end incident management ecosystem. Our mission is for every member and customer to experience Linked In as "always on", every engineer to benefit from a more insightful and proactive site-wide reliability ecosystem, and every business and product owner to be well-informed about service disruptions as they occur.
We own the full incident lifecycle across thousands of services and multiple regions, from incident response and mitigation, through problem management and post-incident learning. The platforms we build are the backbone of how Linked In detects issues, coordinates incident response, captures context, and turns outages and near misses into structured, actionable insights.
By transforming incidents into data and learnings, we enable teams to systematically improve reliability over time. Our work informs engineering priorities, infrastructure investments, capacity planning, and executive decision-making, ensuring the network is dependable when it matters most.
You will be exposed to many different technologies, architectures, and systems hosted in state-of-the-art data centers across the globe.
At Linked In, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a Linked In office on select days, as determined by the business needs of the team.
QualificationsResponsibilities:
- Designing and evolving the core incident management platforms that power Linked In’s full incident lifecycle, from detection and response to problem management and prevention, across thousands of services and teams.
- Serving in a critical on-call rotation, providing expert incident triage and coordination during high-severity outages. Partnering closely with service owners and product teams to diagnose issues quickly, mitigate member impact, and drive timely resolution under pressure.
- Transforming raw, unstructured incident data into clear, actionable intelligence using AI and LLM-based systems, including automated summarization, classification, root cause signals, and mitigation recommendations.
- Building analytics and insights that surface systemic reliability risks, recurring failure patterns, and cross-service dependencies, enabling org-level prioritization rather than isolated, service-by-service fixes.
- Building platforms and tools that enable realistic, fleet-wide stress testing of data center and regional capacity, validating incident readiness across dependencies, traffic patterns, and growth scenarios before they impact a significant production outage.
- Driving consistency, clarity, and quality in how incidents are declared, managed, reviewed, and learned from, raising the reliability bar across a large, fast-moving engineering organization.
- Influencing service architecture, SLOs, and reliability standards through platforms, data, and technical leadership, ensuring improvements are durable, measurable, and adopted at scale.
Basic Qualifications:
- Bachelor’s degree in Computer Science, Engineering, or related technical field, or equivalent practical experience. Many postings also prefer or require an advanced degree (MS/PhD) for Staff-level roles.
- 6+ years of professional experience in software development, distributed systems, or reliability engineering. Some Principal/Staff roles list around 10+ years of experience.
- Several years of experience leading technical…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).