Site Reliability Engineer | Hybrid

LTD GLOBAL, LLC

Site Reliability Engineer | Hybrid

Berkeley, CA
Full Time
Paid
  • Responsibilities

    Job Description

    Job Description

    ** Hybrid — Berkeley, CA**

    ** Assignment: 10/26/2026 – 10/27/2027**

    ** $80/hr**

    Role Summary


    As a Site Reliability Engineer on the Operations Technology team, you'll be part of a round-the-clock crew keeping a national-scale HPC facility accessible, reliable, and secure. Working from advanced monitoring and data collection systems, you'll proactively catch issues before they escalate, triage and resolve alerts across compute, storage, and network systems, and build the automation that makes the whole environment more resilient over time. You'll also collaborate closely with cross-functional teams to coordinate maintenance, improve tooling, and ensure the infrastructure scales smoothly as demand grows, keeping the computational power behind critical scientific research running without interruption.


    What You'll Own

    • Monitor and triage alerts across computer, storage, network, and facility systems in real time
    • Build automation that prevents issues before they become outages
    • Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action)
    • Walk the data center floor to keep power, cooling, and environmental systems humming
    • Coordinate maintenance activities across teams and keep incidents accurately tracked
    • Dig into complex, ambiguous problems and drive them to resolution

    What You Bring

    • Comfort working Owl shift (12am–8am) , 5 days/week, hybrid onsite in Berkeley, CA
    • Solid Linux/command-line (SSH) chops
    • Programming/scripting experience — Python, C, C++, Perl, or Java
    • A self-starter mindset — eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
    • Network security fundamentals (ACLs, firewalls)
    • Strong cross-team communication and collaboration skills

    Nice to Have

    • Experience building or deploying Agentic AI / autonomous automation for technical workflows
    • ServiceNow implementation experience
    • ITSM best-practice know-how

    \nCompany Description

    Great Organization!

    Company Description

    Great Organization!