Java SRE Engineer

Eitacies Inc

Java SRE Engineer

Santa Clara, CA
Full Time
Paid
  • Responsibilities

    Java SRE Engineer

    Onsite San Francisco Bay Area

    Infrastructure Engineer (2 Positions)

    We are looking for an experienced Java SRE / Platform Engineer to support large-scale cloud migrations and production systems on AWS and Kubernetes platforms. This role is focused on infrastructure, reliability, and automation, with Java exposure as a supporting skill.

    Required Skill : AWS, AWS EKS, Kubernetes, DevOps / SRE, Java

    Key Responsibilities:

    Lead large-scale migrations of business-critical applications to AWS and Kubernetes (EKS)

    Design and operate production-grade AWS EKS platforms

    Implement GitOps-based deployment strategies using ArgoCD and Spinnaker

    Build and manage CI/CD pipelines and automated release strategies (blue/green, canary)

    Develop Python-based automation for infrastructure and operations

    Create and maintain Helm charts and deployment standards

    Troubleshoot and optimize Linux-based systems in production environments

    Support production systems including on-call, incident response, and RCA

    Collaborate with SRE and Security teams to ensure system reliability and scalability

    Drive architectural decisions and contribute to long-term platform strategy

    Mentor team members and improve engineering practices

    Qualifications:

    10+ years of experience in Cloud / DevOps / SRE / Platform Engineering

    Strong hands-on experience with:

    AWS (EKS, EC2, VPC, IAM, ALB/NLB, CloudWatch, S3, RDS)

    Kubernetes, Linux systems, Python, ArgoCD (GitOps)

    Spinnaker, Helm

    Experience with Infrastructure as Code (Terraform or CloudFormation)

    Proven experience supporting production environments

    Experience leading or contributing to AWS migration projects

    Strong understanding of distributed systems and networking

    Preferred Qualifications:

    Experience with Akamai CDN and caching strategies

    Experience with Redis and Kafka

    Familiarity with observability tools (Prometheus, Grafana, Datadog, Splunk)

    Experience with service mesh (Istio, Linkerd)

    Knowledge of SRE practices (SLIs, SLOs, error budgets)

    Strong communication and documentation skills