Staff Site Reliability Engineer (SRE) - Remote

Company:  Cribl, Inc
Location: Jackson
Closing Date: 19/10/2024
Hours: Full Time
Type: Permanent
Job Requirements / Description

This Position is for a Staff Site Reliability Engineer (SRE) with a remote work location in Jackson, MS.Seeking a Staff Site Reliability Engineer to join the mission to unlock the value of all observability data. Provides users a new level of observability, intelligence and control over their real-time data. You will join a team of technical engineers who are committed to shipping only high-quality software and enjoying all the goat gifs the internet has to offer. This role is remote and you will be part of the engineering organization where you will contribute in our efforts to envision, create, deploy, test, and ship Company products.Looking for Cloud Site Reliability Engineers and Developers at all levels at the Company, who enjoy being in the thick of it. Fixing things at the operational side should always be the last resort, so SRE engineers are involved from conception to design to development and all the way through production and beyond. You provide your creative input into all things Cloud, Scaling, Reliability, High Availability and much more.As An Active Member Of Our Team, You Will...Engage with teams and improve service delivery and reliability across their entire lifecycleMeasure and monitor all production systems with an eye towards availability, latency and overall system healthSeek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellenceEngage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observabilityHelp Identify and drive down toil with creative innovation and automationOn-call responsibilitiesExtensive experience with enterprise scale continuous delivery environments8+ years of experience with a DevOps or SRE job titleDevelopment with JavaScript/Node.js/TypeScript in a Linux/Mac environmentExperience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, AnsibleExperience with sustainable incident response in a blameless environmentKnowledge of cloud platforms (prefer AWS) and container + orchestration technologiesExperience with APM and Observability and related tools such as, New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc.Background in Linux Systems EngineeringExperience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc.Comfortable with a high level of autonomy and working with a distributed teamPreferred QualificationsKnowledge of Cloud and application securityStrong knowledge of cloud design patterns for scale, data management, resiliency, etc.A love for high quality and a knack for testingOpinions about dashboards, metrics, and SLO'sDiversity drives innovation, enables better decisions to support customers, and inspires change for the better. Building a culture where differences are valued and welcomed, and work together to bring out the best in each other. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other applicable legally protected characteristics in the location in which the candidate is applying.

Apply Now
Share this job
Cribl, Inc
  • Similar Jobs

  • Site Reliability Engineer - Remote US

    Jackson
    View Job
  • Site Reliability Engineer - GCP (Remote)

    Jackson
    View Job
  • Senior Site Reliability Engineer

    Jackson
    View Job
  • Senior Site Reliability Engineer - Automation / Containers

    Jackson
    View Job
  • Site Reliability Engineer - Federal Operations Automation

    Jackson
    View Job
An error has occurred. This application may no longer respond until reloaded. Reload 🗙