Staff Site Reliability Engineer (SRE) - Vancouver, BC
Cribl
This job is no longer accepting applications
See open jobs at Cribl.See open jobs similar to "Staff Site Reliability Engineer (SRE) - Vancouver, BC" Redpoint Ventures.Cribl makes open observability a reality for today’s tech professionals. Our category-defining product suite gives companies the power to control their data and the flexibility to make choices, not compromises. With more than $400 million in funding by top investors including IVP, CRV, Redpoint Ventures, Sequoia, Greylock, and Tiger Global, we continue to grow our revenue and customer base by triple digits, with more than a quarter of Fortune 100 companies now Cribl customers.
As a remote first company, Cribl was recently ranked as the top technology/software company on the Forbes Best Startup Employers list (#7 overall), included in CNBC’s Top Startups for the Enterprise, and has been recognized as a top company for women, diversity, and culture by Comparably. So what's it like to work here? Our culture is rooted in our five core values, which includes Irreverent, but Serious. We like to have fun. We like to make each other laugh. And we love Goats!
We are only considering candidates who reside in and are eligible to work in British Columbia at this time.
Why you’ll love this role:
Cribl Inc is seeking a Staff Site Reliability Engineer to join our mission to unlock the value of all observability data. Cribl provides users a new level of observability, intelligence and control over their real-time data. You will join a team of technical engineers who are committed to shipping only high-quality software and enjoying all the goat gifs the internet has to offer. This role is remote and you will be part of the engineering organization where you will contribute in our efforts to envision, create, deploy, test, and ship Cribl products.
Not often do you get to be part of something that is fundamentally changing a technology. But here at Cribl we are building the next generation of software that puts our customers in full control of their observability data. If this is something that interests you, and you want to be truly at the center of the wheel helping make this work better every day. Then this opportunity might be something you have been waiting for to be a part of making a real impact.
We are looking for Cloud Site Reliability Engineers and Developers at all levels at Cribl, who enjoy being in the thick of it. Fixing things at the operational side should always be the last resort, so our SRE engineers are involved from conception to design to development and all the way through production and beyond. You provide your creative input into all things Cloud, Scaling, Reliability, High Availability and much more.
If reliability is your passion, and you have always had strong opinions on how to make things better and have the desire to build consensus around ideas. Then let's talk!
As An Active Member Of Our Team, You Will...
- Engage with teams and improve service delivery and reliability across their entire lifecycle
- Measure and monitor all production systems with an eye towards availability, latency and overall system health
- Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence
- Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability
- Help Identify and drive down toil with creative innovation and automation
- On-call responsibilities
If You Got It, We Want It
- Extensive experience with enterprise scale continuous delivery environments
- 5+ years of experience with a DevOps or SRE job title
- Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment
- Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible
- Experience with sustainable incident response in a blameless environment
- Knowledge of cloud platforms (prefer AWS) and container + orchestration technologies
- Experience with APM and Observability and related tools such as, New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc.
- Background in Linux Systems Engineering
- Experience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc.
- Comfortable with a high level of autonomy and working with a distributed team
Preferred Qualifications
- Knowledge of Cloud and application security
- Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
- A love for high quality and a knack for testing
- Opinions about dashboards, metrics, and SLO’s
#LI-MV1
#Remote
Bring Your Whole Self
Diversity drives innovation, enables better decisions to support our customers, and inspires change for the better. We’re building a culture where differences are valued and welcomed, and we work together to bring out the best in each other. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other applicable legally protected characteristics in the location in which the candidate is applying.
Interested in joining the Cribl herd? Learn more about the smartest, funniest, most passionate goats you’ll ever meet at cribl.io/about-us.
This job is no longer accepting applications
See open jobs at Cribl.See open jobs similar to "Staff Site Reliability Engineer (SRE) - Vancouver, BC" Redpoint Ventures.