Browsing Tag
Site Reliability Engineering
19 posts
Composite Availability: Calculating The Overall Availability Of Cloud Infrastructure
When designing cloud systems, a major consideration is reliability. As infrastructure architects, what do we want to guarantee…
Snooze Your Alert Policies In Cloud Monitoring
Does your development team want to snooze alerts during non-business hours? Or proactively prevent the creation of expected…
Managing The Looker Ecosystem At Scale With SRE And DevOps Practices
Many organizations struggle to create data-driven cultures where each employee is empowered to make decisions based on data.…
Achieving Autonomic Security Operations: Why Metrics Matter (But Not How You Think)
What’s the most difficult question a security operations team can face? For some, is it, “Who is trying…
Are Your SLOs Realistic? How To Analyze Your Risks Like An SRE
Setting up Service Level Objectives (SLOs) is one of the foundational tasks of Site Reliability Engineering (SRE) practices,…
The SRE Book Turns 6!
It’s hard to believe that it’s already been six years since we published Site Reliability Engineering: How Google…
Application Observability Made Easier For Compute Engine
When IT operators and architects begin their journey with Google Cloud, Day 0 observability needs tend to focus…
Quickly Troubleshoot Application Errors With Error Reporting
Are you familiar with the four golden signals of Site Reliability Engineering (SRE): latency, traffic, errors, and saturation?…
Achieving Autonomic Security Operations: Automation As A Force Multiplier
As we discussed in “Achieving Autonomic Security Operations: Reducing toil”, your Security Operations Center (SOC) can learn lessons…
Achieving Autonomic Security Operations: Reducing Toil
Almost two decades of Site Reliability Engineering (SRE) has proved the value of incorporating software engineering practices into…