📅 Date: 4 December 2024
📍 Location: Howest, BST Campus
🎓 Type: Passive Event
I attended a Tech & Meet session hosted by two Google Cloud Site Reliability Engineers (SREs), Stefaan V. and Tijl Van den Broeck. The session offered deep insights into how Google designs for reliability at scale, not just through code, but through culture.
💡 What Is SRE?
As Stefaan explained:
“SRE is what happens when you ask a software engineer to design an operations team.” Benjamin Treynor Sloss
SRE isn’t just a role. It’s a mindset. At Google, reliability becomes a core engineering goal only after the product is built. That includes planning for failure, automating responses, and constantly measuring how systems behave in the real world.
🧠 Key Concepts We Covered:
- SLI (Service Level Indicator): How you measure performance
- SLO (Service Level Objective): What level of service you aim to provide
- SLA (Service Level Agreement): What you promise users
What stood out to me was the idea that 100% uptime isn’t always the goal. Chasing perfection can cost more than it’s worth. Instead, teams aim for “just reliable enough”, balancing risk and cost.
🧪 Hands-On: “The Movie Guru” Lab
In the final part of the workshop, we acted as SREs for a fictional app. We had to allocate budget between reliability improvements and development speed, simulating the kind of decisions real engineers face every day.
It was a practical reminder that engineering is about trade-offs, not just building features.
🔍 My Takeaways:
- Reliability is a feature, and one that users feel before they see.
- Automation and monitoring aren’t optional at scale.
- SLIs and SLOs are essential tools, not just theoretical ideas.
- SRE blends software engineering with real-time problem solving, and that’s a space I’m excited to explore more in the future.
