Artwork

Content provided by Adam Hawkins. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by Adam Hawkins or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://ro.player.fm/legal.
Player FM - Aplicație Podcast
Treceți offline cu aplicația Player FM !

Incidents & Operations with Dan Slimmon

1:01:48
 
Distribuie
 

Manage episode 433608422 series 2814917
Content provided by Adam Hawkins. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by Adam Hawkins or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://ro.player.fm/legal.

In this episode, Adam welcomes Dan Slimmon, an experienced Site Reliability Engineer (SRE) to discuss aspects of incident response and troubleshooting in software engineering. Dan explains his methodology for clinical troubleshooting, the importance of maintaining a common mental model, and techniques for leading effective incident response efforts. They also delve into the value of continuous ops reviews and ongoing mental model updates to prevent issues, emphasizing the need for structured processes and effective communication.

Want more?

Chapters

  • (00:00) - Incidents & Operations
  • (01:14) - Guest Welcome
  • (01:40) - Dan's Career Journey
  • (02:33) - Evolution of Tech Stacks
  • (04:59) - Clinical Troubleshooting Explained
  • (11:53) - Incident Response Fundamentals
  • (17:41) - Effective Communication in Incidents
  • (26:09) - Training for Incident Response
  • (33:22) - The Essence of Incident Response
  • (33:53) - Balancing Short-Term and Long-Term Fixes
  • (35:01) - The Firefighting Analogy in Software Incidents
  • (37:11) - Postmortems: Learning from Incidents
  • (42:14) - Building a Shared Mental Model
  • (42:41) - Looking for Trouble: Proactive System Monitoring
  • (47:59) - Ops Reviews: Continuous Improvement
  • (54:37) - The Importance of Closing the Feedback Loop
  • (59:40) - Final Thoughts and Resources
★ Support this podcast on Patreon ★
  continue reading

121 episoade

Artwork
iconDistribuie
 
Manage episode 433608422 series 2814917
Content provided by Adam Hawkins. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by Adam Hawkins or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://ro.player.fm/legal.

In this episode, Adam welcomes Dan Slimmon, an experienced Site Reliability Engineer (SRE) to discuss aspects of incident response and troubleshooting in software engineering. Dan explains his methodology for clinical troubleshooting, the importance of maintaining a common mental model, and techniques for leading effective incident response efforts. They also delve into the value of continuous ops reviews and ongoing mental model updates to prevent issues, emphasizing the need for structured processes and effective communication.

Want more?

Chapters

  • (00:00) - Incidents & Operations
  • (01:14) - Guest Welcome
  • (01:40) - Dan's Career Journey
  • (02:33) - Evolution of Tech Stacks
  • (04:59) - Clinical Troubleshooting Explained
  • (11:53) - Incident Response Fundamentals
  • (17:41) - Effective Communication in Incidents
  • (26:09) - Training for Incident Response
  • (33:22) - The Essence of Incident Response
  • (33:53) - Balancing Short-Term and Long-Term Fixes
  • (35:01) - The Firefighting Analogy in Software Incidents
  • (37:11) - Postmortems: Learning from Incidents
  • (42:14) - Building a Shared Mental Model
  • (42:41) - Looking for Trouble: Proactive System Monitoring
  • (47:59) - Ops Reviews: Continuous Improvement
  • (54:37) - The Importance of Closing the Feedback Loop
  • (59:40) - Final Thoughts and Resources
★ Support this podcast on Patreon ★
  continue reading

121 episoade

Toate episoadele

×
 
Loading …

Bun venit la Player FM!

Player FM scanează web-ul pentru podcast-uri de înaltă calitate pentru a vă putea bucura acum. Este cea mai bună aplicație pentru podcast și funcționează pe Android, iPhone și pe web. Înscrieți-vă pentru a sincroniza abonamentele pe toate dispozitivele.

 

Ghid rapid de referință