AWS for Software Companies Podcast
Ep216: Powering AI-enabled Operational Insights with Amazon Bedrock
28 July 2026 23:01 AWS - Amazon Web Services
Listen to episode
About this episode
From alert to root cause in one minute - how PagerDuty built autonomous incident response on Amazon Bedrock, and the future of triage and trust.
Topics Include:
- PagerDuty's agents must perform during 2am outages — stakes are high
- Software shipping accelerated dramatically; production environments largely did not
- A 9:30pm slowdown traced to a race condition solved two years earlier
- The fix was documented — but the context wasn't at hand
- PagerDuty Advance ships four agents: SRE, Scribe, Shift, Insights
- Why four, not one? Focus and predictability in non-deterministic systems
- Saurabh Shanbhag: Bedrock is far more than a model service
- Zero data retention, PrivateLink, TLS — why enterprises pick Bedrock
- Frontier models everywhere burns tokens; classify, route, distill, fine-tune
- SRE agent triages alerts before you even join the call
- One minute to root cause — context beat raw intelligence
- Human surfaces versus machine surfaces: MCP and CLI move fastest
- "The model eats the harness" — every upgrade invalidates foundational components
- Feeding agents everything failed; compartmentalised investigation threads work better
- New York Life's three stages of trust, and the seatbelt override that wasn't
Participants:
- Tom Hogarty - Senior Director Product Management, PagerDuty
- Saurabh Shanbhag – Sr Partner Solution Architect, Amazon Web Services
See how Amazon Web Services gives you the freedom to migrate, innovate, and scale your software company at https://aws.amazon.com/isv/
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity