Webinar Thursday August 13 | 6am & 11am PT

Four Architectural Decisions That Make or Break Production Inference at Scale

A practical look at the four decisions that make production inference reliable and fast, from how you talk to the infrastructure to catching failures before your users do. 

30 min Demo + Q&A
Live

Speakers

Greg Wester

Greg Wester

Director of Product

Greg leads the Serverless product at Runpod. He shipped MIG, FlashBoot, and batch inference, which means he's responsible for most of what's on the agenda today. When he's not thinking about cold start latency, he coaches performance driving.




REST API, SDK, or GraphQL, and how scoped access control changes the calculus once more than one person on your team is touching production: how to decide what to build against, and where each choice leaves you stuck later
Real-time versus batch inference: how to size which one a production workload actually needs, and the cost and latency tradeoffs that only show up after you've shipped

When a model needs more than one GPU: where distributing inference is worth the added complexity, and where it's the wrong tool for the job

Which error-handling and monitoring patterns hold up at production volume, and which ones are just busywork

Bring your toughest questions about scaling inference.