InfrastructureTalk
Serving four hundred models on one GPU fleet
- When
- Monday, October 1211:15 AM - 12:00 PM PDT
- Where
- Main StageSeats 300
- Format
- Talk
- Track
- Infrastructure
- Level
- Advanced
- Language
- English
gpuoperations
About this session
Multi-tenant inference from the operator's seat: scheduling, memory packing, cold-start mitigation, noisy-neighbour isolation, and the observability you need before you can safely oversubscribe anything.
Speaker
LF