AI EngineeringTalk
Evaluation harnesses for production LLMs
- When
- Monday, October 1210:00 AM - 10:45 AM PDT
- Where
- Workshop RoomSeats 80
- Format
- Talk
- Track
- AI Engineering
- Level
- Advanced
- Language
- English
evaluationci
About this session
A working tour of a real evaluation harness: dataset curation, judge prompts you can defend, statistical significance on small samples, and wiring the whole thing into CI so a regression blocks a deploy rather than surprising a customer.
Speakers
TB
HK
Hana works on calibration and abstention — teaching models to recognise the edge of their own competence. She publishes regularly and reviews for three conferences.