Skip to content

AI Engineer Summit 2026

Oct 12–13, 2026Moscone West, San Francisco

Two days for the people building AI systems that have to work on Monday morning. Practitioner talks on evaluation, agents, retrieval, inference infrastructure and the product decisions in between — no keynote fluff, no vendor pitches.

Back to all sessions
AI EngineeringTalk

Evaluation harnesses for production LLMs

When
Monday, October 1210:00 AM - 10:45 AM PDT
Where
Workshop RoomSeats 80
Format
Talk
Track
AI Engineering
Level
Advanced
Language
English
evaluationci

About this session

A working tour of a real evaluation harness: dataset curation, judge prompts you can defend, statistical significance on small samples, and wiring the whole thing into CI so a regression blocks a deploy rather than surprising a customer.

Speakers

TB

Tom Beaumont

Developer Advocate

Ridgeline

HK

Hana Kobayashi

Research Engineer

Institute for Applied ML

Hana works on calibration and abstention — teaching models to recognise the edge of their own competence. She publishes regularly and reviews for three conferences.