Hana Kobayashi
Research Engineer
Institute for Applied ML
Hana works on calibration and abstention — teaching models to recognise the edge of their own competence. She publishes regularly and reviews for three conferences.
Their session1
Monday, October 12, 2026
Evaluation harnesses for production LLMs
A working tour of a real evaluation harness: dataset curation, judge prompts you can defend, statistical significance on small samples, and wiring the whole thing into CI so a regression blocks a deploy rather than surprising a customer.
Speakers
- TBTom Beaumont
Developer Advocate · Ridgeline
- HKHana Kobayashi
Research Engineer · Institute for Applied ML