Session

Keeping AI safe in healthcare - evaluation strategies for production ready agentic systems

Register or Login

Overview

ExperienceIn Person
TrackArtificial Intelligence & Agents
IndustryHealthcare & Life Sciences
TechnologiesUnity Catalog, Databricks Agents, Lakebase
Skill LevelIntermediate
Healthcare organizations struggle with siloed data from claims, hospital records, physician notes, and lab results and now with the latest agentic frameworks on top of Databricks, we are able to unify these data sources and provide access to insights that can vastly improve patient care. However, deploying AI agents in healthcare requires rigorous evaluation that goes beyond traditional accuracy metrics to ensure clinical safety and compliance. In this breakout session, attendees will learn how to implement continuous evaluation workflows that track agent performance across critical dimensions specific to healthcare. We will demonstrate how to leverage MLflow 3 to build a comprehensive evaluation pipeline for a multi agent AI system that includes-SQL-based validation rules for data integrity-Structured UAT protocols with clinical stakeholders-LLM-as-Judge evaluation for scalable output assessment-Human-in-the-loop feedback alignment -Automated production monitoring mechanisms

Session Speakers

Elena Boiarskaia

/AI/ML Practice Lead
Koantek