Session

Building Higher Quality Genie Agents (fka Genie spaces): How to Ensure Accuracy and Trust

Overview

ExperienceIn Person
TrackAnalytics & BI
IndustryEnterprise Technology
TechnologiesGenie
Skill LevelIntermediate

High-quality Genie Agents (formerly known as Genie Spaces) are built on a foundation of clear metadata, curated semantics, and continuous iteration. New capabilities make it faster than ever to bootstrap and develop Genie Agents.

 

In this session, we'll show how to design, refine, and scale Genie Agents that consistently deliver accurate, trusted results. Topics include:

  • Using Unity Catalog Business Semantics (Metric Views) improve accuracy and consistency from the start
  • Leveraging Genie Code and knowledge mining to automate space setup
  • How to define metadata and context in the Genie knowledge store
  • How to use evaluation and Benchmarks to measure accuracy and guide improvements
  • Best practices for monitoring and optimizing Genie Agents at scale

Session Speakers

Speaker placeholderIMAGE COMING SOON

Hanlin Sun

/Product Manager
Databricks

Shah Amini

/Engineering Leader
Databricks

Full Summary

Building production-ready Genie Agents: a practical guide

Databricks' session at Data and AI Summit outlined a practical path for moving beyond text-to-SQL prototypes to production-grade Genie Agents. The core shift reframed agents as focused, evolving "mini-employees" that deliver reliable answers, learn from usage, and work across structured data, documents, and external tools.

FAQ


No. Spaces are not being deprecated. Capabilities such as document analysis and MCP-based tool connections are being added to the experience you already use.

Use SQL expressions and example SQL for calculations, filters, and metric logic so rules apply dynamically per question. Reserve text instructions for behavior, and write them as explicit rules, for example always require a time range for sales questions.

Prefer digital-native PDFs over scanned files so text is machine-readable. Split files larger than 50 MB, reduce overlapping versions of the same content, and give each volume a clear purpose so the agent knows when to consult it.

Benchmarks are proactive tests that validate quality using curated questions, expected answers or SQL, and optional LLM-as-judge criteria, and can run on a schedule via API. Monitoring is reactive, surfacing real conversations and feedback, with Genie Code summarizing issues and recommending updates.

Yes. The demo showed a single agent querying an opportunities table and analyzing win-loss PDF reports in one flow. Keeping both sources under one agent often yields more granular and accurate answers than stitching separate pipelines.