by Hanlin Sun and Richard Tomlinson
Archived. This article has not been updated since the publish date above. The dynamic nature of information means that previously accurate content can become outdated or even obsolete over time. Readers are advised to exercise due diligence and cross-check any information found in this blog post before making decisions or adopting any practices based on said information.
AI/BI Genie is a conversational experience for business teams to self-serve insights from their data through natural language. Genie leverages generative AI tailored to an organization’s data, usage patterns, and business concepts and continuously learns from user feedback. This allows non-technical users to ask questions as they would to an experienced coworker, receiving relevant and accurate responses directly from their enterprise data.
With the growing adoption of Genie spaces, it is essential that users have confidence in the accuracy of the insights provided. This assurance is crucial in enabling them to make the most informed decisions based on the insights Genie delivers.
Data practitioners responsible for authoring and maintaining Genie spaces for their business teams commonly cite two critical requirements:
To address these requirements, we are excited to introduce two new features in AI/BI Genie to help build confidence in the accuracy of answers returned:
Benchmarks allow Genie authors to systematically evaluate the accuracy of their Genie spaces. A well-crafted set of benchmark questions should include the most frequently asked user questions, along with 2-3 variations in phrasing. Authors can then run these benchmarks over time to determine whether edits to the space are effectively improving overall accuracy.
To better assess your Genie space’s accuracy with Benchmarks, follow these steps:



Genie is a powerful tool for exploratory data analysis, allowing non-technical users to ask follow-up questions and get new insights from their data without involving expert practitioners. However, just like analysis in other tools like Excel, you may want a second opinion before presenting your findings as factual.
The Request Review feature enables end-users to complete this review cycle directly in Genie—there is no need for screenshots and back-and-forths in Slack or Teams.


With the introduction of Benchmarks and Request Review, AI/BI Genie significantly enhances user confidence in the accuracy and reliability of the answers they receive. Benchmarks allow for systematic tracking of accuracy improvements over time, ensuring that instruction edits are effective. Request Review provides a seamless way for users to verify critical responses, fostering trust in the insights that Genie generates. Together, these new features empower business teams to confidently leverage Genie to make the critical decisions required in their daily work.
We encourage you all to start creating Genie spaces if you haven't already. Make sure to read through our AI/BI Genie documentation. To see AI/BI Dashboards and Genie in action, check out our demo and take the product tour.
The Databricks team is always looking to improve the AI/BI Genie experience, and would love to hear your feedback!
Subscribe to our blog and get the latest posts delivered to your inbox.