Atlassian's Journey of Using AI/BI Genie From Pilot to Production and What's Next?
Overview
| Experience | In Person |
|---|---|
| Track | Analytics & BI |
| Industry | Enterprise Technology |
| Technologies | Genie |
| Skill Level | Beginner |
| DOWNLOAD SESSION SLIDES | |
This session is not a generic “AI is the future” talk. It is a practical, experience‑based guide to taking Databricks AI/BI Genie from a promising pilot to a trusted, production‑grade capability. Dashboards answer “what,” but not “why” or “what if.” Here is where Genie comes in.
We’ll walk through our journey at Atlassian from pilot to production—designing Genie spaces on curated tables and metrics, building an evaluation framework for accuracy and coverage, and integrating Genie into our everyday tools and Rovo AI agent. We will touch upon how Databricks and Atlassian partnered to bring the power of Genie to our Rovo Marketplace and built the ability to route a user question to the right Genie space.
In the end attendees will have a concrete “0‑to‑95*” playbook for unleashing their data with Genie to reach a point where roughly 95% of everyday questions can be answered accurately, safely, and self‑service freeing data teams to focus on higher‑value work and empowering faster decisions.
Session Speakers
Manav Trivedi
/Senior Product Manager
Atlassian
Prakash Reddy
/Head of Data Engineering
Atlassian
Full Summary
Inside Atlassian's playbook for production-grade self-serve analytics with Genie
Atlassian set out to let any employee ask data questions in plain English and get reliable answers. The hardest work was not picking a large language model. It was making data trustworthy, scoping problems tightly, and building organizational trust.
FAQ
Genie achieved about 60 to 70 percent accuracy on governed gold tables with no tuning. After enriching table and column descriptions, clarifying metric glossaries, and adding few-shot examples, accuracy rose to roughly 80 to 90 percent depending on question complexity.
Tight scope reduces ambiguity and disambiguation errors, keeps prompts within context limits, and makes evaluation practical. Spaces with fewer than ten tables and a single use case consistently outperformed broader ones.
Yes. Visibility into the generated SQL increased trust, even among non-technical users. Transparency signaled that a real system executed a real query rather than a black box making a guess.
A hub-and-spoke model. The hub owned standards, tooling, governance, benchmarks, and the playbook. Domain teams owned their Genie spaces and outcomes. Champions and regular show-and-tells accelerated reuse and quality.
Metadata quality. Rich, precise, governed table and column descriptions, clear metric glossaries, and business context moved accuracy far more than model selection or prompt tweaks. A benchmark evaluation loop made improvements measurable and durable.