Keep business logic in governed data systems
The agent can help interpret intent, but KPI definitions, authorization, and query execution stay in governed platform layers such as dbt, Trino, and Apache Ranger-based controls.
Use metadata before data
The system grounds responses with catalogs, business definitions, schema relationships, query history, and usage patterns before attempting execution.
Retrieve schemas instead of stuffing them
Table selection follows the RASL pattern (arXiv:2507.23104): decompose schemas into indexed semantic units and retrieve per question with relevance calibration. That makes catalog growth an indexing task rather than a modeling or prompt-length problem.
Give tools protocol boundaries
Exposing metadata lookup, execution, and charting as MCP servers keeps tool contracts explicit and versioned, separates platform concerns from prompt concerns, and gives security one review surface per capability.
Validate SQL outside the prompt
Dialect checks, guardrails, and plan inspection run as deterministic pre-execution steps rather than model self-critique, because self-review is not a safety boundary.
Model business questions as structured intent
Questions such as top manager by city, month-over-month trends, and quarterly growth comparisons need explicit metric, geography, time, and entity slots before SQL generation is useful.
Split the agent into subgraphs
Knowledge retrieval, business question answering, query generation, execution, answer reasoning, and visualization each need different prompts, tools, state contracts, and failure handling.
Carry structured state between nodes
Passing structured data between graph nodes made behavior easier to reason about, improved determinism, and avoided spending tokens restating context that the system already knew.
Trace prompts and graph behavior
Langfuse made prompt versions and graph execution observable, which matters when a business-facing data agent needs repeatable behavior instead of one-off demos.
Return references, not unrestricted data
Query results move through persisted references and controlled previews, which keeps downstream visualization useful without expanding the prompt surface unnecessarily.
Split agent reasoning from query execution
Separating the conversational runtime from the governed query engine made cancellation, guardrails, auditing, and authorization easier to reason about.
Treat the chat UI as product surface
Customizing Open WebUI improved the user path for data workflows because business users need richer controls than a plain chat transcript when moving from question to query to chart.