Blog Post

Prep data for AI vs. Fabric Data Agent Instructions: What Goes Where?

,

Making Data AI-Ready, Part 3

(This is the final article in a three-part series on making data AI-ready. Part 1 covered what AI-ready data means, Part 2 looked at the Microsoft tools that support it, and this article focuses on how Prep data for AI and Fabric Data Agent instructions work together.)

You have connected a Power BI semantic model to a Fabric Data Agent, and now you want to explain that People means Customers and Revenue means Net Sales Amount. Several configuration names sound relevant: AI instructions, Data Agent instructions and data-source instructions, and example queries. Should you repeat the explanation everywhere to be safe? I would not take that approach, because these capabilities have different responsibilities and are not all available for semantic models. The useful distinction is between preparing the semantic model for AI and configuring the agent that decides when to use it.

Understand the two different jobs

The Fabric Data Agent architecture separates orchestration from source-specific querying. The agent interprets a question, chooses the relevant source or sources, and invokes the appropriate tool to generate and execute a query. For a Power BI semantic model, that involves DAX; other source types use their own query mechanisms. After the results return, the agent assembles a response for the user. I would describe the division this way: agent instructions address “When should I use this source, and how should the overall agent behave?” while model-specific configuration addresses “Once this model is selected, how should its data be understood and queried?”

The reason this distinction matters is that the Data Agent and the DAX generator have different jobs. Think of the Data Agent as the orchestrator. Its instructions help it interpret the user’s question, decide which source or sources to use, apply cross-source rules, and determine how the final response should be presented. If the agent decides that a Power BI semantic model is the right source, it then hands the question to a separate DAX-generation tool whose job is to understand that semantic model and create the appropriate DAX query.

That DAX-generation tool does not receive the Data Agent instructions. Instead, Microsoft’s semantic-model best practices state that it uses the semantic model’s metadata and its Prep data for AI configuration. This means an instruction such as “Use the Sales semantic model for revenue questions” belongs at the Data Agent level because it helps the agent choose a source. But an instruction such as “When the user says Revenue, use the Net Sales Amount measure” belongs in Prep data for AI because it tells the DAX generator how to interpret and query that particular semantic model.

Another way to think about it is that Data Agent instructions get you to the right source, while Prep data for AI helps generate the right query once you get there. Keeping semantic-model-specific knowledge with the model also makes that context reusable by supported experiences beyond a single Data Agent, including Power BI Copilot.

The easiest way to see the division of responsibility is to compare the capabilities side by side:

CapabilityScopeSupported for a Power BI semantic model source?Used for semantic-model DAX generation?Best use
Prep data for AI – AI Data SchemaSemantic model??Control which tables, columns, and measures AI should focus on
Prep data for AI – AI InstructionsSemantic model??Define terminology, business rules, synonyms, and interpretation
Prep data for AI – Verified AnswersSemantic model??Provide approved question/answer patterns that help guide DAX generation
Data Agent InstructionsEntire Data Agent??Routing, overall behavior, source selection, response formatting, and cross-source terminology
Data Source InstructionsIndividual source?—Source-specific guidance for supported sources such as SQL, KQL, and graph
Schema Object DescriptionsIndividual schema object?—Explain the business meaning of individual SQL tables, columns, and other schema objects; especially useful for ambiguous or technical names
Example QueriesIndividual source?—Show how natural-language questions map to SQL, KQL, or GQL queries

That fourth column is especially important. Data Agent instructions can help the agent decide to use the Sales semantic model, but those instructions are not passed to the DAX-generation tool. Prep data for AI provides the model-specific context used to understand and query the semantic model once that source has been selected.

What should Data Agent instructions say?

Use agent-level instructions for routing, cross-source behavior, and presentation rather than the detailed interpretation of one semantic model. In a hypothetical sales-and-support agent, you might write: “Use the Sales semantic model for revenue, margins, customers, and sales performance. Use the Support lakehouse for support tickets, and consult both sources when a question requires sales results and ticket information.” A separate instruction could request a brief summary followed by a table and clearly identified reporting periods. Common terminology can belong here when it helps the agent understand questions across sources, but that does not replace the model-level definitions needed to generate correct DAX.

What belongs with the semantic model?

Prep data for AI supplies three model-level capabilities: AI data schemas, AI instructions, and Verified Answers. Together with the model’s names, descriptions, relationships, and measures, they provide context for supported AI experiences using that model. Calling Prep data for AI the source-specific configuration layer is a helpful mental model, although it is not the only configuration involved. You can still select semantic-model tables when adding the source to a Data Agent, and the underlying model remains important. I would prepare the model first rather than treating AI instructions as a replacement for unclear measures or incorrect relationships.

An AI data schema narrows attention to relevant tables, columns, and measures. Suppose a model contains Gross Sales, Net Sales Amount, Sales Before Returns, and several internal helper measures; exposing everything without context leaves more room for choosing the wrong metric. Focus the schema on what your agent should answer, while retaining necessary dependencies. Microsoft recommends aligning the agent’s table selection with the model’s AI data schema. Treat this as relevance configuration, not a substitute for permissions: the schema’s behavior varies across Power BI capabilities, and it does not redefine what a user is authorized to access.

AI instructions are where model-specific terminology and analytical rules belong. For our hypothetical model, you could write: “The People table contains purchasing customers; Person_ID identifies a customer. When users ask about revenue, use the Net Sales Amount measure, which excludes returns. GM refers to the Gross Margin Percentage measure, and sales-period questions use Order Date unless the question explicitly concerns shipments.” Notice that these instructions do more than expand abbreviations: they identify the intended objects and interpretation. Use actual model names, confirm the definitions with the business, and keep the guidance specific enough that you can test whether it was applied.

There is also a reuse benefit to keeping this knowledge with the model. Microsoft’s Prep data for AI FAQ explains that these configurations are stored on the semantic model rather than on individual reports, although different Copilot capabilities use different parts of the configuration. That is a reason to maintain shared definitions centrally instead of copying them into every agent. It is not a promise that every AI product automatically consumes every setting. Keep “Use the Sales model” at the orchestration level and “Revenue means this measure” with the model, even when both statements contribute to answering the same question.

Verified Answers are not an example-query text box

Verified Answers associate trigger questions with approved Power BI visuals and their relevant configuration. For example, you could use a reviewed visual showing Net Sales Amount by fiscal quarter and associate it with questions about quarterly sales performance. The value is not simply that someone supplied sample words; the visual embodies an agreed choice of measures, dimensions, and filters. For a question such as “What were sales last quarter?”, verify that the underlying calculation and time interpretation actually match that request. A visual permanently filtered to one historical quarter is not automatically a correct answer to a moving time-period question.

When a Fabric Data Agent uses a Verified Answer, it does not return the Power BI visual itself. Instead, the visual’s measures, columns, and filters help guide DAX generation. That is different from entering a natural-language question and arbitrary DAX into a Data Agent example-query box. Review the current Verified Answer limitations: Microsoft documents limitations for row-level and object-level security and says Power BI Copilot does not return Verified Answers when Fabric IQ is enabled. I would validate behavior in the exact experience you are deploying instead of assuming that a successful Power BI demonstration proves identical behavior in every consuming tool.

What changes for other source types?

Microsoft’s source-configuration support matrix is explicit: Power BI semantic models do not support Data Agent data-source instructions, data-source descriptions, or example queries. Model-specific guidance belongs in Prep data for AI instead. SQL and Eventhouse/KQL sources support source instructions and question-query examples, as do graph sources in preview. Ontology sources support neither source instructions nor example queries. For a SQL source, a description might explain what business area the source covers, while source instructions explain joins, status codes, and calculation rules. Those are different jobs: selecting the appropriate source is not the same as generating the appropriate query within it.

SQL sources can also use schema object descriptions, currently in preview, to explain what individual tables, columns, and other schema elements mean. For example, you could describe People as “one row per purchasing customer,” explain that Person_ID is the customer identifier, or document the meaning and units of an abbreviated column. These descriptions are especially useful when object names are ambiguous or highly technical. They are currently available only when the Data Agent uses the preview runtime.

For a concrete example, SQL source instructions could say, “People contains purchasing customers; join SalesHeader to SalesDetail using SalesOrderID.” An example query could pair “How many customers are recorded?” with SELECT COUNT(DISTINCT Person_ID) FROM dbo.People;, assuming those objects and that interpretation match your source. More complex examples can demonstrate date filters, joins, and aggregations that are difficult to explain in prose alone. These are question-query patterns supplied as context, not a separate requirement to retrain a model. Validate both the query and its business meaning, because syntactically valid SQL can still answer the wrong question.

Avoid contradictions, then test the result

Separating responsibilities reduces confusion, but it does not eliminate every possible conflict. Suppose your model’s AI instructions define Revenue as Net Sales Amount while a Verified Answer uses Gross Sales for the same question; those two model-level signals disagree. Or suppose agent instructions route a sales question to a different source than the one containing the approved calculation. Neither problem is solved by assuming one instruction automatically overrides another. Microsoft’s guidance on writing focused AI instructions warns about conflicting guidance, so I would resolve the definitions and routing rather than depend on an undocumented precedence rule. Instructions are guidance, not guaranteed enforcement.

Finally, test questions with known answers and inspect the generated query, not just the wording of the response. Try variations involving customers, clients, revenue, margins, and date periods, then check which source, measure, filters, and relationships were used. Microsoft provides Data Agent evaluation capabilities to support repeatable testing, but business review remains important. My working rule is simple: agent instructions govern source choice and overall behavior; Prep data for AI explains how to understand and query the semantic model. Combined with clean data and sound modeling, that separation gives you a more maintainable starting point than repeatedly adding instructions to whichever box happens to be open.

The post Prep data for AI vs. Fabric Data Agent Instructions: What Goes Where? first appeared on James Serra's Blog.

Original post (opens in new tab)
View comments in original post (opens in new tab)

Rate

★ ★ ★ ★ ★ ★ ★ ★ ★ ★

You rated this post out of 5. Change rating

Share

Share

Rate

★ ★ ★ ★ ★ ★ ★ ★ ★ ★

You rated this post out of 5. Change rating