Press "Enter" to skip to content

Category: Microsoft Fabric

Spark SQL Temporary Views on Fabric Schema-Enabled Lakehouses

Gerhard Brueckl wants to create a temporary view:

Some time ago my friend Christian Henrik Reich blogged about how to handle schema-enabled lakehouses in Spark temporary views. We already found a good solution leveraging SQL USE keyword to set the schema once and reference tables by name only in our views. However, after some tests, in particular with notebookutils.notebook.runMultiple, I realized that there are some more things to consider as suddenly my SQL and notebooks stopped working when executed in parallel!

But let set the scope for this blogpost first. We recently built a data platform on Microsoft Fabric where we integrated data from different source systems which were combined into a single schema-enabled lakehouse where each source system had it own schema. Naturally, when querying those schemas, we used temporary views to express our business logic in SQL and then continue working with PySpark. We relied heavily on USE to work around the issue describe by Henrik until we realized that this approach does not work in combination with runMultiple. So we had to find another solution which I will describe here.

Read on for that solution.

Leave a Comment

Atomicity and Isolation in Fabric Data Warehouse

Louis Davidson receives a surprise:

Ironically, today’s topic is kind of the opposite. I expected things to be far different in Fabric, but it isn’t really that different, at least not in behavior, except when it is. Since this system is Parquet file based, I didn’t think there would be locking, blocking, etc. I sort of expected it would be sort of locked down when writing data, maybe just single threaded per file, but highly concurrent when reading. Reads would most likely work like time travel and read previous data where it existed. And transactions? Would there be transactions? I guessed not before I got started with it.

Was I wrong?

Read on for the answer. The concurrency model isn’t exactly the same as SQL Server’s, but it’s not that far off.

Leave a Comment

Referencing Assets in Power BI Reports via OneLake URLs

Chris Webb wants to load some images:

Several years ago I wrote a very popular blog post about how to store images for your reports inside your Power BI semantic model. It solved the problem of how you could use display images (for example of products) inside your reports without making those images available via a public URL or personal OneDrive Embed Codes. I was very proud of how efficient the M code to do this was but the code was complicated and storing images as text inside a semantic model makes refreshes a lot slower and increases the size of your semantic model in memory, so it’s not ideal. The good news is that, if you have enabled Fabric in your tenant, the August 2026 release of Power BI brings a much better way of solving this problem: you can now store your images (and indeed other files) inside OneLake and reference them from there. This means you can store your images in a secure location, alongside all of your other data, and make them available for use in Power BI. What’s more this doesn’t just work for images, it also works for other types of files such as GeoJSON files used by map visuals.

Click through to see how.

Leave a Comment

Dynamically Changing Fabric Data Warehouse SQL Pools

Gilbert Quevauvilliers saves some money:

After reading about the new SQL Pools feature for Warehouses in Fabric, I had an idea, if I could change the SQL Pool configuration based on the expected query load, I could then consume less capacity and have better performance.

https://learn.microsoft.com/en-us/fabric/data-warehouse/custom-sql-pools

Here is an Example I thought of below.

  • When the ETL load is running optimize the SQL pool for writing as typically data is being inserted.
  • After the ETL load and for the rest of the day, almost all queries are read by the Warehouse, so change the SQL pool to be read optimized.

Click through for a Python notebook that does this.

Comments closed

INFORMATION_SCHEMA and the Fabric Warehouse

Louis Davidson bangs his head against a wall:

When we decided to use T-SQL and a Fabric Data Warehouse for our ETL, I started thinking about generating the code with the metadata in the system catalog views or the INFORMATION_SCHEMA. Having done this sort of thing before in SQL Server over the years, it seemed really straightforward. And it kind of is, until it isn’t.

In this blog I want to show you a few ways you need to understand how working with metadata and temp tables varies (sometimes wildly) from the comfortable SQL Server environment and language you know very well, and give tips on how to get around these differences.

Click through for some of the fun you can have with a distributed SQL Server-like product.

Comments closed

What to DO When Out of Capacity in Microsoft Fabric

James Serra hits the ceiling:

Microsoft Fabric makes it wonderfully easy to put many analytics workloads on one platform. Power BI, data engineering, warehousing, data science, real-time analytics, Copilot, and other experiences can all share the same Fabric capacity. That is a big advantage, but it also creates an architectural question that does not get much attention until something goes wrong: what should you actually do when a capacity starts running out of room? The answer is not always “buy a bigger capacity.” Sometimes you should optimize, sometimes scale up, sometimes scale out, and sometimes isolate the workload causing the problem.

Read on for an explanation of what it means to “run out of capacity” and a quick overview of four options available to you.

Comments closed

The Pain of Microsoft Fabric Deployment Pipelines

Meagan Longoria has a list:

Fabric deployment pipelines look like they solve CI/CD for Fabric content, especially for people who prefer a GUI over writing code to handle deployments, but the implementation has enough structural gaps that they fall apart for several real deployment workflows. Even when they do work, the UI isn’t always intuitive.

Every piece of active software carries a backlog of feature requests and known limitations, and deployment pipelines get new capabilities on a regular basis. Everything below reflects how deployment pipelines behave as of August 2026. Some of it may have changed by the time you’re reading this, so check Microsoft’s docs for the current state before you plan around any of these.

Click through for the list, as well as a few alternatives that come with their own trade-offs.

Comments closed

Semantic Link Labs UI Updates

Chris Webb takes a look:

There’s so much going on in the Fabric community that it can be hard to keep up with it all. Semantic Link Labs is a great example: in the six months or so since I last had a proper look at it my colleague Michael Kovalsky has done a whole load of cool things and it wasn’t until I had a chat with him recently that I realised how much had changed. Most importantly, for someone old-fashioned like me who still likes tools with a UI, a lot of new functionality has been added which has a UI and is usable with minimal coding.

Click through to see what’s available.

Comments closed

Validating Selective Deployments in Microsoft Fabric

Matt Collins shares some advice:

Selective deployments with the fabric-cicd python library are highly useful for shipping just the items that actually changed in your CI/CD process. Unfortunately, by default, it also tells you that a deployment succeeded when nothing was deployed at all.

This blog showcases an example where we selectively deployed Fabric items to upper environments, not realising that the dev ops pipeline reported as “successful” but did not contain our intended changes. We will then dig into some quality checks that highlighted a bigger error in the way Microsoft Fabric handles item naming in a repository.

You’ll learn how to help safeguard from human error in selective Fabric deployments, as well as create useful build validations in Azure DevOps that help to keep your repository resource names clean and fit-for-purpose. All this is achieved through two simple CI/CD pipelines.

Read on to learn more.

Comments closed

Surrogate Keys in Fabric Data Warehouse

Louis Davidson does some more digging:

Creating the simplest of tables, the first thing I start thinking about is making sure that the data is going to be protected from the user. Users (including myself on my own projects when I have my user hat on) don’t notice when they start inserting poor quality data sometimes. Oops, I hit F5 twice, I wonder if that will affect my data? Without proper constraints, it probably will.

In this second entry, I want to cover a few things about handling surrogate values that will help you avoid some of the (in retrospect) kind of dumb expectations that I had. Fabric Data Warehouse T-SQL feels so much like SQL Server Relational T-SQL that some stuff like choosing a surrogate key makes me think I am missing something.

One of the things Louis mentions is how the values get inserted and how it looks like they’re in different ranges. This makes sense, as each distributed node likely has its own range of identity values, similar to the way merge replication would work with identity keys to prevent overlap.

Comments closed