Press "Enter" to skip to content

Category: Architecture

Defining a Data Science Pit of Success

Bruno Rodrigues thinks about robustness in data science:

Rico Mariani, a long-time performance engineer at Microsoft, coined the idea of the pit of success. A system has a pit of success when the natural, lazy way of using it leads to good outcomes. You don’t need heroics or perfect memory. You fall into the right result, and climbing out to do something wrong takes deliberate effort.

I came across this framing in a talk that I was recently recommended and I further recommend it to anyone interested in the topic: Functional architecture – The pits of success – Mark Seemann. It put a name to something I had been circling around for years.

This is worth a careful read if you’re tied in with a data science team. H/T R-Bloggers.

Leave a Comment

Thoughts on the Shared SQL Server Database

Kendra Little shares some thoughts:

For more than 20 years, many small and medium SaaS companies have built products with complicated business logic and a flexible user experience on the .NET stack using an architecture with a shared SQL Server database. This approach has kept the production environment relatively simple.

These companies have new incentives to move away from this architecture because agentic development lands new business rules on the shared database fast enough to painfully cut velocity and add significant risk to deployments.

I’m not sold on the idea, but I think Kendra’s post is definitely worth the read and some noodling.

Leave a Comment

Bad Data in the Silver Layer

Andy Brownsword answers a question:

When combining and curating data in your Silver layer it’s not uncommon to find entries don’t always fit quite right. It could be missing references, incomplete data, or broken validation rules.

Should that bad data be allowed into your carefully crafted Silver layer? And if not, where should that data go?

Let’s look at some options for handling the bad data and how they compare.

This is a harder question to answer than it would first appear. In a perfect world, you’ve applied all of the data cleansing logic and you have pristine data with no errors. Also, you ride unicorns to work in the Fields of Elysium.

Leave a Comment

The Benefits of Database Cost Optimization

Chad Timms lays out some benefits:

Flexera’s 2026 State of the Cloud Report found that estimated waste in cloud infrastructure and platform spend rose to 29% this year. That is the first increase in five years. Managing cloud costs remains the top challenge for 85% of respondents. Your database estate sits inside that cloud bill through licensing, capacity, and the people who keep it running. It is rarely examined line by line. Database cost optimization seldom fails for lack of effort. It fails because the spend gets treated as a purchasing problem when it is really an operating one.

Renewals get negotiated. Cloud service tiers get compared. Meanwhile, the decisions that actually set the number are made on the ground, often by whoever is on call that week. Below are seven questions a finance or IT leader can put to their own team, or to a provider, along with what a strong answer and a weak answer sound like.

The target of this post is more for managers versus line employees, but it’s good to think about how you would answer the questions in the post.

Leave a Comment

Against Using a Single Database for Everything

Pat Wright lays out a case:

I’ve seen quite a few posts lately about how PostgreSQL can do everything. And, while I do believe it’s the most advanced and fastest-growing relational database available right now, that doesn’t change the fact that – at its core – it’s still a relational database.  

This thinking goes all the way back to the mid-2000s, when I was working on SQL Server and starting to explore the NoSQL movement with Elasticsearch, Hadoop, Hive, HBase, and their respective tools.

Even back then, people said these new technologies could solve every problem. Take it from me: please don’t use any technology for every problem you have.

Read on for Pat’s argument. Pat has tailored this one specifically for PostgreSQL but you can swap out a few things and have it apply to pretty much anything.

Leave a Comment

Ten Decisions to Make Prior to Creating a Fabric Workspace

Meagan Longoria has a list:

Creating a Fabric workspace takes about 30 seconds. Restructuring workspaces after people have built content in them takes a lot longer. Moving content to another workspace usually means redeploying or recreating it, and anything bound to the old item IDs has to be rebound: reports connected to a semantic model, a notebook’s default lakehouse, pipeline activities, and shortcuts. Item-level shares, app content, and links people saved don’t carry over either.

Most of that rework is avoidable if you make a handful of design decisions before anyone starts building. None of these decisions has a single right answer. The right choice depends on your organization’s size, skills, security requirements, and how much self-service you want to support. But I’d much rather see them made deliberately, before the first workspace exists, than left at the defaults and discovered later. For each decision, I’ll cover the options, what should drive the choice, and the platform constraints that narrow it down.

Read on for that list.

Leave a Comment

Building Business versus Data Apps in Microsoft Fabric

Soheil Bakhshi builds an app:

While building that application, I came across a few gotchas and limitations around its architecture. But there is also another Fabric Apps pattern that takes a quite different approach. Instead of creating an operational database, it uses the relationships, measures and business logic we already have in a Power BI semantic model.

Microsoft calls this the Data App template. In this article, I compare it with the operational pattern from my previous exercise, which I refer to as a Business App. They both run on Fabric Apps, but what happens underneath is different. Their security requirements, sharing model and licensing are different too. So, let’s go through what I tested, what caught me out, and where I think each pattern makes sense.

Read on to learn more.

Leave a Comment

A Medallion Architecture Primer

Andy Brownsword takes us through a concept:

Medallion architecture is the go-to for handling analytical data, particularly with the prominence of lakehouses in Databricks and Fabric.

Whilst the layers are well defined, their definitions don’t tell you where particular tasks belong. So here I wanted to present my view from a more practical perspective of what goes into each layer.

What I’ve found interesting is that, despite almost everyone agreeing on the rules of what goes where, you can easily get into debates with other data architects and lakehouse practitioners once you start asking concrete questions around specific datasets and specific levels of transformation. The edges between silver and gold get really fuzzy.

Leave a Comment

Contrasting LTAP and HTAP

Paul Andrew notes another convergence of OLTP and OLAP:

At the Data and AI Summit in June this year, Databricks introduced Lake Transactional/Analytical Processing, or LTAP. My first reaction, I’ll admit, was cynicism, just for a change. Every software vendor must invent new names for old things these days. And for those of us who have worked in data architecture long enough to have the grey hair to prove it, like me, this felt very familiar. Hybrid Transactional and Analytical Processing, or HTAP, a term Gartner coined back in 2014, has long advertised bringing operational transactions and analytics “closer together”. In my experience it never really became a strong implementation pattern in data platform deliveries.

Read on to learn how LTAP and HTAP differ, where Databricks and other competitors (like Microsoft) are going in this space, and how much they have yet to do.

Leave a Comment

Digging into Entity-Attribute-Value Tables

Greg Low has started a series on entity-attribute-value tables. The first post covers what they are:

If you’ve been working with databases for any length of time, you will have come across implementations of Entity-Attribute-Value (EAV) tables (or non-tables as some of my friends would call them).

Instead of storing details of an entity as a standard relational table, rows are stored for each attribute.

The second post covers pros and cons:

In an earlier post , I discussed the design of EAV (Entity Attribute Value) tables, and looked at why they get used. I’d like to spend a few moments now looking at the pros and cons of these designs.

Greg is very much against EAV, and I agree with this. I do like Greg’s alternative of using something like JSON, with the proviso that the database simply become a whole-record storage and retrieval engine rather than trying to strip out and splice in new JSON via T-SQL. Otherwise, spend the time on proper data modeling and take advantage of what the platform can do for you.

1 Comment