Press "Enter" to skip to content

Category: Python

Adding Temporal Reasoning to RAG

Ivan Palomares Carrascosa checks the date:

Topics we will cover include:

  • How to extend standard subject-predicate-object triples into time-stamped quadruples stored in a simple temporal graph.
  • How to calculate recency weights with exponential decay and use them to rank conflicting facts as of a given query date.
  • How to tune the half-life parameter and integrate the temporal graph into a deterministic 3-tiered Graph-RAG retrieval pipeline.

This is especially important if you’re searching over the news or other systems that value more recent information over older information.

Leave a Comment

What’s New in Python 3.15

Russ Hyde has some new release notes:

It’s October, so that means a new stable release of Python has arrived. Here we will review some of the changes introduced in Python 3.15.

Prominent changes in Python 3.15 that we will cover:

  • profiling & statistical sampling
  • lazy imports
  • unpacking * in for comprehensions

Click through to learn more about what’s out. Though I will say that I’m mostly on 3.13 and sometimes 3.14 because it can take a while for package makers to catch up.

Leave a Comment

SQLAlchemy 2.1 and mssql-python Support

Matt Hyon and David Levy share some good news:

We’re pleased to share that SQLAlchemy 2.1.0 is now generally available, with built-in support for mssql-python, Microsoft’s Python driver for SQL Server. You can now use it with SQLAlchemy’s ORM and Core APIs through the first-party mssql+mssqlpython dialect.

If SQLAlchemy is already part of how you build, we want this to feel like a natural next step – not another thing to learn. The goal is simple: help you connect to SQL Server and Azure SQL with less setup, while keeping the tools and patterns you know.

The mssql-python library is worth it over PyODBC.

Leave a Comment

Chain-Ladder Reserving Calculations in Python

Christian Lorentzen digs into loss reserving:

Ask a reserving actuary how they run a Chain-Ladder and you’ll usually hear “Excel” or the name of a pricey specialized tool. It turns out a modern dataframe library handles it just as well — in a few lines, for hundreds of companies at once.

We use the CAS loss reserving data, specifically the other liability line of business (LoB): 233 US insurers (“GRNAME”), 10 accident years (1998–2007), paid and incurred losses at every development lag (1-10). We treat 2007 as our reporting year, i.e. we simulate a year-end closing.

Click through for a demonstration and a comparison against R’s ChainLadder package.

Leave a Comment

Working with Vectors in Python

Bala Priya C avoids the loops:

In this article, you will learn how to think in terms of vectorized operations using NumPy, replacing slow Python loops with efficient array-level computations.

Topics we will cover include:

  • Why Python loops are slow for numeric data and how NumPy’s C-backed engine addresses this.
  • How to apply element-wise operations, boolean masking, and broadcasting to eliminate common loop patterns.
  • How to handle multi-condition branching and axis-based aggregation entirely with NumPy functions.

This is one of those places in which people with database development backgrounds can end up understanding the topic more intuitively than loop-heavy structured programming developers. And depending on how large the loop is and how complex each operation is, there can be a significant performance improvement in applying functions over a vector versus in a loop. It’s part of why matrix operations tend to be much faster than nested loops.

Comments closed

Semantic Link Labs UI Updates

Chris Webb takes a look:

There’s so much going on in the Fabric community that it can be hard to keep up with it all. Semantic Link Labs is a great example: in the six months or so since I last had a proper look at it my colleague Michael Kovalsky has done a whole load of cool things and it wasn’t until I had a chat with him recently that I realised how much had changed. Most importantly, for someone old-fashioned like me who still likes tools with a UI, a lot of new functionality has been added which has a UI and is usable with minimal coding.

Click through to see what’s available.

Comments closed

Merging Data into a Fabric Lakehouse via Python Notebook

GIlbert Quevauvilliers uses a pure Python notebook:

In this blog post I am going to show you how to use a Fabric Python runtime notebook (This is the notebook which only uses Pure Python functions and consumes significantly lower Capacity Units (CUs)).

The pattern is how to get new data and merge it into an existing Lakehouse table. This ensures that if the notebook is run again data will not be duplicated.

Why I am sharing this is I have found that there is not a lot of useful information about how to use a Python notebook to write to a lakehouse table easily. And then also how to use a Merge statement making it easier to insert or update your lakehouse tables. This simplifies the ingestion process, runs faster and consumes the least amount of CUs

Gilbert doesn’t mention it in the blog post but the notebook does use DuckDB to query the data using SQL.

Comments closed

Testing SQL Code in Python

Jamal Hansen writes some tests:

I once had a query that ran fine for months. Then someone added a column to the source table and a SELECT * downstream started returning unexpected data. The query didn’t error. It just silently gave wrong results. A test would have caught it immediately.

Schema changes break queries silently. Refactoring a CTE can shift results in ways you don’t notice. New data patterns expose assumptions you didn’t know you made. SQL deserves the same testing discipline as the rest of your code, and Python makes it straightforward.

PyTest, the library Jamal uses here, is one of my favorites for this kind of work. You can build up tests without a lot of ceremony and it’s pretty easy to deal with for most use cases.

Comments closed

Estimating Probabilities from Unevenly Collected Data

Nina Zumel answers an important question:

In this article, we look at the problem of estimating and comparing probabilities about a population of subjects from unevenly collected observations. Some examples might include:

  • The perceived quality of a movie (how often is a movie positively reviewed) when some movies have far more reviews than others.
  • The effectiveness of various ad campaigns, when some compaigns have had more exposure than others.
  • The efficacy of a certain medical procedure by hospital, when some hospitals have had more cases than others.

For our specific task, we’ll try to estimate the “innate” batting ability (the probability of making a hit when at bat) of major league baseball players in 2023. For the sake of this article, we will take this single season of data as everything that we know about these players and their batting statistics.

It’s an interesting problem because she’s looking at 2023 data as an estimation of the player’s entire career, with the goal of estimating how a player will perform overall given a fairly reasonably sized sample of information collected from one relatively short period of that player’s career. H/T John Mount.

Comments closed

A Look at Tabular Foundation Models

Michael Mayer tries out a neural network model:

Tabular data has had a comfortable life for years. Gradient boosting showed up, got very good at its job, and then quietly became the default answer to almost everything with rows and columns.

In very recent years, a new player has arrived: the tabular foundation model or prior fitted neural network, and suddenly tabular data is sounding a lot less sleepy…

I’ve done a bit with TabPFN and come away fairly impressed. I’ll have to give this a go as well. There are definite limitations to data sizes before things fall over, but for moderate sizes (50k or fewer rows), TabPFN at least worked pretty well.

Comments closed