Data Science – Page 2

The Basics of Bayes’ Theorem

Published 2025-04-09 by Kevin Feasel

In this video, I provide an introduction to Bayes’ theorem, explaining the key concepts and terms, as well as solving a totally realistic problem via Bayesian analysis.

My goal in this video was to explain a counter-intuitive phenomenon: how much a positive test or piece of information moves the needle depends primarily on how frequently the event normally happens and how frequently we generate false positives in tests. The less common the scenario, the less your positive test actually moves the needle.

Comments closed

Converting a CSV to Parquet with DuckDB and Polars in R

Published 2025-04-03 by Kevin Feasel

Michael Mayer makes a swap:

In this recent post, we have used Polars and DuckDB to convert a large CSV file to Parquet in steaming mode – and Python.

Different people have contacted me and asked: “and in R?”

Simple answer: We have DuckDB, and we have different Polars bindings. Here, we are using {polars} which is currently being overhauled into {neopandas}.

Click through for the comparison.

Comments closed

Calculating a Matrix Inversion in SQL Server

Published 2025-04-03 by Kevin Feasel

Sebastiao Pereira performs matrix math in-database:

There are numerous applications to obtain a Matrix inverse for a given Matrix. Is it possible to do it using only SQL Server? Read on to learn how to build a matrix inverse calculator using a set of SQL Server custom functions.

I expect this to be extremely slow in comparison to GPU-based methods using a language like C, but this approach maximizes style points.

Comments closed

Solving Linear Equations in SQL Server

Published 2025-03-11 by Kevin Feasel

Sebastiao Pereira implements a function:

Solving linear equations is essential for solving real-world problems in Science, Engineering, Data Analysis, Machine Learning, Economics, Finance, and other areas. Is it possible to have a tool to solve linear equations directly in SQL Server? We will look at how to create a Gauss-Seidel method function for SQL Server.

This is one way to solve a series of linear equations, and it’s a pretty neat implementation.

Comments closed

Lagrange Interpolation in SQL Server

Published 2025-03-05 by Kevin Feasel

Sebastiao Pereira creates a function:

It is very usual to have a set of discrete data points but sometimes it is necessary to estimate values between those points. Is it possible to create a function to do this in SQL Server?

It turns out that the answer is “yes.” Click through to see how.

Comments closed

Bass Product Diffusion and Data Science

Published 2025-03-04 by Kevin Feasel

John Mount does a fun analysis:

This is a graph of the percentage of Stack Overflow questions tagged with data science terms such as R, Pandas, and so on. It seems to show exploding interest in R and Pandas, and maybe even Tensorflow. Pandas was likely chosen as a proxy for interest in Python for data science (versus a general interest in Python). I’d prefer view counts over question percentages as a proxy of interest, but it is what it is.

Then I thought, let’s see if they have newer data. They do, and it is horrifying (though not unexpected to those of us in the industry).

Click through for the analysis, as well as an important note in the comments.

Comments closed

Inflation in Medieval China

Published 2025-02-10 by Kevin Feasel

Richard Vale digs into a dataset:

In this post, I would like to draw attention to a very interesting data set collected by Guan, Palma and Wu as part of the replication package for their paper The rise and fall of paper money in Yuan China, 1260-1368. The paper describes inflation, money and prices during the Yuan Dynasty era in China.

First, a little historical background.

Read on for the analysis. H/T R-Bloggers.

Comments closed

Running a Regression Analysis in T-SQL

Published 2025-02-10 by Kevin Feasel

Sebastiao Pereira performs an analysis:

Regression analysis is a statistical technique to identify relationships between a dependent variable (outcome) and one or more independent variables (features). Is it possible to do this in SQL Server without using any external tools?

Read on for the answer. I’d file this under “neat but probably not something I’d want to rely upon.”

Comments closed

An Overview of HyperLogLog

Published 2025-01-27 by Kevin Feasel

Bhala Ranganathan talks about a powerful algorithm:

Cardinality is the number of distinct items in a dataset. Whether it’s counting the number of unique users on a website or estimating the number of distinct search queries, estimating cardinality becomes challenging when dealing with massive datasets. That’s where the HyperLogLog algorithm comes into the picture. In this article, we will explore the key concepts behind HyperLogLog and its applications.

HyperLogLog is the algorithm that SQL Server users in the APPROX_COUNT_DISTINCT() function to make it so much faster than a regular COUNT(DISTINCT) while still providing correctness guarantees within a fixed percentage error: they guarantee a 2% or lower error rate with a 97% probability.

Comments closed

The Distribution of P-Values under the Null Hypothesis

Published 2025-01-24 by Kevin Feasel

David Lindelöf asks a question:

I sometimes use this fun interview question for aspiring data scientists:

How are p-values distributed assuming the null hypothesis is true?

Read on for four incorrect answers, followed by the actual answer. I actually got the answer right, but in the sloppiest and laziest way. But hey, it was still the right answer. H/T R-Bloggers.

Comments closed

M	T	W	T	F	S	S
	1	2	3	4	5	6
7	8	9	10	11	12	13
14	15	16	17	18	19	20
21	22	23	24	25	26	27
28	29	30	31

Category: Data Science