DuckDB: In-Process Analytics That Replaces Spark for Single-Machine Workloads

DuckDB runs inside your Python or R process with zero setup, queries Parquet files directly with SQL, and outperforms Spark on datasets under 100GB on a single machine.

Mahmudul Haque Qudrati — CEO & ML Engineer at Pristren

Mahmudul Haque Qudrati

CEO & ML Engineer

March 10, 2026
7 min read
DuckDB: In-Process Analytics That Replaces Spark for Single-Machine Workloads

What Is DuckDB and Why Should You Care

One AI engineering post, weekly

LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.

DuckDB is an in-process OLAP database - it runs inside your Python, R, Java, or Node process without a server. No installation, no Docker, no cluster to manage. Just pip install duckdb and you have a full SQL analytics engine.

The core insight: most analytics workloads that teams throw at Spark or BigQuery fit comfortably on a single modern machine with 32-128GB RAM. DuckDB handles these cases with dramatically less complexity.

Query Parquet Files Directly with SQL

DuckDB's killer feature is reading files directly:

python
import duckdb

# Query a single Parquet file
result = duckdb.sql("SELECT region, SUM(revenue) FROM 'sales.parquet' GROUP BY region").df()

# Query all Parquet files in a directory
result = duckdb.sql("SELECT * FROM read_parquet('data/*.parquet') WHERE date >= '2025-01-01'").df()

# Join two files  -  no loading into memory required
result = duckdb.sql("""
    SELECT o.order_id, u.email, o.total
    FROM read_parquet('orders.parquet') o
    JOIN read_csv('users.csv') u ON o.user_id = u.id
    LIMIT 1000
""").df()

DuckDB pushes predicates and projections down to the file reader - it only reads the columns and row groups you need.

Team workspace

Ship faster with chat, meetings, and projects in one place — Zlyqor.

Start free

DuckDB vs Spark

DuckDBSpark
SetupZeroComplex (JVM, cluster manager)
Data sizeUp to ~500GB single machinePetabyte-scale distributed
Query speed (<100GB)Often fasterOverhead from shuffle
CostFreeCompute cluster cost
SQL supportFull SQLSparkSQL (limitations)

For data teams processing 1-100GB per job, DuckDB eliminates the operational overhead of Spark entirely.

Python API

python
import duckdb

con = duckdb.connect("analytics.duckdb")  # persistent database

# Create table from Parquet
con.execute("CREATE TABLE events AS SELECT * FROM read_parquet('events/*.parquet')")

# Query with pandas interop
df = con.execute("SELECT date_trunc('day', ts) AS day, count(*) FROM events GROUP BY 1").df()

# Register a Pandas DataFrame as a virtual table
import pandas as pd
users_df = pd.read_csv("users.csv")
con.register("users", users_df)
result = con.execute("SELECT * FROM users WHERE country = 'US'").df()

MotherDuck: DuckDB in the Cloud

MotherDuck is managed DuckDB with a cloud UI, team sharing, and hybrid execution (local + cloud). It uses the same DuckDB SQL dialect. Useful for teams that want DuckDB's simplicity with cloud storage and collaboration.

DuckDB + Polars Combination

python
import duckdb
import polars as pl

# Read with DuckDB, convert to Polars for transformation
result = duckdb.sql("SELECT * FROM 'large.parquet' WHERE status = 'active'").pl()

# Register Polars LazyFrame in DuckDB
lf = pl.scan_parquet("data/*.parquet")
duckdb.register("my_data", lf.collect().to_arrow())
result = duckdb.sql("SELECT category, AVG(value) FROM my_data GROUP BY category").df()

When to Use DuckDB

DuckDB wins for: ad-hoc analytics on files, replacing SQLite for analytical queries, embedded analytics in Python apps, replacing pandas groupby/merge on large files. Use Spark when: data exceeds single-machine RAM, you need distributed fault tolerance, or you already have a data platform built on it.

Resources: DuckDB docs, GitHub.

#duckdb#analytics#sql#columnar#olap

// discussion

Comments

0/4000
Mahmudul Haque Qudrati — CEO & ML Engineer at Pristren

Mahmudul Haque Qudrati

CEO & ML Engineer

Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.

PristrenZlyqor