Skip to content

euspinolia logo euspinolia logo

euspinolia#

A small CSV/table library for Python with its engine in Zig. Typed columns, filter, sort, groupby — faster than pandas at most of it, with no dependencies and a library that links nothing, not even libc.

pip install euspinolia

Read the API How it works GitHub

In thirty seconds#

>>> import euspinolia
>>> df = euspinolia.read_csv("people.csv")
>>> df[df["age"] > 30].sort_values("score", ascending=False)[["name", "score"]]
    name  score
0    ada   91.5
1  grace   88.0

[2 rows x 2 columns]
>>> df.groupby("city").agg({"score": "mean"})
       city  score
0    London   91.5
1  New York   88.0
2     Paris  73.25

[3 rows x 2 columns]

Or without writing any Python:

$ euspinolia stats data.csv
data.csv: 500,000 rows x 5 columns, 18.0 MB, parsed in 108 ms

column  type      min     max        mean  distinct
------  ------  -----  ------  ----------  --------
id      int         0  499999    249999.5   500,000
name    string                              500,000
dept    string                                    5
salary  int     40000  199999  119926.142   152,970
score   float     0.0   100.0      49.973    10,001

What's inside#

  • Typed columns


    Every column is int, float or string, inferred on parse and stored as one flat array. Numeric columns are read from Python without a copy.

  • Filter, select, sort


    df[df["age"] > 30], df[["name", "age"]], and a stable sort_values that radix-sorts numeric keys.

  • Group and reduce


    groupby(...).agg({...}) hashes keys into dense ids and folds each group in one pass; sums stay exact in 64-bit integers.

  • In and out


    CSV, TSV or any single-character delimiter, both ways, round-tripping types — plus from_dict and to_dict for data already in Python.

Against pandas#

500,000 rows, 5 columns, 18 MB, best of five on one laptop. Shorter is faster; each pair is scaled to the slower of the two.

read + parse
114 ms
183 ms
write CSV
63 ms
478 ms
sort by an int
41 ms
70 ms
groupby, mean
7.5 ms
23.6 ms
filter half
10.5 ms
8.7 ms
euspinoliapandas 3.0

Filtering is the one loss: the result copies its strings so it can outlive its source. The benchmarks page explains every row.

Why it exists#

It is a teaching project with a fixed scope — one header row, three column types, the operations above — small enough to read end to end in an afternoon: about 3,700 lines of Zig with tests, and 1,250 of Python. The internals walk through it module by module, from the CSV scanner to the C ABI.

The name is Euspinolia, the genus of the velvet ant known as the "panda ant".