AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Polars says version 2.0 makes its streaming engine the default for LazyFrame queries and enables initial spill-to-disk support, changes intended to reduce memory pressure. The release also expands SQL support and adds a Map data type; performance comparisons cited by the developers come from their own benchmarks and have stated limitations.

Polars has released version 2.0, making its streaming engine the default when users call collect on a LazyFrame and enabling initial spill-to-disk support. The changes affect how queries execute and handle memory, while the release also promotes SQL as a first-class interface and adds a native Map data type, according to the Polars project.

Under the new default, lazy queries run through the streaming engine, which Polars says can provide memory and performance improvements on many workloads. The trade-off is that some operations, including joins, group-bys and unpivots, do not guarantee observable row order by default. Users who need that behavior can set maintain_order=True for supported operations, the release post says.

Version 2.0 also enables out-of-core processing by default for operations that currently support it, including sorts, window functions and many expressions. Polars says the system begins spilling data to disk at about 80% of RAM, a threshold the team says may need tuning, and sets a default disk budget of 64 GB. Joins and group-bys are not yet supported for spill-to-disk; the team says those are planned for later.

The release introduces a Polars Map dtype that directly supports Arrow MapType data. Previously, the project represented that data as a list of structs containing key and value fields. The updated type includes dictionary-like operations such as retrieving a value by key, checking whether a key exists, and returning a map’s keys or values.

At a glance
announcementWhen: Released; the source report does not gi…
The developmentPolars has released version 2.0, changing the default execution engine for lazy queries and introducing initial out-of-core support alongside SQL and data-type updates.

Memory and Ordering Changes

The execution changes may matter most to users running queries on datasets that put pressure on available memory. With supported operations able to spill data to disk, a query may be able to finish without keeping all intermediate data in RAM. Polars describes this as making the engine more resilient for high-memory workloads, but the release does not claim that every query or operation can now run out of core.

Making streaming the default also means existing code can behave differently where results previously depended on row order. Users with order-sensitive workflows may need to review joins, group-bys and unpivots and request order preservation where required. The practical effect will depend on each workload, and the project’s release notes do not provide a universal performance guarantee.

For teams choosing a query engine, the expanded SQL interface and benchmark results offer a reason to test Polars against their own data. The project reports strong results against DuckDB and DataFusion on selected TPC-H and TPC-DS tests, but those findings are based on a specified setup and are not independent evaluations.

Amazon

high performance data processing laptop

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

SQL Benchmarks and Test Setup

Polars says the 2.0 release was not intended to be a major feature release, but describes it as a point where SQL becomes a first-class citizen alongside its existing engine. The team highlights optimizer and engine work, including join reordering, common-subplan elimination and dynamic predicates or bloom filters, as part of the effort to support SQL workloads efficiently.

For its comparison, Polars ran SQL queries generated with DuckDB’s TPC-H and TPC-DS query tools against Polars, DuckDB 1.5.6, a DuckDB 2.0 development build, and DataFusion 54.0.0. Tests used two AWS machine configurations: one with 16 virtual CPUs and 32 GB of memory, and another with 192 virtual CPUs and 384 GB. Each query ran five times in a hot setting, and the best time was used; the file cache was cleared between engines and benchmarks, but not between queries.

Polars reports that it completed all queries and was fastest on all but one benchmark in its default configuration. The project also reports that DataFusion timed out on TPC-DS query 72, timed out once on query 67, and ran out of memory on TPC-H query 18 on the smaller machine; those queries were excluded from the comparisons for all engines. Polars notes that its default configuration had overhead on the 192-thread machine for small queries and says limiting it to 32 cores was competitive or faster across the benchmarks. The team has published a repository for reproducing the tests.

Amazon

SSD external storage drive 64GB

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits Still in the Release

Out-of-core support is partial: the release post lists sorts, window functions and many expressions as supported, while joins and group-bys remain on the roadmap. The project does not give a release date for those additions or quantify how often supported workloads will spill at the stated threshold.

The benchmark results should also be read within their stated conditions. They reflect particular hardware, query-generation methods, data and cache settings, and a best-of-five timing method. Some DataFusion queries were excluded after timeouts or an out-of-memory result, and the source does not establish how the results translate to other systems, datasets or production workloads. Polars says it has diagnosed the scaling overhead it saw on the 192-thread machine, but says a fix is hoped for in a future release.

The source report does not provide a calendar release date or detailed compatibility guidance for existing projects. Users will need to check the version’s full documentation and test order-sensitive queries and workloads before upgrading.

Amazon

large capacity RAM disk for data analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing and Planned Engine Work

Polars has invited readers to reproduce its benchmark results through a repository linked in the release post. For users evaluating the update, the immediate next step is to test representative queries under version 2.0, paying particular attention to memory use, runtime and any reliance on row order.

The project says it plans to extend spill-to-disk support to joins and group-bys, though it gives no schedule. It also says it hopes to address the overhead observed when scaling to 192 threads in a future release. Until those changes arrive, both the supported operation list and the benchmark caveats remain relevant to decisions about adopting the new defaults.

Amazon

SQL data analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main change in Polars 2.0?

LazyFrame queries now use the streaming engine by default when collected. Polars says this can improve memory use and performance for many queries, though behavior and results depend on the workload.

Does Polars 2.0 spill every query to disk?

No. The release enables spill-to-disk support for certain operations, including sorts, window functions and many expressions. Joins and group-bys are not yet supported, according to the project.

Can results appear in a different row order after upgrading?

They can for some operations. Polars says the streaming engine does not guarantee observable row order by default for operations such as joins, group-bys and unpivots. Users can set maintain_order=True where they need order preserved and the operation supports it.

Are the Polars 2.0 performance results independent?

No. They are benchmarks reported by the Polars project, using specified TPC-H and TPC-DS tests, hardware and run conditions. The team published a repository for reproducing them; results may differ on other workloads and machines.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Bottleneck Moved: Inside Anthropic’s Expansion of Project Glasswing

Anthropic is adding about 150 organizations to Project Glasswing after partners found 10,000-plus severe flaws with Claude Mythos Preview.

The Compute Reckoning: Anthropic Finally Admits What Customers Suspected for Ten Months

Anthropic reveals that recent customer experience issues were due to compute shortages, now addressed by a major partnership with SpaceX and other providers.

SpaceX Starship V3’s first test flight was largely successful

SpaceX’s Starship V3 completed its first test flight, despite engine issues, marking a significant step toward future lunar and Mars missions.

OpenAI Connects The Dots

A Platformer columnist praised OpenAI’s new Dots agent after limited testing, while questions remain about access, privacy and reliability.