Skip to main content

One post tagged with "Apache DataFusion"

View All Tags

We Benchmarked Apache DataFusion Comet on Iceberg — 1.51x Faster on 45% Less CPU, and the Version Number That Almost Fooled Us

· 12 min read
Cazpian Engineering
Platform Engineering Team

Native acceleration on Iceberg, measured

Every native execution engine for Spark arrives with the same headline: 2x, 3x, sometimes 5x faster. The benchmarks are usually real. What they rarely tell you is whether they describe your data, your table format, and your version matrix.

So we measured it ourselves. Apache DataFusion Comet, against 200 million rows of Apache Iceberg data, on the Spark version we ship. We recorded wall-clock time, CPU seconds, peak memory, and — the part most benchmarks skip — whether the answers came back the same.

The result: 1.51x faster on 45% less CPU, with every result identical to standard Spark. On the best query in the suite, 3.18x faster on 80% less CPU.

We also nearly published the exact opposite conclusion. Our first full round measured zero acceleration on Iceberg — not "a little", literally zero accelerated operators across five configurations. That result was real, reproducible, and completely misleading, and the reason it was wrong is worth more to you than the speedup number.