Arrow Database Connectivity for Analytical Workloads
ADBC eliminates the serialization overhead that makes analytical queries slow on the client side.

A columnar database finishes a query in seconds. The client application, waiting on the other end of a JDBC or ODBC connection, doesn't see the results for 40 seconds more. That gap, between when the engine is done and when the data actually arrives, is the subject of this piece: what causes it, what it costs, and what Arrow Database Connectivity (ADBC) does to close it.
Why JDBC and ODBC impose a serialization round-trip on analytical results
JDBC and ODBC were built for a different era of database work: transactional systems returning a handful of rows to an application that needed them right away. Their result-transfer model reflects that job. Both standards move data row by row, one record at a time, because that's how a bank balance or an order lookup behaves. An analytical engine doesn't store or produce data that way. It works in columns, scanning and computing over whole vectors of values at once. That mismatch, row-oriented transport serving a column-oriented engine, produces the serialization round-trip described throughout this piece.
A columnar engine finishes its scan and holds the answer in column vectors, values for a single field sitting together in memory, and that's the trace worth following closely. To hand that answer over JDBC or ODBC, the engine first pivots those columns into rows, one record at a time, so the wire protocol can carry them. The driver then serializes each row field by field across the network. On the other end, the client, whether that's pandas, Polars, Spark, or any other analytical library, takes those rows and rebuilds them back into columns, because that's the shape the downstream code actually needs. Data that was columnar when the engine produced it and columnar again when the application consumes it spends the trip in between flattened into rows, twice converted, for no reason tied to the data itself.
This cost is invisible in the places people normally look for performance problems. Each stage, the pivot, the serialization, the rebuild, runs on CPU and mostly on a single thread per connection. Server-side query timing stops at the point the engine finishes producing results. The database reports the query as done. The application hasn't received a single row yet. Monitoring built around server-side metrics simply doesn't see the 40 seconds that follow, because by the server's clock, the work is already over.
A second cost is buried in the same design. A JDBC or ODBC connection is a single stream. None of that parallelism survives the handoff to the client: everything funnels through one socket, and a distributed execution that took seconds in parallel becomes a serial download on the way out. A 2017 CWI paper, "Don't Hold My Data Hostage," measured exactly this across common databases and found that client-side serialization and driver overhead dominated total end-to-end time for large results, in some cases by an order of magnitude more than the query itself took to run.
What the transpose tax costs in CPU and memory, at scale
Call it the transpose tax: the CPU and memory spent converting columns to rows and back, for data that never needed to leave its columnar shape in the first place, in a huge number of small, cheap steps repeated for every single cell in the result set. A row-oriented driver has to box each value, create an object for it, track it, eventually garbage-collect it, and do that once per cell. The allocation count climbs fast when multiplied by the number of rows and columns in a real analytical result. Memory pressure builds. Garbage collection pauses start eating into wall-clock time. None of this depends on a slow query or a badly tuned engine; it's a direct function of how many cells need boxing, and it scales linearly with result size.
This is a common failure mode that shows up under ordinary conditions. A notebook pulling a feature table for model training hits it. Any ETL job extracting data out to Parquet hits it. Any dashboard pulling raw detail down for local aggregation hits it. These are everyday analytical patterns, and at production volumes, the driver's row-by-row conversion becomes a real fraction of total time spent waiting.
A similar cost shows up one layer down, at storage, because the shape of the problem is identical. ETH Zürich's DaMoN '26 paper, "Should I Hide My Duck in the Lake?," measured what happens when analytical queries read data straight off Parquet files. The parallel to the transpose tax is close to exact. Parquet decoding is a format-translation cost paid between disk and compute. Row-boxing in JDBC is a format-translation cost paid between compute and the client. Both tax the CPU at a boundary where one columnar system has to prove its data to another, through a format neither one actually wants to use internally.
The same research found that pre-filtered, pre-decoded data dramatically outperforms scanning raw Parquet files end to end. The lesson generalizes past storage: when a translation step at a boundary gets removed instead of optimized, the cost tied to that step drops close to zero. That's the logic ADBC applies to the client connection. Removing the row format at the network boundary eliminates the CPU time spent boxing and rebuilding values.
How ADBC eliminates the round-trip: the zero-copy path from engine to application
ADBC answers the transpose tax by refusing the row format at every boundary. Data stays columnar from the engine's internal representation, across the network, through the driver, and into the application. There's no pivot to rows and no rebuild back into columns, because nothing ever leaves the columnar shape to begin with.
Apache Arrow's in-memory format is what makes this work. Two systems that both speak Arrow hand over a pointer to the same bytes rather than translating data between each other when they exchange it.
Two mechanisms make that handoff real rather than theoretical. Arrow's C Data Interface lets two libraries inside the same process share Arrow data with zero copies, nothing allocated, nothing rebuilt. There's no parsing step on arrival, because the format on the wire and the format in memory are the same bytes.
ADBC is the application-facing piece built on top of that foundation. The data arrives typed, columnar, and ready for vectorized processing, with no transpose step anywhere in the path.
Arrow Flight SQL and ADBC solve different parts of this problem. Flight SQL is the wire protocol, the way queries and results actually travel over gRPC. ADBC is the API a developer calls to talk to the database. The zero-copy path is most complete when both layers speak Arrow end to end, but ADBC still delivers a real gain against non-Arrow backends, because the driver handles whatever conversion is still needed in optimized native code (C, C++, Go, or Rust) instead of inside a Python or Java client.
Arrow Flight also fixes the second problem JDBC's design creates: the single-stream bottleneck. On the ADBC side, a call to AdbcStatementExecutePartitions returns partition descriptors mapped to those same Flight endpoints, and a client can read locality information off them to schedule workers close to where the data actually lives.
Where in the ADBC specification and driver ecosystem this is available today
None of this matters if it only exists on paper. It doesn't: the ADBC specification is stable, and the driver list covers most of the databases teams are already running.
ADBC lives under the Apache Arrow project, governed by the Apache Software Foundation. It first shipped in 2023, and the specification itself has reached version 1.1.0, adding query cancellation, richer error metadata, and a statistics API along the way. The driver libraries version separately from the spec and move fast: a new release landed July 28, 2026, following version 22 in January 2026 and version 23 in April 2026. That pace points to a project in active production use rather than one still finding its footing.
Recent releases have filled in gaps that matter for real deployments. A statistics API landed in the Python bindings.
The driver list itself is the strongest evidence of maturity. A March 2026 Spark feature request proposing a native ADBC data source lists mature native drivers already available for PostgreSQL, SQLite, DuckDB, Flight SQL, Snowflake, BigQuery, MySQL, SQL Server, and Databricks. That's the list a production engineering community judged complete enough to build a Spark integration on top of, not a wishlist of future coverage.
Adoption outside the Arrow project itself backs this up. Dremio, working with Microsoft, adopted ADBC for Power BI's connector to Dremio, putting the zero-copy path inside an enterprise BI product running at real scale. Apache Doris added Flight SQL support in version 2.1 and reported tens-fold speedups for data transfer compared to its prior connectivity path. Put together, a practitioner can reach for ADBC today for the databases that matter most, and trial it against a real workload in an afternoon rather than wait on a roadmap.
Which workloads gain the most and which should stay on JDBC
ADBC isn't a universal replacement for JDBC, and the case for it only holds where the transpose tax actually bites. The tax is a function of result size and destination format. Large columnar pulls landing in a columnar consumer pay it heavily. Small transactional lookups pay little of it.
The workloads where ADBC's gain is large and immediate share a shape: notebooks pulling feature tables into pandas or Polars, ETL jobs extracting data out to Parquet, dashboards pulling raw detail down for local aggregation, observability pipelines ingesting time-series telemetry. In each case the driver itself is the bottleneck, and the size of the win scales with the size of the result.
The workloads where ADBC doesn't help much, or isn't the right fit, share a shape too: OLTP applications, small lookup queries, CRUD operations, anything where the result is a handful of rows.
The decision rule follows directly from that split. Any JDBC or ODBC connection feeding an analytical workload, where the result lands in a DataFrame, an Arrow table, or some other columnar pipeline, is a reasonable candidate for an ADBC swap. Connections feeding transactional application logic generally aren't.
Spark is a useful edge case here. ADBC integration fits naturally into Spark's Data Source V2 columnar read path, and it can help external accelerators like Comet. The engineer behind the March 2026 Spark feature request described the effect as "not as dramatic, but still noticeable. That's a fair summary of what to expect anywhere the consuming system itself isn't fully columnar.
How the zero-copy path applies to high-volume real-time and observability pipelines
Everything above treats the transpose tax as a one-time cost paid on a single pull. In real-time analytics and observability pipelines, the transpose tax is a tax paid on every single query in a continuous stream, and the cumulative cost over a day or a week dwarfs what any single-query benchmark would suggest.
Observability telemetry is naturally suited to this problem, because it's already columnar in shape before it ever touches a driver. Metrics, traces, and logs arrive as append-only event streams, where each row is a timestamped event, and the queries running against them aggregate across millions of those events at once. That's close to the ideal case for ADBC's columnar path: large results, columnar destinations, repeated constantly.
Dashboards firing multiple sub-queries in parallel, monitoring tools running continuous aggregations, risk engines checking data against thresholds in real time: all of these issue queries at high frequency, and a driver overhead that would be a rounding error on one pull becomes the dominant cost once it's repeated hundreds of times a minute. The storage-layer finding from the DaMoN '26 research applies here almost without modification. Removing that decoding step through Arrow-native connectivity frees up CPU budget that was previously spent on format translation, available instead for the actual computation the pipeline exists to do.
A fast engine isn't enough on its own. If the connectivity layer between the engine and the application still moves data through a row format, part of whatever speed advantage the engine bought gets spent again at the client boundary. A columnar engine that produces results faster than a row-oriented driver can transmit them has simply moved its bottleneck downstream instead of removing it. Any analytical database that stores data in columnar form internally and then has to serialize it into a row-oriented wire protocol for JDBC or ODBC is working against its own storage engine at exactly the moment results leave the building.
What changes when AI agents are the query issuer rather than a human analyst
Everything above assumes a human analyst running a query, waiting, and reading the result. AI agents change both halves of that assumption. An agent exploring a dataset, testing hypotheses, or building up a feature set doesn't run one query and read the answer. The per-query driver tax that was already compounding in high-frequency dashboards compounds faster still under agentic access, because the query frequency has no human pacing behind it.
The second change matters just as much. A human analyst reading query results eventually looks at a table or a chart, so the exact format the data arrives in is somewhat forgiving. An agent consuming results downstream, feeding them into another model call or another step in a pipeline, benefits from a typed, columnar structure it can process directly, rather than a row-oriented, loosely typed structure like JSON that needs its own parsing and interpretation step. For agentic workloads, the gain from ADBC is a result format built for a consumer that reads data by the millions of rows and acts on it immediately, not one that reads a handful of rows and hands them to a person, beyond just the removed transpose.


