How Agent Query Patterns Break Assumptions Built for Human Traffic
Agent query patterns shatter databases built for human traffic.

Agent traffic already outweighs the human traffic that most database systems were built to handle, and the gap is widening fast. Automated traffic grew 8× faster than human traffic year over year, and agentic AI traffic alone (systems that reason and act autonomously, not simple bots) grew 7,851% in the same period. Monthly AI-driven traffic volumes grew 187% between January and December 2025, nearly tripling inside a single calendar year. Gartner expects 40% of enterprise applications to embed task-specific AI agents by the end of 2026, up from a small fraction the year before. Uber's QueryGPT already handles around 1.2 million interactive queries a month, and OpenAI runs an internal data agent across 70,000 datasets totaling 600 petabytes, figures that put this well beyond the application layer. None of that is a pilot program. It's the shape of the workload every production data team is now planning around, whether they've noticed yet or not.
What agents do when they query a database, compared to what humans do
Start with the raw volume. MotherDuck's own query history analysis found agents running 29 times more queries than humans within a single observed month, a gap that has roughly doubled month over month, with the average organization now running twice as many agents as human users. That's not a marginal shift in traffic composition.
Now look at pace. The median gap between queries is 4 seconds for an agent and 60 seconds for a human, a 15× tighter loop that produces far burstier load profiles. A human clicks, reads, thinks, clicks again. An agent doesn't read anything. It observes a result, decides on the next step, and fires again, often within the same breath as an internal reasoning cycle. That loop is called OODA (observe, orient, decide, act), and agents run it continuously; in multi-agent systems, every tool inside the workflow runs its own nested version of that loop, in parallel with the others.
One prompt from a person can trigger hundreds of internal database operations underneath it. Each reasoning step an agent takes might involve a context read, a write of intermediate state, a memory lookup, a tool call, and a logging entry, stacked one after another. Datadog's 2026 State of AI Engineering report found a notable share of agentic application requests making three or more service calls each, and each of those calls can fan out into its own database hit. A single user question doesn't map to a single query anymore. It maps to a tree of them.
The shape of that tree depends on which design pattern the agent is running. ReAct-style agents alternate reads and writes in tight, high-frequency loops. Plan-and-Execute agents front-load a burst of exploratory reads during planning, then concentrate writes during execution. Multi-agent systems fan out in parallel, with several agents hitting the same tables at the same moment. Reflection-based agents re-read the same data repeatedly as they critique and revise their own output. Humans have natural governors on their query rate, including cognitive load, the time it takes to read a result, and the pause between navigating screens. Agents have none of that. There's no fatigue, no reading time, no reason to slow down.
Where connection pooling and concurrency models stop holding
Connection poolers were sized for humans. Engineers sizing connection pools for humans picked a number based on expected concurrent sessions, expected idle time between queries, and a rough sense of how many people would be logged in at once, whether or not they realized that assumption was baked in.
Agent workloads violate all three inputs: more concurrent initiators, far shorter idle periods (4-second gaps vs. 60-second gaps), and bursts that spike without the gradual ramp human traffic produces. That last part changes how the spike shows up: it can appear in the space of a single reasoning step instead of ramping up over minutes. A human traffic spike ramps up over minutes as people log in throughout a morning. An agent traffic spike can appear in the space of a single reasoning step.
When multiple agents complete a reasoning step at roughly the same moment, they all issue their next query at roughly the same moment, and they hit the database together, creating a thundering herd directly. It's the same pattern that has overwhelmed CDN origins for years, except now it's happening one layer deeper, at the query layer, where there's often no cache in front to absorb the shock. In multi-agent architectures, this gets worse because each tool inside a workflow may hold its own connection open. A moderately complex agentic system can exhaust a pool that was sized correctly, on paper, for the human workload it replaced.
This isn't a hypothetical risk sitting in a whitepaper somewhere. AI-related incidents rose sharply from 2024 to 2025 according to the AI Incidents Database, with monthly incident counts climbing steadily over several years, and infrastructure failures tied to agentic concurrency now make up a measurable, growing slice of that total. Some mitigations help directly: per-agent connection limits, queue-based admission control so bursts get smoothed rather than dropped, and isolated compute reserved per agent class so one workflow's storm doesn't take down another team's queries. None of those are a complete fix on their own. They buy headroom. The deeper architectural answer needs a different foundation entirely, and that appears later in this discussion.
Why caching hierarchies built for repeated human queries fail under agent exploration
Caching works on a simple bet: the same query, or close to it, will come back around.
Agents don't behave that way. They explore. Every reasoning step tends to produce a slightly different query than the last one, a different filter, a different time window, a slightly different slice of the same table, because the agent is probing for information it doesn't have yet rather than re-checking something it already knows. That's the opposite of the repeat-access pattern caches are tuned for, and hit rates collapse under it. The Reflection pattern offers a partial exception, since an agent re-reading its own prior output to critique it can produce genuine cache hits, but those hits sit interleaved with a stream of novel queries that thrash the cache in between. A little bit of reuse surrounded by a lot of noise doesn't move the average much.
A structural floor underlies all of this, and it produces real latency costs that are easy to miss until they hit someone directly. That sounds small until it's stacked. An agent running a 4-second query loop, with fan-out across several hops per reasoning step, hits that floor over and over, and each hit adds directly to what the end user experiences as delay. There's no way to average that away when the loop is that tight.
Which is why NVMe cache-mesh approaches that keep hot data in memory, never routing reads through S3, are architecturally necessary, not optional, for agentic sub-second requirements. For agentic workloads that need sub-second responses, it's a baseline requirement. That's a genuinely different design posture than the one caching hierarchies were built around for human traffic, where a slower cold path was an acceptable fallback rather than a constant tax.
How fan-out and p99.9 latency expose the limits of human-grade reliability targets
Think about what fan-out does to probability. If a single query has some small chance of hitting a slow response, that chance looks negligible in isolation. But an agent reasoning step doesn't fire one query. It fires several, often in parallel across multiple hops, and each of those calls carries its own independent shot at hitting the slow tail. Stacking enough of those calls together makes what looked like a rare event visible on a routine basis, a regular feature of the user experience rather than an edge case.
User experience in agentic systems depends on p99.9 and p99.99, so engineering teams focusing on p99 response time are under-targeting for production AI.
The variance driving that tail comes from sources that are each individually forgivable, including a cache hit rate that dips for a few seconds, a garbage collection pause, a compaction or vacuum cycle running in the background, or a hot shard getting rebalanced. None of those, on their own, would bother a human user clicking through a dashboard once a minute. Stacked across the many hops an agent's reasoning step touches, they compound.
Plan-and-Execute agents make the exposure especially visible. They issue a burst of queries during the planning phase before any action gets taken, and a single slow response inside that burst delays the entire plan before execution even starts. One tail-latency hit, early in the sequence, and the whole downstream chain waits on it. Systems that deliver flat, predictable latency regardless of access pattern absorb that kind of fan-out far more gracefully than systems that are fast on average but wobble at the tail. Average speed was always a proxy metric. For agents, it stops being a useful one.
Where query cost models break down when agents run short, high-frequency workloads
Billing granularity is where this appears first, and it hits hard. MotherDuck's analysis shows that for an agent firing many short queries throughout the day, a per-minute billing minimum can turn into a dramatic cost multiplier compared to per-second billing. The query volume agents generate multiplies the gap between what was actually used and what gets billed, widening it fast. Agent cadence exposes the discrepancy directly.
S3 pricing carries a similar trap, just spread across more dimensions than most teams track. For an architecture routing continuous agent reads through S3, egress alone can end up dominating the total bill, not storage. Most teams size their cost model around the storage rate and quietly ignore request charges, retrieval fees, and data transfer, five additional cost dimensions altogether that frequently account for a large share of total spend. A Crayon study found 94% of IT leaders still struggling to get cloud costs under control, and agent-driven query volume is only going to make that harder.
Behaviors change alongside billing, too. Per-query pricing charges more the more an agent asks. It directly penalizes the exact behavior an agent exists to perform. Curiosity, probing, checking an adjacent slice of data just to rule something out, all of that gets taxed under a per-query model. Compute-hour pricing breaks that link entirely: cost tracks time spent processing, not the number of questions asked, so an agent can explore freely without triggering a surcharge for every additional thought. That's a meaningfully different incentive structure, and it matters more for agent workloads than it ever did for human ones. Running a production database in-house carries its own overhead, patching, backups, monitoring, failover, that self-managed setups routinely underprice against a managed option, and it's a cost most teams leave out of the comparison altogether. S3 pricing is multi-dimensional in ways teams consistently underestimate: AWS S3 Standard costs $0.023/GB-month, and data out to internet costs $0.09/GB for the first 10TB/month, $0.085/GB for the next 40TB, and $0.07/GB for the next 100TB, meaning that for S3-routed architectures serving agent reads continuously, egress alone can dominate the bill.
Why eventual consistency breaks when agents need a coherent view of shared state
Correct decisions depend on a coherent view of reality, and eventual consistency (the guarantee that data will converge eventually, just not necessarily right now) stops being good enough once the reader making decisions off that data is an autonomous system rather than a person. Humans absorb staleness without much cost. Someone glances at a dashboard, notices a number looks off, refreshes the page, and moves on with the corrected figure. An agent doesn't have that reflex. It reads once, acts immediately on what it read, and if that read was stale, the staleness propagates straight into whatever decision came next, with no human in the loop to catch the drift.
Multi-agent systems make this sharper still. Several agents can read and write shared state at the same time, and without strong consistency guarantees, one agent's write may simply not be visible yet to another agent's read happening in the same reasoning cycle.
Enterprise data fragmentation piles onto this problem. Answering one agent question can require pulling from several systems at once, each running its own consistency model on its own schedule, which means the coherence problem isn't confined to a single database. It's a property of the whole stack the agent has to reason across.
Specific storage engine behavior can quietly widen the gap. In systems using ClickHouse's MergeTree storage engine, DELETE and UPDATE operations execute in the background without blocking reads, which is architecturally useful for throughput, but agents running tight read-after-write loops may read pre-mutation state. Frequent small inserts carry a related risk: too many small parts written in quick succession can trigger merge storms and rising insert latency, which directly undermines the freshness of the data an agent reads back moments after writing it. Isolation offers a partial guardrail here. Giving agents read-only compute over shared data, separated from the write path, at least keeps the consistency risk of an agent's reads from turning into a corruption risk on the data everyone else depends on. That same isolation principle carries directly into the access control problem that follows. 72% of organizations store data in disparate silos, and 82% report those silos disrupt critical workflows, meaning that answering a single agent question often requires integrating across several systems, each with its own consistency model, and enterprise data fragmentation compounds this.
How access control models built for human roles fail when agents are the users
Traditional role-based access control was built around a simple chain in which a user belongs to a role, the role defines what that role can touch, and a specific human is accountable for what happens under it. That chain assumes a person on the other end of every session, someone who can be asked why a query ran.
Agents break that chain in a way permissions frameworks weren't built to absorb. An agent might act on behalf of several users at once, or none in particular, and its queries can span sessions, time zones, and data domains that no single human role was ever scoped to cover. Fan-out compounds the exposure: a read-only agent running thousands of queries a day, all under one permission grant, has a far larger effective reach into the data than a human analyst carrying the exact same role ever would. The permission looks identical on paper. The blast radius isn't.
Safety has to get built into the access layer directly rather than assumed from good behavior. That means hard-coded blocks on destructive operations, approval gates in front of schema changes, and full audit logging of every statement an agent executes, three things that human-facing tools rarely ship with all at once by default. None of that is about restricting what agents can do. It's about making sure what they did do is visible after the fact.
Accountability, not restriction, is the right frame here. Every agent query logged in full, tagged with the agent's identity, the exact query text, the data it touched, and the timestamp, gives engineers the same forensic trail that application logs have always given for human-triggered bugs. Debugging an agent's behavior without that trail means guessing. Sandboxed branches extend the same logic one step further: an agent gets isolated, read-only compute over a copy of the shared data, runs its experiments there, and none of it writes back to production. Schema changes and pipeline operations an agent proposes should get expressed as reviewable, version-controlled code rather than freeform statements executed directly, so a human can look at exactly what's about to run before it runs.
What infrastructure has to look like to make agent access reliable and auditable at scale
Putting the pieces from every section before this one together produces a fairly specific shape. Latency has to stay flat at p99.9 and beyond as data volume grows, not just fast on average, because fan-out math turns rare tail events into routine ones the moment an agent's reasoning step touches multiple hops at once. Compute has to sit separate from storage, with the hot data actually living in memory or on NVMe rather than routing every read through a system with a built-in hundred-plus-millisecond floor.
Pricing has to track compute time rather than query count, so an agent's habit of asking one more question doesn't turn into a line item working against it. Consistency guarantees have to be strong enough that an agent reading data seconds after another agent wrote it gets the current version. And access has to carry full identity, full logging, and isolated experimentation space by default, not bolted on after an incident forces the question.
None of these requirements is exotic on its own. Taken together, they describe a system built for a user that never sleeps, never paces itself, and never tolerates a stale answer without acting on it first. This is what the user database infrastructure is now serving, and the workload model has not caught up to that fact yet.
Sources
- Best Analytics Database for LLM & AI Agents (2026 Guide) | MotherDuck
- Can AI Agents Answer Your Data Questions?A Benchmark for Data Agents
- Measuring the AI-Driven Internet with The 2026 State of AI Traffic & Cyberthreat Benchmark Report - HUMAN Security
- Agents Don't Query Like Humans Do | MotherDuck
- When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers


