More tokens, more chips! Kimi K3 demonstrates the "Jevons Paradox": increased efficiency instead leads to greater resource consumption.
Kimi K3 brings AI investors back to a core question: Does cheaper, more efficient models mean lower demand for chips and memory? Citi and Bank of America Securities both lean toward “no.” Efficiency improvements may unlock more usage, driving up total resource consumption.
According to Chasewind Trading Desk, Moonshot AI's Kimi K3 has 2.8 trillion parameters, supports 1 million token context windows and long-cycle intelligent agent tasks. Citi semiconductor analyst Peter Lee believes even if Kimi K3 is widely used, general memory demand for server DDR5 and eSSD will still increase.
Bank of America Securities semiconductor analyst Vivek Arya also pointed out in a July 17 report that top US AI labs are more likely to increase compute power investment, not decrease it. If Chinese open-source models continue to close the gap, OpenAI, Anthropic, and Google will need even larger training scales, heavier inference, and faster product iteration to maintain differentiation.
This means market pricing logic for Kimi K3 shouldn’t just look at single inference cost reductions. More crucial is whether low-cost models bring more calls, longer task chains, and more token generation. If the answer is yes, GPU, HBM, DDR5, eSSD, high-speed networks, and inference systems may all continue to benefit.
Low Cost, High Performance, Triggering the “Jevons Paradox”
Citi sees Kimi K3 as a potential case of the “Jevons Paradox” in the AI industry chain. The core of this paradox: when technology efficiency reduces unit cost, usage may increase significantly, resulting in total resource consumption rising rather than falling.
Kimi K3’s appeal comes from the combination of low cost and high performance. Public pricing shows cache-hit input costs $0.3 per million tokens, and output costs $15 per million tokens. Its complete weights are scheduled for disclosure on July 27, targeting long-cycle intelligent agent work.
Citi’s view is that lowering model prices won’t automatically reduce hardware demand. On the contrary, lower costs may boost willingness among developers and enterprises to call models, driving more AI agent deployments. Agent tasks aren’t one-off Q&A; they continuously generate, read, and process tokens during tasks. As call count and task length rise, lowered unit costs are translated back to higher total resource consumption.
Therefore, the impact of Kimi K3 on the semiconductor chain focuses not on “whether single call is cheaper,” but rather on “whether the total token volume expands.” This is why Citi favors server DDR5 and eSSD demand.
Long Context Inference Transfers Pressure to Memory
Kimi K3 uses three key technologies: Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. Kimi Delta Attention lowers the cost for 1 million token context windows, Attention Residuals selectively retrieve representations across different model depths, Stable LatentMoE increases sparsity, activating only 16 of 896 experts per token.
According to Moonshot’s technical notes, this architecture gives Kimi K3 2.5 times the scaling efficiency of K2. High sparsity and long context ability are key foundations for lowering operating costs.
But Citi stresses this does not mean inference resource pressure vanishes. Kimi K3 is still not a lightweight deployment solution; it requires multi-node clusters and super-nodes with over 64 GPUs. More importantly, long context and agent tasks ramp up KV Cache usage, increasing inference-side memory burden.
KV Cache demand is directly related to server DDR5 and eSSD. DDR5 deals with frequent data access, while eSSD benefits from larger cache and storage demands. For memory manufacturers, the key variable isn’t whether a model saves compute, but rather how many times the model is called post-price-cut, how many tokens are generated per call, and how long agent task chains are.
Model Convergence Pushes Up the Compute Bar
Bank of America Securities reaches similar conclusions to Citi but focuses more on GPU and AI infrastructure. Vivek Arya believes that stronger Chinese open-source models won’t necessarily cause US AI giants to cut compute investment. Rather, as the model gap closes, the cost to maintain an edge increases.
Bank of America Securities notes that if open-source models keep closing the gap, OpenAI, Anthropic, and Google must rely on larger training, more reinforcement learning and synthetic data loops, heavier inference, and faster product release cycles to preserve differentiation.
This logic doesn’t depend on which model leads temporarily. Large model rankings may rotate quickly, but enterprise buyers are really shopping for stable, low-latency, highly available AI output, and lower “effective output per unit cost.” As model capabilities converge, competitive pressure transfers to underlying infrastructure—including GPU, HBM, high-speed networks, and inference systems.
Therefore, Bank of America Securities thinks, models like Kimi K3 shouldn’t simply be seen as “models are more efficient, so fewer chips.” The more realistic path is intensified model competition, as leaders continue increasing compute investment to maintain their lead.
MoE Architecture Changes Bottlenecks: VRAM and Interconnect Matter More
Kimi K3 uses an MoE architecture, which separates total parameter count from the actual parameters activated each inference. This change eases some compute pressure, but also shifts infrastructure bottlenecks from pure compute power to VRAM access, expert routing, response latency, and interconnect capability.
Bank of America Securities stresses that MoE inference requires systems to move data more efficiently, schedule expert modules, and connect bigger compute clusters. Nvidia claims modern MoE inference needs larger GPU domains. Their measurements show GB300 NVL72 offers up to 25 times better per-watt performance on leading open-source models versus the Hopper platform.
CoreWeave’s tests around Kimi K2.6 also show that open-source MoE models, even if activating only part of parameters each time, still need optimized Nvidia GB300/GB200 NVL72 infrastructure to stay ahead in speed and cost-performance.
This indicates that model weights may become gradually commoditized, but GPUs, high-bandwidth memory, network interconnects, and inference systems required for running models will retain value. On the contrary, as model call volumes expand, these components may become even more competitive focal points.
Token Usage Still Expanding, Demand Not Yet Peaking
Bank of America Securities also cites OpenRouter data, showing model call volumes are still increasing rapidly. As a third-party API platform connecting multiple model labs, token usage on OpenRouter keeps growing, and Chinese AI lab models have already surpassed non-Chinese models in token usage.

This data doesn’t directly represent all enterprise and consumer scenarios, but at least shows developers’ choices in the open model ecosystem are shifting. Falling model prices and rising capabilities may drive ongoing expansion in call volume, rather than offsetting it through single-model efficiency gains.
Paid enterprise usage is also rising. Ramp data shows, as of June 2026, about 55% of US firms have paid for AI model, platform, or tool subscriptions—much higher than the US Census Bureau BTOS survey's estimate of 21%. By model, Anthropic’s enterprise adoption rate is 42.4%, OpenAI’s is 39.5%.
However, AI spending remains highly concentrated. The top 1% of enterprise users spend about $4,833 per employee per month on AI, the top 10% spend $516, while the overall median is just $11. Subscription rates in tech and media hit 79.8%, large enterprises 65.5%, higher than medium enterprises at 61.3% and small firms at 48.7%.

This data supports the judgment: AI usage is still in the diffusion phase. If low-cost models lower entry barriers, future incremental demand could come from more enterprises, developers, and agent applications.
Shift from “More Efficient Compute” to “Greater Workloads”
Citi and Bank of America Securities both agree that Kimi K3’s significance is not whether a single model reduces per-inference cost, but whether low cost unleashes larger scale workloads.
For Citi, the most direct beneficiaries are server DDR5 and eSSD. Long context, KV Cache, and agent tasks raise memory access and storage demand. For Bank of America Securities, the focus is on GPU, HBM, high-speed networks, and inference systems, since model competition pushes leaders to keep investing in infrastructure.
Bank of America Securities also notes that Kimi K3’s “45nm open-source EDA” shouldn’t be viewed as replacing commercial EDA. Rather, this shows chip design still heavily relies on EDA tools. In advanced processes, commercial EDA firms like Cadence and Synopsys remain pivotal.
The risk lies in another scenario: if model compression, inference optimization, and hardware efficiency accelerate faster than new workload growth in the long run, expansion of AI infrastructure may temporarily cool off. But under the current framework of both investment banks, Kimi K3 looks more like a catalyst for increased usage, rather than a signal for declining chip demand. After efficiency gains, the focus should be on more tokens, more inference, and more fundamental hardware consumption.
~~~~~~~~~~~~~~~~~~~~~~~~
The excellent content above is from Chasewind Trading Desk.
For more in-depth analysis, including real-time interpretations and frontline research, please join【Chasewind Trading Desk▪Annual Membership】
Risk Warning & DisclaimerThe market carries risks, and investing requires caution. This article does not constitute personal investment advice, nor does it take into account the specific investment objectives, financial situations, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their particular situations. Investing based on this is at your own risk.