Bitcoin Brief

AI Just Got 10x More Efficient. It Will Need More Power.

Vera Rubin produced up to 10x more DeepSeek-R1 tokens per megawatt. Cheaper intelligence will increase total power demand and reprice every chip already installed.

10 min read
AI Just Got 10x More Efficient. It Will Need More Power.
Share
TFTC - Truth for the Commoner

Bitcoin Brief

Sup, freaks.

Yesterday we looked at the Henry Adams Curve and America's need to build more energy. This morning NVIDIA gave us a useful reminder that the demand side is not standing still.

The megawatt just got more productive. It did not get less valuable.


LEAD STORY

AI Just Got 10x More Efficient. It Will Need More Power.

CoreWeave turned on NVIDIA's newest rack-scale AI system and produced a result that should force every datacenter operator, utility planner, chip lessor, and AI investor to revisit the assumptions sitting inside their spreadsheets.

I think the second-order effects matter more than the headline.

On a DeepSeek-R1 inference workload, CoreWeave reports that Vera Rubin NVL72 delivered up to 10 times more tokens per second per megawatt than GB200 NVL72. The comparison held user responsiveness constant.

That last point matters. A datacenter can produce impressive throughput by forcing users to wait. CoreWeave compared the systems at matched interactivity, measured as tokens generated per second for each user. At that target, Rubin produced dramatically more useful output from the same power envelope.

For a fixed DeepSeek-R1 workload, this is a real energy-efficiency gain. If the result holds across deployment, the same volume of inference at the same responsiveness could require roughly one-tenth the power.

That does not mean the AI industry will consume one-tenth the electricity.

It means the industry can produce much more intelligence from every megawatt it secures.

Rubin's gain comes from a complete system, not one magical transistor. Each NVL72 rack combines 72 Rubin GPUs, 36 Vera CPUs, HBM4 memory, NVLink 6 fabric, BlueField-4 infrastructure processing, liquid cooling, and NVIDIA's software stack. CoreWeave used expert parallelism, NVFP4 precision, multi-token prediction, disaggregated prefill and decode, TensorRT-LLM, and Dynamo.

The result comes from hardware and software codesign at rack scale. Dropping a Rubin GPU into an ordinary server does not recreate it.

It is also a narrow benchmark. CoreWeave is an NVIDIA partner selling the capacity. The test covers one reasoning model and inference, not training. Public materials do not fully disclose the request mix, context length, concurrency, output-quality controls, absolute power draw, or whether the megawatt includes cooling, networking, storage, pumps, and other facility overhead. No independent operator has replicated it publicly.

Use the phrase "up to 10x." Do not turn it into a universal constant.

The measured direction is still important. Reasoning models do not improve only by growing their parameter counts. They can spend more compute after receiving a question. They generate longer reasoning traces, search multiple paths, call tools, check results, and revise their answers. Lowering the cost of those tokens makes more test-time computation economical.

Agents multiply the effect. An ordinary chatbot waits for a human prompt. An agent can run continuously, retrieve information, call software, monitor a system, coordinate with other agents, and report only when something changes. One human request can spawn a tree of machine activity.

NVIDIA says agentic applications can consume up to 15 times more tokens than conventional AI applications. A 10x hardware-efficiency gain can disappear inside a workload that generates 15x more tokens.

Jevons paradox is now reaching machine intelligence. When the cost of producing something falls, people consume more of it. Better steam engines did not end coal demand. Cheaper bandwidth did not reduce internet traffic. More efficient chips did not stop the growth of computing.

Bitcoin mining already ran this experiment. I wrote last night that mining moved from CPUs to GPUs to FPGAs to purpose-built ASICs. An Antminer S9 used approximately 100 joules per terahash. The S21 XP Hydro operates at 12 J/TH, an efficiency improvement of more than eight times.

Bitcoin mining ASICs and AI GPUs do different work, but the industrial progression is the same. Turn electricity into valuable computation, then relentlessly lower the energy cost of each unit of output. Those mining gains did not cap aggregate demand. They lowered break-even thresholds, attracted more capital, made more power projects economical, and intensified the search for cheap energy.

AI will follow the same path.

Better AI infrastructure can lower energy per token while raising total energy consumption.

The condition is straightforward. Total energy use falls only if demand grows more slowly than efficiency improves. Nobody knows that elasticity yet. The explosion of coding agents, research agents, voice systems, video generation, autonomous operations, and personalized software suggests demand will put up a fight.

This changes what operators should measure. Raw GPU count is becoming less useful than tokens delivered within a power, latency, and quality envelope. A 100-megawatt site running Rubin can produce more saleable output than the same site running Blackwell. That makes the site more valuable. It can strengthen the case for securing the next 100 megawatts rather than eliminate it.

Power remains the denominator.

Rubin also changes the economics of chips already installed. Blackwell does not become obsolete. It is deployed, CUDA-compatible, financed, contracted, and capable of running training, fine-tuning, and inference for years. Supply constraints will keep multiple generations working. Different models will receive different benefits from Rubin.

But technical life and economic life are not the same thing.

A Blackwell system may continue running perfectly while losing the premium reasoning workloads to Rubin. Rental rates can fall. Utilization margins can compress. Hopper can move further into batch inference, smaller models, research, and workloads where power is cheap. Blackwell-heavy fleets with long depreciation schedules and short customer commitments face more risk than fleets protected by take-or-pay contracts.

Every datacenter model should separate the chip's ability to function from its ability to earn an attractive return after power and price competition.

The capex boom also has a credit layer. Bloomberg macro strategist Simon White reports that hyperscalers have issued more than $500 billion of gross debt over roughly the last decade, with issuance accelerating as AI infrastructure expands. His leverage measure is at a high while record-low single-stock correlation suppresses index volatility.

That calm can be deceptive. These companies increasingly share one giant factor: the belief that demand for GPU-powered compute will remain insatiable. White's simplified stress test estimates that the Nasdaq VIX could move above 36 from approximately 28 if implied correlation returned to 60%. The regression explains only 34% of the variation, excludes operating leases, and uses S&P correlation as a Nasdaq proxy. It is a stress test, not a forecast.

The warning is sound. AI can become more efficient at the workload level while becoming more leveraged and fragile at the portfolio level.

Yesterday's energy argument still stands. Efficiency is not a substitute for abundance. Better hardware increases the output and potential revenue attached to each megawatt. It lowers the cost of applications that could not exist before. It makes secured generation, transmission, substations, cooling, and interconnection more valuable.

Rubin does not cancel the grid buildout. It pulls the demand curve forward.


SIGNAL

ENERGY / NATURAL GAS

America May Be Promising the Same Natural Gas Molecule Three Times

Matthew Smith's model warns that LNG exports, new gas-fired generation, and AI datacenters may claim more future U.S. gas than producers and pipelines can deliver by 2028. Its strongest insight concerns infrastructure, not geology. Gas underground is useless to a power plant without production, processing, firm transportation, and storage.

The forecast is aggressive. Chronometer assumes roughly 35.7 Bcf/d of LNG nameplate demand by 2030, while EIA projects approximately 27.7 Bcf/d of export capacity. Michael Spyker counters that $4 to $4.50 gas can unlock Tier 2 Marcellus inventory, deeper Permian zones, refracs, other basins, and Canadian imports. High U.S. prices would also suppress LNG demand.

Working storage will not literally reach zero without prices, drilling, curtailments, project delays, or politicians reacting first. The stress test still matters. Anyone underwriting gas-fired compute should identify the molecule, processing plant, pipeline capacity, storage, and firm contract. A turbine without fuel is not firm power.


ARTIFICIAL INTELLIGENCE / INFRASTRUCTURE

Culture Is an AI Infrastructure Constraint

SemiAnalysis reports that fragmented ownership and short-term incentives pushed Meta toward expensive systems its own model teams did not always want. Its examples include storage-heavy Grand Teton designs, Ariel with a reported 14% total-cost premium, DSF with a reported 11% premium, and integration problems following the Rivos acquisition.

The internal cost models and personnel accounts are SemiAnalysis reporting, not independently established facts. The broader lesson is stronger than any single percentage. AI infrastructure is not simply chips plus power. Model teams, networking, storage, software, compilers, cooling, and datacenter operations have to optimize for the same workload. A technically elegant component can destroy value when it solves the wrong problem or shifts costs elsewhere.

Meta can buy accelerators and acquire chip talent. The harder task is creating an organization that turns those inputs into useful, reliable compute. Culture sits inside total cost of ownership.


ARTIFICIAL INTELLIGENCE / TOKENIZATION

Gigatoken Speeds Up AI's Front Door

Models do not read words. A tokenizer breaks text into pieces and maps each piece to a number before the model sees it.

Stanford researcher Marcel Rød released Gigatoken, an MIT-licensed CPU implementation that replaces regex-heavy preprocessing with hand-written state machines, SIMD instructions, caching and low-contention parallelism.

I did not ask Martin, the AI agent we run at TFTC, to benchmark it. I sent him the link. He recognized that our server could run the software, downloaded the package and a 200 MB Stanford OpenWebText sample, built a comparison against Hugging Face, and validated the outputs across 20,401 documents.

On our 16-core AMD EPYC server, Gigatoken processed 2.33 GB/s and ran 59 times faster than Hugging Face. Martin did not merely summarize the author's benchmark. He autonomously designed and ran an experiment, checked the result, and brought back original evidence before we published.

Gigatoken remains beta, and one raw tiktoken path has a known special-token bug. That does not make models reason 59 times faster. The autonomous test is more interesting.


BITCOIN / ETF FLOWS

The ETF Bid Has Flipped, Not Roared

The institutional bid has returned after the worst month in the TFTC ETF tracker's history. U.S. spot bitcoin funds absorbed $203.1 million on Tuesday, their sixth consecutive inflow session. The streak has brought in $930.4 million, while July has moved to a $630.2 million net inflow after June lost $4.5 billion.

The reversal is real, but the damage is not repaired. July has recovered only 14% of June's outflows, and Tuesday's demand was heavily concentrated. Farside shows BlackRock's IBIT accounting for $163.9 million, or 81% of the day's net inflow. Bloomberg macro strategist Simon White argues that washed-out sentiment, poor trailing returns, improving ETF flows and bitcoin's recent strength against falling stocks create a favorable setup. The bid is back. Its breadth and durability still need to prove themselves.


ARTIFICIAL INTELLIGENCE / COLLABORATION

Buzz Gives AI Agents a Seat at the Table

Block's Buzz is an open-source workspace designed around humans and AI agents sharing channels, tools, project memory, signed identities, work history, and an audit trail. Slack and GitHub bolt agents on as bots. Buzz gives them the same basic participation model as human teammates.

The sovereignty claim needs precision. Each workspace currently relies on one authoritative relay. There is no peer-to-peer event exchange, gossip, or replication across independent relays. Buzz is self-hostable and identity-portable, not fully decentralized.

I have already downloaded Buzz, onboarded without much friction, and talked with agents inside the workspace. Today I am connecting Martin, our Hermes agent, and asking the TFTC team to onboard. The plan is to coordinate our development projects in Buzz, keep the work attached to open Git repositories, and use shared channels and the audit trail to manage handoffs between humans and agents.

The test is no longer whether the demo works. We are going to put real TFTC development through it and find out whether an agent-native workspace improves how the team ships software.

Sponsored

CASH APP

Cash App lets users buy, earn, send and spend bitcoin.

For a limited time, new customers can get $21 added to their balance. Just use code TFTC10 when you sign up, and send at least $5 to a friend in the first two weeks. Terms apply. Bitcoin services by Block, Inc. See the Bitcoin disclosures at cash.app/legal/podcast.

Get started with Cash App
Sponsored

UNCHAINED

Protect your bitcoin with collaborative custody from Unchained. Hold your own keys, remove single points of failure, and build a setup that can survive real life.

Watch: The Age of Debasement

⚡ FREEDOM TECH CORNER

bitcoin Does Not Care Whether the Buyer Is Human or Machine

Why it matters: Agents need native money, not bank accounts built around human identity and permission.

Lightning Labs released Wavelength, an alpha toolkit that gives applications and AI agents self-custodial bitcoin wallets through web and mobile SDKs, REST and gRPC APIs, a command line, and MCP. Users hold their own keys, can exit on-chain unilaterally, and can send interoperable BOLT 11 Lightning payments while Loop manages liquidity.

The public testnet and signet launch is real. Mainnet remains invite-only. Stablecoins and broader mobile support are future work. Users pay ordinary routing and on-chain fees plus a one-basis-point alpha service fee. Public documentation also leaves important Ark-style coordinator and VTXO-expiry questions incomplete.

The direction matters. Software agents can already communicate, retrieve information, and perform work. Wavelength gives them a path to receive and spend bitcoin without asking a bank or card network to recognize them as legal customers.


DATA SNAPSHOT

As of July 22, 2026, approximately 8:12 a.m. ET

Bitcoin price$65,849
Bitcoin market cap$1.32T
Bitcoin dominance56.7%
Sats per dollar1,519
Network hashrate887 EH/s
Recommended fee2 sat/vB
Block height959,126
ETF net flow, July 21+$203.1M
Current ETF streak6 inflow sessions

Join the TFTC Roundtable

The Roundtable is where builders, operators, founders, and curious freaks work through AI, bitcoin, markets, and company-building in real time.

Join the Roundtable

⚡ Find wallets, mining hardware, privacy tools, books, and other products built for a bitcoin standard.
Browse BitcoinProducts.com

See you tomorrow.


Marty Bent on X: https://x.com/MartyBent?ref=tftc.io

TFTC on YouTube: https://www.youtube.com/@TFTC

TFTC Podcasts: https://www.tftc.io/tag/podcasts/

News and analysis, not financial, investment, legal, or tax advice. Figures and quotes are verified against primary sources where possible. See our editorial and financial disclosures.

Keep reading

All of TFTC

The Bitcoin Brief

Bitcoin, markets, energy, and the tech reshaping all three.

A daily brief on the freedom tech building a parallel economy, written for the curious and the convicted alike. Signal, not noise. Truth for the Commoner.

Free, daily. Unsubscribe anytime.