From Hot Storage to Cold Memory: Decentralized Storage Under the AI-Era Storage Boom
Author: 0xjacobzhao | https://linktr.ee/0xjacobzhao
Today, CXMT, officially debuted on the ChiNext board, sending shockwaves through the market with a staggering 500% surge. Despite lingering volatility from recent pullbacks across the broader memory sector, AI storage continues to undergo aggressive valuation repricing by capital markets, swept up in the prevailing tech narrative. In stark contrast, decentralized storage within the Web3 space has languished in prolonged dormancy and neglect. Why, despite both bearing the name “storage,” is there such a stark dichotomy in their market performance? Ultimately, the fundamental answer lies in the radical divergence of their underlying value functions.
The repricing of storage in the AI era is, at its core, a celebration of “hot-data efficiency”: it serves the maximization of compute utilization and commercial monetization. Decentralized storage, by contrast, defends the value proposition of “cold-data trust”: data fairness, censorship resistance, and the long-term memory of human civilization. The former is an efficiency system for hot data; the latter is a trust system for cold data. Today, capital markets clearly stand on the side of efficiency. Yet human civilization will ultimately still need a memory substrate that cannot be arbitrarily tampered with. The long-term value of trusted cold storage has not disappeared; it is merely dormant on the dark side of the cycle, waiting to be repriced by the next era.
Why Storage Has Become a Focal Point in the AI Value Chain
In the traditional IT era, storage was a “capacity business.” Enterprise CIOs cared about cost per unit of capacity, drive reliability, disaster recovery, archive strategy, and 3- to 5-year hardware refresh cycles. Storage was treated as an accessory to server procurement.
The current storage boom is not a conventional cyclical recovery. It is AI repricing the ability of data to flow. In the large-model era, the logic of storage has shifted from “capacity first” to “efficiency first,” with extreme focus on GPU feeding rate, checkpoint writes, and ultra-low-latency RAG. This marks the elevation of storage from “the final parking lot of data” to “the high-speed channel through which data enters computation.”
The evolution of resource bottlenecks in AI infrastructure is essentially a battle to fill the shortest plank in the barrel. Real compute utilization is not the linear sum of individual assets, but a strict multiplier effect: real compute utilization = GPU x HBM x DRAM x SSD x network x file system. A weakness in any link can collapse overall utilization. In the AI era, storage has for the first time moved from a “cost center” to an “efficiency engine.” This is the fundamental logic behind storage repricing.
AI Storage Architecture: From HBM as the Bandwidth Organ to the Data Lake Foundation
AI storage is not a simple pile of hardware. It is a tightly coupled, layered, and scheduled system. Within this system, industrial value and capital attention are highly concentrated around HBM, enterprise SSDs, SSD controllers, NVMe/CXL protocols, and high-performance storage systems. To clarify the value flow, we divide AI storage architecture into four core layers from top to bottom:
Near-compute memory layer (bandwidth core): HBM is the absolute core, supplemented by DRAM and CXL memory pooling. This layer sits close to the GPU/CPU package or bus, aiming to break the “memory wall” and forming the first gate that determines whether compute can be fully released.
High-speed persistent storage layer (I/O hub): The core formula is enterprise SSD = NAND flash + SSD controller + NVMe/PCIe data path. This layer supports high-frequency checkpoint writes, large-scale training dataset loading, and RAG hot-data caching, making it the clearest incremental persistent storage layer in AI data centers.
Low-cost mass-capacity storage layer (capacity foundation): This layer consists of HDDs, cold storage, and data-lake archive systems. In the face of exponentially expanding multimodal raw data, historical logs, and compliance backups, it still provides irreplaceable TCO advantages.
AI storage systems and data software (scheduling brain): This includes high-performance parallel file systems, distributed object storage, vector databases, and the RAG data governance layer. What AI actually consumes is not bare hardware, but data availability that has been efficiently organized, indexed, permissioned, and governed by the software stack.
As an ecosystem extension, decentralized storage does not directly enter the millisecond race of AI hot data. Instead, it anchors itself in public dataset notarization, AI training data provenance, and long-term cold-memory archives, establishing its unique niche as a “trusted cold layer.”
HBM: The “Bandwidth Organ” Closest to Compute in the AI Storage Chain
High Bandwidth Memory (HBM) is not traditional storage. It is a high-bandwidth memory layer near the GPU. Its core mission is not to preserve data, but to continuously feed compute with data at extremely high bandwidth. HBM is the most compute-adjacent and most certain segment of the AI storage chain. It directly determines whether a GPU can be “fed,” making it one of the most critical supply-chain bottlenecks today.
HBM’s core architecture is “3D DRAM stacking + 2.5D advanced packaging.” Through TSV vertical stacking and CoWoS heterogeneous integration, it compresses the distance between storage and compute and delivers a generational leap in bandwidth. Its industrial barriers are not only about DRAM design, but also DRAM process technology, TSV, ultra-thin stacking, packaging, thermal management, testing, and customer certification. A yield defect in any link may cause the entire HBM stack to be scrapped.
At present, only SK hynix, Samsung, and Micron can stably mass-produce high-end HBM globally, supported by a triple moat of top-tier DRAM process capability, packaging capability, and NVIDIA/AMD customer certification.
DRAM and CXL: The System Memory Foundation and Memory-Pooling Engine
HBM addresses the extreme bandwidth required near GPUs. DRAM forms the foundation of server system memory. CXL attempts to break physical boundaries and reorganize memory resources across data centers.
DRAM: DRAM primarily supports CPU-side caching, data preprocessing, intermediate-state buffering, and system operation. It is the most fundamental system memory layer in servers. The global DRAM market is highly concentrated among SK hynix, Samsung, and Micron, while ChangXin Memory Technologies (CXMT) is the core variable for domestic substitution in China.
CXL (Compute Express Link): CXL is a next-generation cache-coherent interconnect protocol for data centers. It aims to overcome the limits of traditional DIMM slots, local memory capacity, and fragmented server memory islands, pushing memory architecture toward expansion, pooling, and sharing. CXL is still in the early stage of moving from platform support toward scaled deployment, but it has meaningful long-term architectural value. Representative companies include Astera Labs and Montage Technology.
Enterprise SSDs: The Data Hub Built by NAND, Controllers, and NVMe
Enterprise SSDs are the most important high-throughput persistent storage increment in AI data centers. With high throughput, low latency, and stable QoS, they continuously feed data to GPUs across the full lifecycle of training dataset loading, checkpoint writing, RAG retrieval, inference caching, and log backflow.
In AI storage architecture, an SSD is not an isolated hardware component, but a tightly coupled system. It can be summarized by an industrial formula: enterprise SSD = NAND flash + SSD controller + NVMe/PCIe data path. These three layers represent distinct segments of the value chain:
NAND flash (raw material layer): determines storage density and unit cost, while the controller governs performance release and lifetime management. Representative companies include Samsung, SK hynix/Solidigm, Micron, Kioxia, Western Digital, and YMTC.
SSD controller (performance enablement layer): determines performance release, error correction, QoS stability, and wear leveling. Representative companies include Phison, Silicon Motion, Marvell, and Maxio.
NVMe/PCIe (data-path layer): determines the efficiency of data transfer from storage to compute. Combined with GPUDirect Storage, data can reduce CPU memory bounce buffers and CPU involvement, significantly alleviating the I/O bottleneck. Representative companies include Broadcom, Marvell, and Astera Labs.
HDD / Cold Storage / Archive: The Low-Cost Foundation of the AI Data Lake
AI will not eliminate HDDs. As multimodal models demand more video and image data, and as enterprise compliance logs and historical datasets expand exponentially, the demand for low-cost cold-data storage is rising in parallel. In AI storage architecture, SSDs and HDDs coordinate based on business value: SSDs handle hot data and high throughput, while HDDs handle low cost and long-term retention. Representative companies include Seagate, Western Digital, and Toshiba.
AI Storage Software Stack: The Scheduling Hub of Data Availability
AI never truly consumes bare disks. It consumes “data services” carefully organized by the software stack. This architecture turns underlying hardware into knowledge assets that AI can directly call. It has four layers:
High-performance storage systems (feeding systems): centered on concurrent throughput and low latency, they use parallel file systems to solve the “data hunger” of GPU clusters and ensure rapid data flow for training and inference. Representative companies include VAST Data, WEKA, and Pure Storage.
Object storage (raw data lake): organized around objects, keys, and metadata, object storage carries massive unstructured data. It does not pursue extreme low latency, but uses low cost and cloud-native characteristics to build the capacity foundation. Representative product: AWS S3.
Vector databases (semantic index layer): vector databases store, index, and retrieve vectors generated by embedding models, allowing AI to locate semantically relevant content within massive knowledge bases. Representative companies/projects include Pinecone and Milvus.
RAG data layer (knowledge invocation layer): beyond simple retrieval, it covers data chunking, cleaning, permission control, and citation provenance, ensuring that enterprise data can be safely, accurately, and traceably invoked by large models. Representative company: Databricks.
From AI Hot Storage to Decentralized Cold Memory: Efficiency Maximization vs Trust Maximization
AI storage is an efficiency-driven system. Its value function focuses on maximizing compute output. HBM bandwidth determines whether a GPU can be fed, SSD throughput determines dataset and checkpoint read/write efficiency, and low latency matters for real-time RAG and inference. These metrics ultimately converge into GPU utilization and unit token cost, directly determining the commercial profitability of AI applications. The ultimate goal of AI storage is not preservation, but acceleration; it serves productivity.
Decentralized storage has a fundamentally different value function. It asks whether data will still exist ten years from now, whether it has been tampered with, and whether it can resist single-point censorship. Through cryptographic proofs and distributed networks, it builds an open-access public data foundation for permanent preservation. Its ultimate goal is to defend the absolute authenticity and sovereign independence of data, serving fairness, censorship resistance, and civilizational memory.
AI storage is the “hot storage” that fuels future productivity. Decentralized storage is the “cold memory” that preserves undeletable historical records for human civilization. The former serves efficiency and pursues extreme speed; the latter serves trust and defends silent memory. The former determines how fast models run; the latter determines whether memory can be erased. Today, market mechanisms reward productivity efficiency. AI storage stands at the center of the storm, while decentralized storage appears to be undergoing valuation collapse and narrative drain in silence.
The Vision and Reality of Decentralized Storage
There are many decentralized storage projects, but by industry mindshare and ecosystem depth, the core representatives remain Filecoin and Arweave. Although both are labeled “decentralized storage,” their underlying architectural philosophies follow almost entirely different paths: the former uses market contracts to approximate AWS-like elasticity; the latter uses a one-time social contract to approximate the permanence of a library.
Filecoin: Filecoin uses PoRep and PoSt to build one of the most complete verifiable economic systems. It should not continue to compete with AWS in consumer-grade cloud drives. Instead, it should move toward AI data provenance, public dataset hosting, and compliance archiving, providing verifiable chains for model audits and copyright proof. Its necessary path is to package itself behind S3-compatible APIs and fiat payment, upgrading from a “cheap storage market” into “verifiable compute infrastructure.”
Arweave: With the narrative of “pay once, store forever,” Arweave uses Blockweave and SPoRA to incentivize miners to preserve and rapidly access as much historical data as possible, especially scarce historical data. Its best position is as the public memory foundation of humanity: preserving human-rights records, evidence of war crimes, cultural classics, legal and financial history, and permanent long-term memory for AI agents. Arweave’s value is not speed, but the ability to carry civilizational memory across cycles.
The difficulties faced by decentralized storage projects such as Filecoin and Arweave do not come from a wrong value proposition. They come from long-term mismatch among productization, retrieval experience, real demand, and token incentives. This reveals the enormous gap between hacker ideals and mainstream commercial adoption:
Supply-demand incentive mismatch: Early networks represented by Filecoin used tokens to rapidly expand capacity, but failed to build a sufficiently strong paid demand side. This led to massive capacity but insufficient utilization and paid conversion. The reward was for “I can store,” not “you need me to store.”
Missing enterprise-grade service capabilities: AWS’s moat is not hard drives, but a “data operating system” composed of APIs, SLAs, permission management, compliance auditing, and technical support. Enterprises buy peace of mind, not experimental infrastructure that requires them to manage keys and select nodes themselves.
Retrieval-experience shortcomings: “Stored” does not mean “retrievable with stability and low latency.” Distributed nodes, complex topology, and the lack of unified SLAs make decentralized storage difficult to use for AI hot-data workflows. It is better suited for trusted cold archives and data provenance.
Privacy and compliance gaps: Enterprise private data cannot simply be written into a public permanent network. The right to erasure and permanent immutability have natural conflict. Decentralized storage is better suited to public data and long-term archives than to indiscriminately hosting core private enterprise data.
Token economics amplifies cycles: In bull markets, financialization hides insufficient demand. In bear markets, falling miner ROI exposes commercialization weakness. Tokens can cold-start supply, but cannot automatically create demand or sustainable revenue.
Other decentralized storage projects mostly focus on specific ecosystems or vertical niches: Storj/Sia have weaker cross-cycle industry mindshare and Web3 narrative influence than Filecoin/Arweave; BNB Greenfield/Walrus are tied to specific BNB or Sui ecosystems; Celestia/EigenDA are data-availability (DA) layers, serving rollup transaction confirmation rather than long-term archive; 0G and similar AI/DA hybrid narrative projects attempt to integrate storage, data availability, compute, and AI agent settlement into an AI-native modular infrastructure, but their real demand, developer adoption, and commercialization loop remain to be proven.
Future Opportunities for Decentralized Storage: The Long-Term Pendulum Between Efficiency and Trust
During periods of technological surplus, capital chases efficiency aggressively, assigning extremely high premiums to assets such as GPUs and HBM. Decentralized storage, which advocates trust and fairness, naturally becomes marginalized. Yet the pendulum of history will not stay forever on the side of efficiency. Arbitrary bans and content deletion by super platforms, AI copyright lawsuits forcing proof of data sources, geopolitical conflicts over data sovereignty, the disappearance of public archives caused by data monopolies, and regulatory pressure to audit model training data could all create conditions for repricing “trusted storage.” The future opportunity of decentralized storage may still express itself in several differentiated directions:
AI data provenance: combining cryptographic proofs to build “data lineage proofs” in response to regulatory and audit pressure.
Public datasets and civilizational archives: anchoring censored archives and cultural heritage to build irreplaceable, undeletable memory.
Trusted archiving and compliance notarization: using hash-based attestations to provide trusted self-proof and high-grade digital notarization.
Integration with ZK / TEE / DID: easing privacy tensions and upgrading from a single “storage protocol” into “trusted data infrastructure.”
Invisible product route: providing S3-compatible APIs and fiat payment, allowing users to directly purchase “trusted archiving” services.
AI storage and decentralized storage represent two different instincts of the data world. One pursues extreme efficiency and provides the fuel for us to rush toward the future; the other defends silent memory and protects our right to look back. Today, the market rewards efficiency without hesitation, and decentralized storage therefore appears quiet or even collapsed. But when the AI era further amplifies data monopolies, copyright disputes, and the fragility of historical memory, decentralized storage may be re-understood as a “trusted cold layer.” Memories that cannot be easily erased by platforms, companies, or any single authority may evolve from idealistic romance and marginal belief into necessary infrastructure.
Disclaimer: This article was created with assistance from AI tools including Claude Opus 4.8, ChatGPT-5.5, and Qwen 3.7. The author has made every effort to review the content and ensure that the information is true and accurate, but omissions may still exist. Thank you for your understanding. This article is for information integration and academic/research exchange only. It does not constitute investment advice and should not be regarded as a recommendation to buy or sell any token.







