NVIDIA CMX KV cache offload creates write demands that standard enterprise SSDs weren't built for. ScaleFlux's AI-optimized SSD platform targets that gap with support for more than 200 Flexible Data ...
ScaleFlux is also developing a trace-driven simulator that models KV cache movement across GPU HBM, host memory, and SSD tiers. The simulator generates replayable SSD traces for evaluating placement, ...
Nvidia researchers have introduced a new technique that dramatically reduces how much memory large language models need to track conversation history — by as much as 20x — without modifying the model ...