>"Alongside Sony, AMD has developed “Universal Compression”, which should help compress data within a GPU to deliver higher levels of memory efficiency. This reduces memory usage and memory bandwidth usage. This frees up memory bandwidth for other tasks. Furthermore, it effectively gives next-generation hardware more memory space to work with.
Note that Mark Cerny states that this compression technology will have “synergies” with AMD’s Neural Arrays and Radiance Cores. After all, smaller, more compressed/optimised data will fit better onto caches and optimise workflow.
Universal Compression – a system that evaluates every piece of data headed to memory, not just textures, and compresses it wherever possible. Only the essential bytes are sent, dramatically reducing memory bandwidth usage. This means the GPU can deliver more detail, higher frame rates, and greater efficiency. "
Great Idea!
By putting discrete, automatic Hardware Compression abilities into GPU's, it should be possible to compress regions of a GPU's memory, which should result in freed memory, which should result in an upgrade of capabilities at a given level of memory...
Observation: It will be interesting to see if automatic, general, hardware-based (which should be very fast!) compression on future GPU's (and potentially AI accelerators) can yield memory space savings when loading LLM's into memory while being performance compatible with older GPU's that do not do this.
If so, then we may have unlocked or "leveled up" our future Graphics Cards / AI Accelerators yet again...
Sort of like what DSpark did as a software optimization for LLM's, hardware-based GPU compression of graphics card / AI accelerator data could be a major win, and a major step forward, enabling future GPU/AI cards to have greater capabilities at the same base memory level IF the compression works for a given LLM's data!
Crossing my fingers that Universal Compression for future GPU's will work with LLM's!
A promising idea, to be sure!