| |
A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation
Researchers propose FIBER, a new GPU architecture that decouples execution threads from private register ownership to improve tensor computation efficiency for modern AI workloads. By enabling dynamic parallelism scaling and fine-grained register-level scheduling, FIBER achieves 2.25x end-to-end speedup on Ampere GPUs and up to 2.49x kernel-level gains, addressing bottlenecks from fixed parallelism and coarse-grained scheduling in current GPU designs.
Read Full Article →
← More Tech news