BREAKING: ByteDance just dropped a paper that should make
🚨 BREAKING: ByteDance just dropped a paper that should make every GPU infra team and NVIDIA investor do a double-take.
An AI agent that now writes faster CUDA kernels than torch.compile the gold-standard compiler every ML engineer relies on.
Not close. 2.11× faster overall

