Fusing a 27B ternary LLM's whole decode step into one CUDA kernel(twitter.com)3 points by Jr23_xd 4 days ago | 0 comments